
The tweets published in a state of intoxication can be calculated automatically , and the data on the location of drunken users will help improve the healthcare system .
The dissemination of information on the Internet is often compared with the spread of infections: “Viral video”, “Viral marketing”, “Media viruses”. But also real physical diseases also leave traces on social networks. The field of epidemiological studies using open Internet data gains more and more popularity and even gained a separate name-inframiology.
Among the latest achievements of inframiology - the definition of postpartum depression on a change in activity on Facebook [1] and the prediction of when the user is infected by the flu, based on the analysis of the tweet of his friends and neighbors [2]. Researchers from the Rochester University applied the methods of inframiology to the process of alcohol consumption and presented a number of curious observations [3].
As the initial data, all Twitter social networks were taken for the year, which are binding to the map in New York or in the Monroe district. Researchers needed to solve two main tasks: to allocate tweets related to alcohol use, and to determine exactly where the user uses is at home or not, and if not, then at what distance from the house.
To exclude abstract thoughts about alcohol, observation of others, memories and plans for the future, relevant tweets were determined in three stages. At the first stage, records were determined that had at least some attitude to alcohol, then those where alcohol was used to be used directly by the user himself, and at the third stage, tweets were selected from the already selected, where the use is described in the present time.
At each of the stages, the same machine learning algorithm was used - the method of support vectors. Training and test samples consisted of manually analyzed records, and as parameters it was taken into account what words and emoticons contained a record. The typos in the analysis were not taken into account - a word written with an error was counted in the same way as correctly written. The sensitivity and accuracy of the automatic method obtained were quite high - both were more than 82% in each of the stages.
To solve the second problem - determining whether the tweet was sent from the house - several considerations were also used. The machine learning algorithm took into account how often the user writes tweets from this place, at what time the tweet is written and whether he has words like “house”, “sofa”, “TV”, “bathtub”, etc. According to estimates, all this made it possible to accurately evaluate the location of the authors of tweets, in 80% of cases the error was no more than 100 m.
Further, on the basis of the data obtained, a thermal map of the density of the tweets about the libations was compiled and on the same map, alcohol sales points were noted. It turned out that the share of Twitter users who drink at home in the city is higher than in the suburbs. This is despite the fact that the density of the bars in the New York city in question is much higher than in the suburbs in question-Monroe County. There, according to the study, a significant part of users are drunk at a distance of more than a kilometer from the house. In general, the share of drinking users in the city is higher, and the more points of alcohol sale, the higher the density of the “drunk” tweets in the adjacent territories.
The ratio of the results with reality remains a big question. Firstly, the sample-Twitter users-are very unconstructive, the distortions are known by age and social status. Secondly, it is not clear the direction of causal relationships: they drink a lot, because there are a lot of bars, or many bars, because they drink a lot. The work has great methodological value: from publicly available data using machine learning methods, it turns out that you can get plausible assessments even for such a non -trivial process as alcohol consumption. The use of ready -made open data significantly reduces the cost of the study.
The authors suggest that the developed approach can be used to study social components of alcoholism, and the patterns of clarified patterns can be used to prevent it - the text of the article even mentions an anonymous alcoholics as an example of an organization working with social factors. In addition, the authors believe that with the help of such an approach, one can study the spread of any other hidden states and identify “typhoid mary” - hidden carriers of certain diseases. Or beliefs. Still, ideas are similar to viruses.