Let's play a children's game and imagine ourselves as aliens from the distant planet Chi-2, who are not at all interested in our funny political passions. And let these “hidvams” be given 100 thousand tablets with unfamiliar words and numbers and ask them to figure out these numbers: did they arise on their own, or were they “a little invented”?
Perhaps they would do this: they would count how often the last digit appears in these tablets, because it changes more often than all other digits, especially the first. We would reason with caution: of course, we should not expect that all digits will occur exactly the same number of times - there will always be some spread around the average value (1/10 for 10 digits). But what kind of spread is possible and what kind is not? The spread is measured by the value of a (sigma), and it can be calculated using the school formula V(pq/N), where p = 0.1 is the probability of encountering the selected number, and q = 1 - p = 0.9 is the probability of encountering any other. The spread will be such that 2/3 of the results will not exceed a, and 99.7% of the results will remain within “3 sigma”.
This is approximately what the possible result looks like, obtained for the last digits using a mental “roulette”, throwing out random numbers from 100 to 2000, which was scrolled 100 thousand times (Fig. 1).
The green corridor in this and the following figures precisely shows the boundaries of “3 sigma”. (When modeling such graphs, one must very carefully use truncation of the fractional part of the random number to avoid artifacts).
But “3 sigma” is only a guideline, because you also need to take into account what happens with the remaining numbers: suddenly only one of them is an exception, albeit a rare one, but the rest are located in a normal way. For this purpose, about 100 years ago, the English statistician Karl Pearson came up with the x2 (chi-square) criterion: this value is greater, the greater the total sum of squared deviations from the expected value. This, of course, is much more accurate than judging by only one number, without noticing the behavior of the others. Using x2, you can also calculate the probability that the result obtained is pure chance. All this has long been known and applied almost without hesitation.
But can this be applied to elections? How to check? This is what the voting results look like in the 2010 Swedish parliamentary elections (Fig. 2).
The probability calculated by x2 that this is just an accident is 80% (but not 100%, because there are still not enough numbers in the narrower corridor of 1 sigma). But there is an 80% chance that this could happen.
Our Russian elections
Mr. Churov’s department, in obedience to the law, made available the election results for each of the nearly 100 thousand polling stations in Russia. True, it hid this data so deeply, so carefully divided it into 3000 tiny fragments (why do you think this was done?), that the average and even non-ordinary Internet user will not get to it, and if he does, it will only be to individual pieces and stop there. So you will either have to patiently put the puzzle together or ask for help.
At each polling station, a protocol containing about 20 numbers is filled out. But most of them are small numbers, some are just zeros. Of course, it is impossible to count such numbers, so we will take only those graphs where the numbers are the largest. There are five of them: 1) the number of voters in the lists; 2) the number of voters who received ballots at the polling station; 3) the number of ballots recognized as valid (i.e., not spoiled), and 4) and 5) the number of votes for each of the two largest popular alternatives: parties and candidates. We will also limit ourselves to only three-digit and four-digit numbers.
The number of voters on the lists is not as obvious as it seems. The voter lists change right during voting: those who vote by absentee ballots come in, those who were missed when compiling the lists come in (for example, a person has recently moved), the dead are excluded, etc.
And this is what happens if, for simplicity, we take all five columns together (as was done with Sweden). In Fig. 3 - the last elections to the State Duma in 2011, and in Fig. 4 - the just held presidential elections in 2012.
Pictures 3 and 4 are similar and not similar. Both show that zero is a favorite number, but nine and seven are not. But with the help of x2, hidvams also see the difference between them! In the 2011 elections, the probability of random deviations of this magnitude was no less than 1/1027 (the denominator of this fraction is one followed by 27 zeros!), and in the 2012 elections it was only 1/1025. As we can see, stunning progress has taken place: 2 zeros have disappeared at once! True, there are still 25 left.
But there is also a subtlety that needs to be mentioned: for some reasons, which are not always valid, large numbers may still occur a little less often than small ones. Just in case, let's take this into account. What happens if we take into account such a drop is shown in Fig. 5 with a blue line.
Now x2 will test the difference not with a uniform distribution, but with an inclined background. Alas, even such an account will not make the result good: the probability will become approximately 1/1010 - one ten-billionth part. But there is also a funny circumstance: it has become better visible that not only zeros, but also fives are preferred, but there are few neighboring numbers. We will meet such love for excellent grades again.
But perhaps some other subtle features of the distribution of voters among precincts, across villages and cities, across the Caucasus and the Far East have an effect? Then you can “move” the zeros from their home in the decimal system by using other number systems (see Fig. 6).
In Fig. 6 the same data as in Fig. 3, only in the fivefold system (above) and in the septenary system (below). In the fivefold system, where fives and zeros are the same number “0,” the effect of unevenness remained, but in the sevenfold system it just disappeared: people do not see zeros in the sevenfold system and cannot prefer them. And at the same time, we will make sure that in the septenary system the suspicious “background” noticeable in Fig. 1 has disappeared without a trace. 5.
Where do magic and miracles come from?
The question arises: where exactly do such miracles come from? You can look at the data for all regions and for each of the five columns separately and find the record holder: this is Dagestan (Fig. 7).
Pay attention to the scale: zeros are found here one and a half to two times more often than other numbers! And, of course, my favorite five. Moreover, this can be seen for all graphs together, and for each individual graph in particular. (Note that in the 2011 elections, zeros in Dagestan were three times (!) more common than other numbers, so there is obvious movement for the better). Here the hidvams would probably shake their heads in disapproval.
Now let’s compare Dagestan with another region that occupies the very middle in the ranking of election reliability, 43rd place out of 85. This is the Vladimir region (Fig. 8). As we see, the hidvams cannot make a claim here: the probability is quite ordinary, 31%.
But what if it’s all about the number of PECs? Let's look at another region and the most “sensitive” column to zeros - the number of valid ballots (Figure 9).
No love for some of the numbers that was shown in Fig. 7 for Dagestan, not visible, no preferences. Where is the passion for zeros at the end of numbers or even fives? Where is the hatred for nines? Complete indifference.
The election results in the Caucasus at the very top of the power vertical were naively explained by the special teip structure of society and respect for elders (bosses). A society in which everyone votes as the elder says. So be it, let's believe it. But, as we see, then we will have to explain why teips predominantly consist of dozens of people who live in dozens (but not 9 people each), and go to elections in dozens, and make their choice in dozens? What are these miracles?
Maybe we should explain this differently?
But overall, the elections have become noticeably cleaner, we must give them their due. So, back in December, Dagestan showed mind-blowing reliability: 1/10204, 10 thousand googols, and now only 1/1064. In Moscow, according to the column “valid ballots” there were 3%, and now 78% are honest. Reliability has increased in almost all regions, and it has also increased throughout the country.
But, alas, the reliability in St. Petersburg dropped sharply and amounted to only 1.5%.
As A.S. wrote Pushkin: “It has arrived - who helped us here? The frenzy of the people, Barclay, winter or the Russian god?
But, of course, the analysis of the frequency of occurrence of the last digits is rough. It cannot by itself reveal more subtle methods of possible influence: “carousels”, forced voting, administrative pressure, bribery, etc., it only shows the very fact of “assault”. But there are also trends that can be identified by comparing regions with each other and obtaining unexpected results.
Regions and trends
Exactly the same analysis of the frequency of occurrence of the latest figures can be done for all regions of Russia. Each of the 85 will have its own probability value: from completely implausible, as in Dagestan, to quite reasonable, as in the Vladimir region, where the probability calculated by x2 is more than 80%.
But first, let's talk about rounding.
Suddenly, people in election commissions get a little tired of counting ten to two or three thousand? So they correct it a little: one up, one down. Sin, of course, is breaking the law, but isn’t it a mortal sin? (Chairman of the Moscow City Election Commission V. Gorbunov said this at the election rehearsal on February 25: they say they rounded up in December. A little bit.)
Let's check this too. After all, if the predominance of zeros and the lack of ones and nines is the result of only a small, simple and innocent rounding, then this should not affect the turnout and voting results, should it? Well, it can’t be that a cashier in a store, accidentally rounding up and making mistakes when giving out change, lives without denying herself anything? After all, she makes mistakes in her own pocket, then in the buyer’s pocket alternately and does not become richer.
Let's do it like the hidvams: we'll arrange the regions in order of probability and we'll sequentially exclude from the calculations of the voting results regions in which the probability calculated by x2 is too small. And then you can look at the results (Fig. 10). Here, the points with a reliability of, for example, 20% show the share of voters (right scale), turnout and voting results (left scale) in regions in which the reliability is better than 20%. (To avoid confusion, we note that turnout here is determined by the number of valid ballots, and the winner's result is determined in relation to their number.)
At the very left, on the Y axis, there are points showing the election results, if we take all regions indiscriminately, all 100%.
But if we take only those regions in which, for example, the probability is no worse than 20%, then it turns out that there is a lower turnout and a worse result for the winner. Surprisingly, the points stubbornly go down: the more carefully they were calculated, the less often they rounded, the... the less money our cashiers had at the end of the day. You can even see how much the “cashiers” were able to earn due to innocent rounding: in the end, 10 percent, isn’t it?
If we take that half of Russia in which the votes were counted by more accurate (or still more honest?) “cashiers,” then the result will be 5 percent worse, and the turnout will be 2 percent less. What if you take the neatest quarter? Ultimately, the green curve tends to the point where about 14% of the members of all election commissions live.
That's all honor and praise to them.
Questions instead of conclusions
As you can see, we, as befits Khid-Vamas, did not engage in politics. We just looked at the numbers, and even then only at the latest...
But there are three questions that I would like to ask the Chairman of the Central Election Commission V. Churov:
1. Why don’t he and the people in his department do this or a more complex analysis?
2. Why doesn’t the Central Election Commission send strict inspections to regions where the reliability of the results drops to mind-bogglingly low levels?
3. What are the Central Election Commission and its apparatus doing, to which taxpayers pay a lot of money precisely for organizing fair elections and, first of all, fair counting of votes?
Independent expert S.V.
The author thanks Alexey Shipilev, who provided the original data, Maxim Pshenichnikov for friendly criticism, and Vadim Kaymanovich, whose constant interest stimulated the work.