Rubric: Our networks
Mr. Scooter, bibliophil and idiot
Anton Nosik <anton@beseder.com>
In the last issue of our heading, we talked about the principles of the work of the Altavist search engine (http://altavista.digital.com/). Today we continue and complete the conversation on this topic - in the hope that the reader received enough information to independently continue to work on the Altavosty.
Yesterday we wrote about how a scooter (a robot replenishing the Altavistan information base) harvests its crop. In particular, an example of how to add our resource to the information base in one second, which was updated six months ago no more than once a month. Let us return to this example, for the HTML file that we published yesterday needs some comments. This listing:
<Html> <head>
<Title> Proverka Altavisty </ Title>
<Meta name = keywords concent = "golovozhopoe, proverka, altavista">
<Meta name = description control = "Proverim altavistu na vshivost">
</head> <body>
ETA Stranica Ne Neset Nikakoj Smyslovoj Nagruzki. Vse Pretenzii po etomu povodu k <a href = mailto: nosik@usa.net> nosiku </a>.
</body> </ html>
In it, as the reader can notice, there are two teams of the META class. Both of them do not carry any semantic load from the point of view of the “ordinary” visitor of our site (neither in the explorer, nor in the explorer, nor in the linx they are simply not visible), but these META commands are addressed only to indexing systems - all kinds of robots, which, in addition to the Altavist scooter on the network, are several tens. Indexing systems use the META command "Description" in order to announce the message about our page in the output search results. If there is no META command of the Description in our file, the first 512 characters of its visible text will be displayed on the screen as a sample of the file contents (HTML commands are discarded during indexing). As for the META command "Keywords" (Keywords), its contents are directly entered into the robot database index.
However, it is worth immediately making a reservation that not all robots readily swallow these two bait. For example, Excite search engine simply ignores all META commands and analyzes only the visible text of the page using its intelligent system of semantic analysis. My own opinion about the exat and its search technology is as follows: a brilliant idea, compromised by a very vulgar performance - both with a lexical and a technical point of view. The reader is probably not too interesting to know about the technical disabilities of the Architext search engine, based on the ex -technology, who has already unloaded this package for himself and fucked his joys. And lexically, "semantic annotation of pages" in the database of the exate does not even closely reflect the contents of the documents. If, for example, you argue in your document about the color scheme, then the word “color” can appear once or twice, and even then in the META list of keywords, while in the document you will be operated by concepts such as “palette”, “spectrum”, “red”, “green”, “blue”, etc. After indexing by excite technology, your page will be displayed during the key search for the word "red", and will not be displayed when searching for the word "color". If this is called an intelligent contextual analysis, then what is called oligophrenia at the stage of idiocy?
Speaking of idiocy. Remember our story about Scott Pakin and his novel "Idiot" in the issue of week ago? The novel consists of very short chapters, each of which is a paraphrase of one statement: that an idiot can be puzzled for many hours if it is offered to his attention Internet information databases created using multimedia, serving the only goal of puzzling idiots. At the end of each chapter, a link was contained, offering to read further. Each next chapter was generated by a special script, by changing the order of words and the substitution of synonyms for the head of the previous one. You can not read it sequentially, but request a chapter with any most unthinkable number - for example, chapter 98605043, located, in accordance with the name of the chapters, in a file with the name Chapter98605043.html. Therefore, it is easy to understand how the independent researcher Stanislav Malyshev from Jerusalem (frodo@sharat.co.il) was taken aback, when he found a link to such a curious file in Altavist:
http://wwww-csag.cs.uiuc.edu/individual/pakin/idiot/chapter256231258547854578217348561762475844524854856147856175124583.HTML
Stas’s imagination quickly painted the following picture: the scooter consistently reads the chapter behind the head of the Pakinsky novel, enters into the database each next stuffed file, and over time, the entire database begins to consist of 99% of non -existent (but generated by the script at the first request) variations on the topic "How to puzzle the idiot for many hours."
I wrote a letter on this subject to the department of technical support, the Altavists, but I did not wait for an answer to this day. However, Scott himself appeared during this time (the text of our note over the past Wednesday will soon appear on the website of its automatic complaints generator, in the section "Publications"). Pakin explained that the Links on his "Idiot" with astronomical numbers of the heads were taken by a scooter from the pages of different people who have filled the "idiot" manually. “The scooter itself is indexing only one page in one pass,” Scott Pakin recalled, “and the next chapter should be indexed only when the line comes to it, that is, about a week after indexing the previous one. In this way, to the astronomical figure, which is indicated in some of the existing lincum Altavists on the“ idiot ”, the case will not reach ...”
It is difficult to understand why such an obvious consideration did not occur to me myself. However, about four hundred links on the "idiot" scooter managed to absorb in this way - consistent - in the way. In total, Altavista took into account about nine hundred chapters of the novel.
“Nevertheless, I took measures,” writes Scott Pakin, “I put the instructions in the Robots.txt file, which prohibits Altavist and all other robots of a similar design to index chapter*.html in the idiot directory.”
For those readers to whom this topic is not close, let's explain: the Robots.txt file allows the owner of the www server to prohibit robots indexing individual documents or directory on their website.
In the conclusion of our story, we will tell one story, which partly rehabilitates the scooter after such a straightforward accusation of idiocy. The word to Levon Delitsyn, the owner of one of the largest meetings of Russian literature on the Internet, the holder of the American server of Roman, the founder of the International Tenet competition, publisher of the literary magazine Delitzyne.
"For several years, looking at the lists of visitors to my site, I invariably found that some Mr. Scooter from Digital again and again visited me again and again. Unlike all normal visitors who read two or three pages, this strange gentleman with a non-Russian surname in every parish read all the documents on my site without exception. And then he returned again and then return I read it strongly.
According to Delitsyn, on some sites, the share of robot calls from the total number of HTTP queries reaches 60%. This statement is very easy to verify arithmetic. If the Sharat server has 600 HTML documents, and every day 1000 “living” visitors enters the server, reading on average 3 pages per nose, then this means 90,000 requests from living users every month. If at the same time, 30 indexing robots enter the same server every week, each of which brings all 600 documents on the list, then in a month these robots send Sharat 77.400 requests for documents. Thus, the share of robot requests is in this particular example about 46% of all contacts to the site. And, of course, for many, indexing robots are the only readers ...
The next week we will start with a trip to the shitty cakes. Then there will be a summary of the Internet actuals and another story about the Russian network.