The network turned out to be an unreliable repository, and many data were lost irrevocably-why it happened and whether it is possible to somehow prevent it.

Imagine: you are looking for information on the Internet about a performer who has risen to the top of the hit parade in the summer of 2001, or an event that occurred at the end of spring 2002. Of course, you go to Wikipedia and read general information, but they are not enough. In an effort to learn more, you click on all links at the end of the articles, but over and over again you get on non -working sites.
The thing is that the Internet is dead. Trying to find the sources of the events of those times, you seem to go to the cemetery: here it gives out an error, here - the domain is expired, there was never nothing at all. It seemed that all photos and records uploaded to the network would remain in it forever, but now the links of the beginning of the century either disappeared without a trace or work only through the web archive.
20 years ago, the Internet looked completely different: sites were distinguished by variegated, background music and a lot of animation. Then Yandex was already working (though it looked completely wrong), but users also looked for the necessary information through Aport and Rambler, talked in the chat “Cotle” rooms and read the news on the “Web Planet”.
The listed sites (again, except for Yandex and Rambler) no longer work. If you follow the link Aport.ru , you can see not a search engine, but the price-aggregator, similar to Yandex.Market. In this form, it has existed since 2012, when the company was bought by the director of the site of Mamba.ru Andrey Bronnetsky for 150 thousand dollars.

The “crib” is also dead, according to the old link “hangs” a memorial plate with the inscription “krovatka.ru. 1996-2020. It was the best time. " The site worked as an online chat with 25 channels, among which-“Dating”, “Love”, “For 30”, “Art” and “Computers”. The order was monitored by the moderators who blocked the violation of user rules.
The reasons for the closure of the site were not declared, but in 2016, one of its authors Andrei Kul told the “secret of the company” that with the advent of social networks, the chat left the “lion's share of users”. “When LiveInternet appeared, people began to leave slowly, then Skype appeared. The audience was still, but progressive growth ceased. Then social networks began to appear - “VKontakte”, “classmates” - the people went there, ”he said.

The Web Planeta, created by Denis Kryuchkov, who later discovered one of the largest IT communities of Habrahabr, is not updated even longer. The online publication closed at the end of 2011. The reasons for the decision were not announced, but, according to Lenta.ru, the case is the unprofitability of the project. The editor-in-chief of the project Lyokha Andreev told the publication that "all living things are different, that someday dies."

The sites of the zero are closed every year - sometimes whole "packs". In April 2019, Google closed the Google+ social network for eight years. She possessed all typical attributes, for example, the opportunity to update the status, post photos in the tape and call up on video communication. In the first weeks after launch, several million users registered in the service, but the social network did not become popular.
“I open the news tape, but I see an empty page on which nothing happens. This is a huge wasteland, which is distinguished by an abundance of registered people and who did not begin to use the service, because they did not understand its work, ” wrote Forbes journalist Paul Tassi shortly after the launch of the platform.

The owners justified the closure of the low popularity of the service and problems with user data protection. Experts said that the social network was inconvenient: the interface was constantly changing in Google+, people with pseudonyms were blocked and brands were removed. The consultant for working with social networks Matt Navarra noted that because of this, "the unenviable fate of the service was a foregone conclusion from the first day."
Yahoo closed several projects: in 2001, they stopped working on the Internet radio Broadcast.com, which has existed since 1995 and redeemed three years earlier. The company has acquired a successful project with 570 thousand users, each of which was estimated at ten thousand dollars, but turned off the service due to the decline in the popularity of the industry.
In May 2021, Yahoo closed one of the oldest services and answers to Yahoo Answers. An understandable analogy for the Runet culture is as if millions of discussions from the “ Mail.ru answers” irrevocably deleted from the network. You can’t go to the site anymore - users were allowed to download part of the questions and answers by preliminary application, but this function was closed in June. At the same time, the general archive was refused to create the service. The reason for the termination of work in Yahoo !!!! They again called the fall in popularity.

Messengers also disappear - naturally, along with the information stored in them. A striking example: Aol browser client, who worked for exactly 20 years, closed in December 2017. In a farewell letter, Oath Vice President Michael Albers admitted that since the 1990s, the means of communication have changed, and the messenger himself lost the struggle of SMS, WhatsApp and other social networks.
Sometimes the loss of sites is associated with historical processes: for example, this happened with Yugoslavia. According to the director of the Institute of Web Nauki at the University of Southampton, Dam Wendy Hall, the .yu domain of the upper level for Yugoslavia, ceased to exist after the collapse of the country. “There is a researcher who is trying to restore what was there,” the specialist notes.
Sometimes sites do not die, but modernize that in the person of users it looks like “partial death”: during the reform, entire sections disappear, and with them information. So, for example, it happened with Myspace: in 2019, due to the unsuccessful transfer of the server, all the contents of the profiles and all the music, loaded until 2015, disappeared.
Owners of accounts on Flickr and Webshots had to worry about the loss of pictures - but due to the change of owners. When Smugmug came to the first company, users were ordered to buy a paid subscription or “part” with all photos except the last thousand. Buzzfeed assumes that as a result they deleted the “huge number of photos”, many of which were laid out by people “not worried about their loss”.

Webshots, which successfully worked in the 2000s as a photo exchange service, has turned American Greetings into a desktop wallpaper. In only two months, users learned that all their files would be deleted if they do not buy a paid account. The same story happened with the platform with the reviews of the Xanga books and music-in 2013, the service deleted user blogs who did not pay for the Pro account.
Information can disappear from the Internet and automatically: this applies to e -mail and instant messengers. If you do not go to Telegram for several months, then it will delete the account, and with it - all correspondence and files. Since November 2019, Twitter has the same policy - an account should not attend six months.

The disappearance of information leads to the fact that users cease to trust the Internet and companies that owns sites. In case of unprofitability, a change in a business course or claims by state authorities and large companies, your profile and all information can be deleted without a return.
Often the extinction of links leads to serious consequences - for example, when one site closes, and in its place deliberately or accidentally, another appears. All this leads to a situation where you can not be confident in your personally affixed links.
In 2010, the American judge Samuel Alito expressed a special opinion regarding the cancellation of the ban on selling “cruel” video games for children in California and accompanied him with a link to a detailed explanation of his opinion. Soon after the publication of the text, everyone who crossed it saw at all what the judge wanted.
“Are you not glad that you have not quoted this web page in the report of the Supreme Court in the Brown case <...>. If you did this, as Judge Alito did, the initial content would have disappeared long ago, and someone else could come and buy a domain to comment on the speed of related information in the Internet era, ”the report said in a message.
Around the "dead" links on large resources, a whole shadow industry has been built . If such a link leads to a non -existent site, then it can “reanimate” it to order with the same domain and the same addressing to a particular page. But instead of original information, on this page they can place an advertisement or page with the exact opposite information.
But this is only one example. According to a study published in Harvard Law Review in March 2014, 50% of links from the judicial conclusions of the Supreme Court since 1996, when the hyperlink was used for the first time, they no longer work. The same thing happened with Harvard Law Review: scientists have found that 75% of links from the magazine cannot be opened.
The journalists of The Atlantic and The New York Times analyzed about two million external links published in the articles on the NYT website, and found that 25% of them no longer work. The older the article, the less likely it is that you can “go” somewhere: 72% of links do not work in 1998 materials.

This situation leads to a break in the chains of information, which the Internet is strong in its ideal form. By going to any site, you can go to another site, and then another, thereby finding the sources, causes and sources of any knowledge. The disappearance of links violates this order and often affects, for example, special scientific knowledge. The situation is complicated by the fact that the information analogues of information storage are everywhere abandoned, focusing on digital format.
For example, as scientists from Princeton University found out back in 2001, the number of URL addresses in scientific articles is growing every year, but 53% of them do not work. The work of 2014, uniting 3.5 million articles on science and technology, showed that every fifth of them does not indicate the original source.
The extinction of links violates the integrity and evidence base of scientific research. It is difficult for scientists to influence this, because they are not responsible for the safety of resources, but site owners. Attempts to fight independently to ineffective solutions: for example, in the journal Cancer Research it is forbidden to put links to the URL, and in Russian publications it is necessary to put a label on the date of the last appeal to the resource.
The scale of the disappearance of links demonstratively demonstrates the project of The Million Dollar Homepage Alex Tum. The 21-year-old student created it in August 2005 to raise money for training. On the site with a net of 1000 per 1000 pixels for one dollar, exile places were sold. All pixels were sold in 138 days, but by 2014 22% of them already led to dead web pages.

The problem of the disappearance of links also applies to TJ - articles that have been published in the first years of the site’s existence are available, but there are no photos in them. All due to the move of the pictures to another server. For example, you can show the text about the project “Million Pixel”, published in March 2014, but the widget will be ugly-precisely because of the lack of illustrations.
The main reason for the extinction of web pages is the decentralization of the Internet. For the safety of information, the owners of specific sites are responsible, which close them, change the structure and links, and sometimes they simply forget to update the registration of the domain.
Content becomes unavailable as a result of deliberate actions: for example, in 2015, Buzzfeed removed more than a thousand materials that advertisers and partners complained about. This affected articles criticizing the advertising content of AXE, Microsoft Internet Explorer and Twitter.
Media, and sometimes entire sites are deleted at the request of the authorities: for example, in the summer in Russia the publications “MBH Media” and “Open Russia” were blocked , and “Project” was recognized as an “unwanted organization”. Due to the status of the last publication, other media are forced to delete materials with links to his articles at the request of Roskomnadzor.
The extinction of links is included in the scenario of the “Digital Dark Age” - a theory in which all electronic data that does not have paper equivalents will disappear from the world. The concept appeared back in the 1990s and refers to the era of the Middle Ages, which was distinguished by an almost complete absence of written evidence. The main argument of the theory is precisely that all digital data constantly disappear.

For example, in 1986, the BBC launched the Judgment Day project in honor of the 900th anniversary of the Book of the Last Court-a set of materials collected by order of William the Conqueror about the possession of his kingdom. The publication asked the residents of Great Britain to document their native cities-more than a million people participated in the action, they gathered photos, cards and video tours. But by the beginning of the 2000s it turned out that all physical carriers of the project were broken or lost, and the data was lost.
It is noteworthy that the original book of the Last Judgment from 1086 is not lost, but is stored in the state archive in Kew and anyone can get access to it. “It is ironic that the 15-year version is unreadable, and the ancient version is still suitable for use. We were lucky that Shakespeare did not write on the old PC, ”said Paul Witley, computer specialist in a conversation with The Guardian.
The prospect of losing all digital information does not inspire humanity, so society is trying to solve the problem of storing data. In 1997, they published an international OAIS standard that defines approaches and solutions in the field of electronic archiving. Following him, several more documents were adopted, including Trusted Digital Repository, Digital Preservation Network (DPN), Interpares Project and Pronom.
Standards have established seven main archive strategies for digital materials:
The most curious object for maintaining information is the Arctic World Archive, opened in March 2017 at the Spitsbergen archipelago. In the bunker, called the media, the “Second Day of Judgment Day”, there are reserve data in case the originals are damaged due to wars or natural disasters.

All information is stored in a shelter in a super -resistant film covered with iron oxide powder. According to the manufacturer, it is able to withstand up to 750 years in normal conditions and up to two thousand years in a cave with a low oxygen content.
In October 2019, Microsoft began transferring the entire source code from GitHub to the "Day of Judgment Day". The first bobbin was recorded by the code of the Linux and Android operating systems and six thousand other important Open-Source applications. By July 2020, the entire archive of the site in the size of 21 terabyt (or 186 coils) was transferred to the bunker.

Web archivists are engaged in preserving directly links and sites. The first to the problem of “death of links” was drawn the attention of Bruster Cale. Still studying at the Massachusetts Technological Institute, he did not accept the closeness of information: to get into Harvard’s legal library and gain access to business for his work, he used the professor’s certificate.
In 1996, Cail founded the non -profit organization Internet Archive, the purpose of which was to preserve the knowledge on the Internet. According to him, the main difficulty is that everything is constantly changing in the network: the average life of the web pages is 90 days, after which they change or disappear.

Only the service administration had access to the information for the first five years - all the data was stored on the servers of the Archive. Since 2001, archivists opened access to saved data to everyone. Initially, the organization worked only as a web archive, but gradually they began to preserve books, audio, texts of Open Library and software. In December 2021, there are more than 635 billion pages in the archive.
Web pages are preserved using the Wayback Machine service, the “spider” of which regularly explores available sites and saves them on specialized servers. Each new copy of the page does not rewrite the previous one, but is saved separately with the date of addition. Links can be added manually if the “spider” has not reached the desired page.
Internet Archive is known for several large projects: for example, in 2000, archivists, along with the Congress library, collected information about the political campaigns of candidates for the US presidential election, and in 2001-about the attack in New York . Two collaborations with Wikipedia are also interesting: with the replacement of several million dead links to archival copies and the development of the function of pre -examination of books.
Хранение обеспечивается с помощью системы зеркальных сайтов, расположенных в отдалённых друг от друга местах. Все файлы сохраняются в формате ARC. Копии Wayback Machine находятся в Сан-Франциско, Ричмонде, Александрии и Амстердаме.

Какую часть интернета удалось сохранить архивистам, неизвестно. «Я бы выглядел идиотом [если бы попытался оценить]. Потому что никто не может точно определить размер интернета. Бесполезно беспокоиться о том, что вам неподвластно», — говорит Брюстер Кейл.
Работа «Архива» изменила отношение к ссылкам в интернете — в мире стало появляться множество программ по архивированию сайтов. К процессу массово подключились государственные организации — например, Библиотека Конгресса и национальные библиотеки Австралии, Швеции и Норвегии. В 2013 году Европейский союз запустил проект EU web archive, где сохраняются сайты ЕС.
Веб-архив Библиотеки Конгресса сохраняет миллиарды объектов — от сайтов правительства США до культурно значимых мемов. Уже более 20 лет этим занимается Эбби Гротке — руководитель группы веб-архивирования. «Мы просто пытаемся зафиксировать изменения во времени», — описывает свою деятельность специалистка.
Созданием архива сайтов российских организаций и учреждений с 2017 года занимается президентская библиотека. На периодической основе специалисты архивируют такие ресурсы, как сайты президента России и правительства России — копия создаётся каждый день.

Ещё одно крупное некоммерческое объединение энтузиастов — Archive Team — занимается сохранением частей интернета с 2009 года, когда компания Yahoo закрыла Geocities — веб-хостинг с сайтами пользователей. Проект создал историк технологий Джейсон Скотт, приводивший в числе причин «чувство гнева и бессилия», возникающее у пользователей.
Мы позволяем компаниям решать за нас, что выживет, а что умрёт. Но это не наша работа выяснить, что ценно и что значимо. Мы действуем на основе трёх добродетелей — ярости, паранойи и клептомании.
Джейсон Скотт
Первоочередная задача Archive Team — сохранить контент, размещённый на онлайн-сервисах из группы риска. Так специалисты занимаются архивированием, например, Yahoo! Video , Google Video ,Splinder , Friendster , FortuneCity и сокращённых URL-ссылок. В ноябре 2019 года команда запустила инициативу «Twittering Dead» по сохранению твитов умерших людей. Заявки оставляют пользователи, передающие ссылки через Google-формы.
В Archive Team входят независимые пользователи и авторы. Процесс сохранения сайтов выглядит так: архивариусы загружают страницы в виртуальную машинную среду Warrior, после чего она появляется в хранилище The Internet Archive. В 2019 году «Архив Интернета» и Archive Team подписали соглашение о сохранении публичных постов с закрывшейся соцсети Google+. За первые четыре недели архивации специалисты собрали 1,56 петабайт данных.
Исполнительный директор института веб-науки при Саутгемптонском университете Дам Венди Холл подчёркивает важность архива: «Если бы не они, то у нас не было бы ни одного из ранних сайтов. Если бы Брюстер Кейл не создал архив и не начал сохранять ссылки, не дожидаясь разрешения, мы бы всё потеряли».
Работа веб-архивистов ценна ещё и тем, что они сделали то, чем должны были заниматься обычные архивы и национальные библиотеки — но «растерялись» из-за быстрого роста значимости интернета. «Британская библиотека должна иметь копию каждой местной газеты. Но когда газеты перешли из печати в сеть, архивирование приобрело другую форму. Являются ли эти веб-сайты таким же важным источником, как и предшествовавшие им газеты?», — спрашивает Венди Холл.
Сотрудник веб-архива Британской библиотеки Джейсон Веббер считает важной проблемой то, что, несмотря на усилия архивистов, «большая часть интернета нигде не хранится». «Сохранение интернета началось только через пять лет после появления первых веб-страниц. Не осталось ничего из той эпохи. А первая веб-страница, созданная в 1991 году, больше не существует, сохранённый в архиве вариант — её копия», — говорит специалист.
Цифровой мир очень эфемерен, мы смотрим на свои телефоны, материал на них меняется, и мы не задумываемся об этом. Но сейчас люди всё больше осознают, как много мы можем потерять.
Джейсон Веббер
#лонгриды #технологии #медиа #истории