
Photo: Era
The ChatGPT revolution that began in 2023 introduced another fundamental adjustment. The popularity of the so -called basic training models (LLM) of generative AI on big data captured the whole world. And it took only six months so that half of the employees of leading world companies began to use large language models of the GPT-4 type in their work processes, and hundreds of the company began to offer all new products with built-in generative AI.
As a result, the filling of the Internet, which has already become the main data storage for everything in the world, has radically changed: from culinary recipes, jokes and life hacks for repairs to statistics, patents, scientific articles and all kinds of professional and analytical information.
It is important to understand two things. Until 2023, most of the content on the Internet was created by people. It was this content that was used to teach AI. From this year, the increasing share of content filling the Internet will be created by AI. It's not only about texts - but also about numerical information, images, photos, audio and video.
It is extremely important to understand where it all leads . The newly published preprint of a new study of the group of authors, led by Ross Anderson, warns of a huge ambush that awaits the world when filling the Internet with LLM products.
The result may be tremendous damage to business safety, as well as for the intelligence of mankind.
Ross Anderson, as the Royal Society of Great Britain notes, of which he is a member, is a “pioneer and world leader in the field of security engineering”. He is one of the best specialists in the world to detect security systems in security systems and algorithms, a member of the Royal Engineering Academy and a professor at the Safety Department of Safety and a computer laboratory of Cambridge University, as well as one of the most famous industry consultants in the field of infection. His work was laid the foundations for building threat models for a wide range of applications, from banking to health care. And now Ross Anderson with his colleagues warns of a new, global threat to all of mankind - collapse of large language models (LLM).
Scientists suggest that the following will happen:
As the Internet is filled with the fruits of the GPT models, each new model will more and more study on the content generated by previous models.
This will cause irreversible defects.
Later generations of models will begin to produce samples that would never be produced by the original model, that is, they will begin to incorrectly perceive reality, based on the mistakes made by their ancestors.
Remember the comedy “Multiple” with Michael Keaton in the main role in which a person cloning himself, and then clones clones? At the same time, each new clone becomes stupid than the previous one.
LLM will happen the same. If you teach Mozart’s musical model, you expect the result to be similar to Mozart: albeit without that brilliance (and therefore we will call this Salieri model), but it looks like it. But when then Salieri will teach the next generation, otherwise the generation is the next and so on how the fifth or sixth generation will sound? Obviously, everything is worse and worse.
A similar process of intellectual degradation of models is called Rossa Anderson and his colleagues on the study of the “Model Collapse”.
As a result of such a collapse, the Internet will be more and more clogged with nonsense - garbage data and garbage information.
But that's not all. For it will be not just garbage (nonsense that does not have information value), but a “radioactive” garbage, the use of which will be dangerous for the results of the activity and cognitive safety of users.
The main danger to the business will stem from the permanent “radioactive background”. Already used ChatGPT or similar tools to receive answers to non -trivial questions, know that sometimes they give out absolutely incorrect information. In addition, such AI systems often do not disclose sources of information or refer to non-existent sources of their so-called "Hallucinations." Operational and reputational damage to business and individual specialists who make decisions based on such information can be colossal.
The main threat to the cognitive safety of people using such AI systems will be that not only LLM will be drunk from the Internet of nonsense in ever-increasing volumes. People will be supplied with the same nonsense.
The growing harmfulness of filling the Internet will appear diverse. People will inexorably stupid , and " intellectual blindness " will increase in society. It will become more difficult to distinguish the truth from lies , so problems with critical thinking will begin. Excessive doses of “radioactive information garbage” will provoke an increase in cognitive distortion of both individuals and the whole society. Under the influence of this process, people's representation about the world will become more and more crooked.
No matter how terrible the above prospect is, this is only a warning, not a sentence.
You should not be like a naive techno-Pessimists focusing in their forecasts only at the exorbitant price of technology progress, ignoring the tremendous benefit from their use.
However, in a similar way, it is not necessary to like a naive techno-optimists coming exactly the opposite.
As an antidote from the conversion of the Internet into a landfill of “radioactive” information garbage, the study of Ross Anderson and his colleagues offers two specific ways to prevent a model collapse.
The first method is to mandatory a copy of the original set of data created by a person, and to prevent pollution of this copy generated by LLM. The second method is to include new, pure data generated by people in the training process.
There are other important tasks: the development of a policy for assessing the accuracy of models and their thorough testing, as well as building a reliable system for ensuring the quality of models and the results they generate.
None of the named, unfortunately, is not yet in the priority list of the most important tasks of any of the governments. And it is very dangerous. For here, unlike the challenges of global ecology, humanity will not have decades.