
LLM is a type of AI that uses large volumes of data and machine learning algorithms to analyze and understand the language (both natural and, for example, programming language). Such models are trained on huge volumes of text data, including those that are freely available on the Internet, and are able to apply knowledge in practice - for example, to generate text for answers to user questions.
One of the most famous large language models is the GPT (Generate Pre-Trained Transformer), developed by Openai (among its investors-Ilon Musk, Peter Til and other famous businessmen). It is trained in a huge amount of text data and can generate coherent texts, answers to questions, translate from one language to another and perform other tasks related to the natural language.
The Microsoft Corporation has already fitted into the history of language models. In January 2023, she announced investments in Openai - Bloomberg sources rated them at $ 10 billion in the next few years. Technologies of the latest version of the language model (GPT-3) Chatbot ChatGPT immediately began to integrate into commercial products- Bing , Edge browser and the Microsoft 365 office package, where you can already get acquainted and chat with this.

Such a sharp move from Microsoft attracted attention to Bing: if earlier technological analysts more often laughed at the obsessive attempts by Microsoft to convince at least Windows users to use their search engine (in January 2023 there were only 8.85% of the global search market, while Google had 84.69%), now now is seriously. They became interested in them and even discuss that ChatGPT’s integration can help Microsoft win the search market.
However, one should not think that other leaders of the technological market are lagging behind: they have long begun to integrate language models into their products. For example, Google uses the Bert model in its search engine to improve requests and results. It is used in the Google Assistant and Google Translate applications. And in early February, Google released (so far only for beta testing) the Bard chat boot, a full-fledged competitor of ChatGPT.
The same Facebook has long have a Roberta model (optimized Bert version) - it is used to improve the quality of recommendations, personalization and text analysis. They plan to integrate the new LLAMA model there in WhatsApp, Messenger and Instagram.
Amazon develops the Comprehend machine learning service, which uses BERT and other models to analyze text data, including users, social networks and news articles. Snapchat, Nvidia, Chinese Baidu and Huawei, as well as several startups, also have their own developments.
Large language models are considered resource -intensive technology - that is why large corporations (or companies that attract large investors from the start, like the same Openai) received a head start in entering the market.
At the same time, the technology of large language models in itself is not new. This is discussed in the same META - if you listen to not the first person and PR managers of the company, but the real scientists who are engaged in the project. So,
The well -known engineer of machine training, Yang Lekun, who leads the AI research in the field of AI, called ChatGPT - probably the most highlight that is now in the popular technological agenda - “not particularly innovative” and “not at all revolutionary” technology.
The statement seemed a scandalous wide audience, but definitely not to specialists. As Lekun himself explains, all existing large language models, although they differ in size, datasets (arrays of information on which they train), optimization algorithms, etc., are based on old developments.
So, the Transformer network architecture, created by Google by 2017 (it is built on it, including GPT-3), according to Lekun, uses the achievements of Canadian mathematician Yoshua Benjio, who created his large language model “about 20 years ago”. Not to mention the fact that Google and Facebook themselves have been using them for years - for example, in real commercial products.
Why do we hear so much about language models right now - at the beginning of 2023?
One of the reasons is the emergence of a convenient and free interface that made it possible to quickly test the technology to millions of people. ChatGPT was discovered on November 30, 2022, and the pace of penetration into the audience turned out to be unprecedented. The chat boot scored 100 million unique users per month already in January-that is, in just two months. For comparison: Tiktok took nine months, and Instagram has more than two years.
After that, a standard cycle of interest in new technologies, well described by Gartner, started. Her curve hype (Gartner Hype Cycle) reflects five stages of adoption by society:
“Big expectations”: New technology, product or service appear on the market and causes interest and enthusiasm among consumers and investors.
PEAK of Inflated Expectations: interest and attention to new technology reach their peak, there is a lot of noise and hype around it, but there are no or few practical applications yet.
“The Ground of Disorder”: A fall in interest in the new technology, since it did not justify the promises and did not give the expected results.
“Slope of Enlightenment): At this stage, the technology begins to reveal its potential, new ideas and methods of use appear.
Plateau of Productivity: The new technology becomes widely accepted and begins to bring significant results in business and society.

It seems that in the case of large language models, we confidently go to the peak of high expectations. Companies are in such a hurry to make new announcements about their future breakthroughs in this area that sometimes they themselves put themselves sticks in the wheels.
Here is a pair of examples. Google in early February lost $ 100 billion of market capitalization after his Bard Chatbot made a mistake right during the marketing presentation. In the commercial, promising that the new technology will simplify complex topics for users, Bard, in response to the question of discoveries made using the James Webb orbital telescope, confidently ascribes to him the pictures taken through the Tery Large Telescope of the European Southern Observatory.
The integration of ChatGPT in Bing also looked hasty:
The article by The New York Times, the author of which, having tested the technology, came to the conclusion that it has not yet been ready for use without control by a humans:
The bot in communication allowed aggression, threats, lies, and could take a certain political position.
Against this background, it is not surprising that Mark Zuckerberg came out with his statements later than the rest and at the same time did not show any technology actually: Facebook had already had its own failures with large language models, however, last year, when attention to them was not so high. The Blenderbot, published in August 2022, simply worked poorly , and the Galactica model-it was supposed to write scientific works- they closed it three days later , because she wrote mostly nonsense.
At the sight of a wave of hype around the language models and chat bots, the former, barely managed to end, the wave-metavselnaya come to mind. What the speed of META, which spent billions of dollars on this technology with almost any visible effect, switched in its public agenda on AI, surprised many observers. Do not forget that investments in the technology of metavseli and AR (augmented reality) also announced Google with Microsoft (but managed to get out of the news agenda on this topic a little earlier).
Within the framework of Hip, this is logical - and does not say anything about the prospects of a particular technology. Corporations have to act especially quickly in a situation where any careless statement can bring down quotes, and investors are already alert and expect a serious reduction in expenses from companies, which has already resulted in mass reductions among all leaders of the technological market. The same Meta now needs to continue to dismiss thousands of people to fit into their financial purposes. That is why one - already unprofitable - the narrative was quickly forgotten and replaced by another, more promising.
But perhaps this time everything is really more serious. Metavselnaya, at least in the interpretation of META, caused a comic effect rather, looked like a computer game like SIMS (and developed several years ago), and assumed largely purely marketing use - a potential showcase for brands.
Large language models offer a more concrete path: if no one actually needed a metavselnaya inside Facebook, then the AI assistant inside the search engine is obviously in demand now-despite the fact that he can still make mistakes.
The promise bar is also high.
In Cyberpanka novels, AIs are often described, which become self -sufficient and capable of learning. LLM may be the first step to create such intelligence.
Among other potential options for application - the use of AI on the basis of LLM to create long -term plans for the development of business or even states: these plans take into account many factors, such as economics, demography, politics and technology, the creation of new languages that can be used to communicate between people and machines, and much more.
There are many concerns: this time, robots may not first threaten to replace the working class, but “white collar” and creative workers - it is still unclear how society will be able to adapt to such reality. There are questions to what new opportunities the technology will give to disseminate misinformation - and how to deal with it.
In any case, in a year or two, the corporation will probably switch to something else-it is good if during this time the risks can be minimized, and the technology will begin to work for the benefit of users as we may not even imagine.
One of the optimistic scenarios of its application was described by the Slovenian philosopher with the glory of Zizhek, answering the question of whether he will bury the educational system: “No! My student brings me my essay, written by AI, I connect it to my AI for evaluation, and we are free! While "training" takes place, our super-ego is satisfied and we are free to learn everything that we want. "