

On May 13, Openai announced a new flagship model of generative artificial intelligence called GPT-4O, where “O” means OMNI. The developers say that the model can process text, speech and video and will be available to users within a few weeks.
GPT-4O allows people to communicate with him as a real interlocutor. For example, users can ask ChatGPT a GPT-4O basis and interrupt ChatGPT during an answer. The model provides responsiveness almost in real time and can catch the nuances in the user's voice, in response to generating voices in “a number of various emotional styles” (including singing) and trying to determine in what emotional state a person is in.
About what GPT-4O is, and how the development of large language models will affect human society, The Insider was told by anonymous Russian independent researcher of artificial intelligence, who for several decades worked as the head of projects in large international IT companies:
“Her capabilities as a language model have not changed, it simply became multimodal. GPT-4 is a purely text model, they talked with it, introducing requests on the keyboard. The new GPT-4O model supports all the modality of the conversation: sees you, hears you, can communicate with you as you like-with the help of a voice, video image, and so on.
Just tell her: “Hello, let's talk about this topic with you,” and she will speak with you not as “Alice” or Siri, just reading out text answers, but understanding the emotional coloring of your voice and reacting accordingly. Video with a demonstration of her work makes an indelible impression.
All this has a tremendous meaning. The fact is that in the last year discussions were conducted about what these language models have practical use, since for the most part people are simply playing with them. Everything is wonderful, but what to build a business on? Now the answer to the question is received - this is a cardinal change in the form of communication of people and computers.
Now the smart assistant will work for you, launch certain applications, will decide in what sequence and what to do. This multimodal bot can communicate and communicate in any modality, and together with Microsoft means, it will make a complete revolution in communication between people with intellectual devices.
If you simplify, then someday you can ask: “Where am I going to do your glasses?”, And he, who watched your actions with the help of the camera, will say: “When you went behind milk, he put them on the nightstand.”
40 years ago, to write course or thesis, you went to the libraries, delved into forms, watched microfilms, examined catalogs. Google is now looking for information for us. Our information agenda is formed by smart algorithms - search, recommender and so on. We look and listen not to what we want, but what the algorithms recommend.
Personally, I belong to the category of grandfathers, and I do not know how to type text with two fingers on a smartphone. My generation does not have this "cognitive gadget." If we want to dial something on a smartphone, then poke with an index finger in it. In the young generation, the recruitment skill with two fingers is fixed from childhood, and it already becomes the essence that spends more in the digital world than in a dream.
The second tremendous change is due to the fact that we not only lose the decision skill, since algorithms take them for us, but also to separate the lie from the truth. And the third - the algorithms will take on 90% of any intellectual work. Behind the person, as a result, there will be something like a cluster role for models. But now special models are already being trained, which make tips to models better than people. ”
Subscribe to regular donations