At the OpenAI presentation in San Francisco, a new version of the language model with generative artificial intelligence GPT-4o was presented. The developers call it a step forward towards more natural human-computer interaction. The model can perceive any combination of text, audio and visual data and generate the same combinations in responses. And most importantly, GPT-4o has become even more human-like in communication.
The main part of the presentation was devoted to demonstrating the voice capabilities of the new model. They were available before, but now the delay in responses has significantly decreased and is on average 320 milliseconds, which is comparable to the speed of human reaction (in previous versions of GPT, this figure varied from 2.8 to 5.4 seconds). At the same time, when interacting with ChatGPT, the developers constantly interrupted it, but this did not affect the quality of the responses.
Programmer Robert Lukoshko noticed that most of GPT-4o's responses began with introductory words, and suggested that they were being reproduced by another, simpler model while the new version prepared a full response. In this way, the developers could not only create the appearance of an instant response, but also bring GPT-4o closer to the model of communication of real people. However, the programmer soon changed his mind after watching a video of two models singing, continuing phrases one after another.
Yes, GPT-4o can sing, as well as change the intonation of its voice (the chatbot can make it more dramatic or, on the contrary, deliberately speak like a robot upon request) and recognize the user's emotions. And also analyze visual information. The presentation showed how the model reads an equation written on paper through a smartphone camera and gives hints on how to solve it, correcting the user if he suggests incorrect options.
Last year, a breakthrough in solving elementary-school-level math problems was cited as a possible reason for Sam Altman’s temporary resignation as CEO of OpenAI. The results were rumored to have been achieved with the secretive Q* project, which was touted as a major step toward creating general (or strong) artificial intelligence . The developers reportedly informed the board of directors that the project was dangerous due to unpredictable consequences, and that Altman was not paying enough attention to risk assessment.
The model can also work as a translator from foreign languages. The company's technical director, Mira Murati, spoke to one of the developers in Italian, and he responded to her in English. ChatGPT recognized these phrases and immediately translated them into the desired language.
All these features, combined with the updated interface, are reminiscent of the sci-fi film Her, in which Joaquin Phoenix's character falls in love with an AI voiced by Scarlett Johansson, The Verge writes . OpenAI CEO Sam Altman (who did not appear at the presentation) hinted at this similarity by publishing a laconic tweet with the original title of the film.
The creators also showed a new ChatGPT app for macOS, which can not only communicate with the voice assistant, but also show it information on the screen by pressing a certain key combination. During the presentation, the model not only recognized the code on the screen and explained what exactly it does, but also explained the meaning of one of the functions.
It's not just programming that works. When GPT-4o was shown several temperature charts by month, it was also able to analyze them, describe them, and answer follow-up questions. The official press release says that in the future, ChatGPT will make conversations even more natural. For example, a user could show a live broadcast of a sports game and ask for an explanation of the rules.
Users with a ChatGPT Plus subscription have already begun accessing the app, with a wider release planned in the coming weeks. A Windows version of the app is expected later this year.
The GPT-4o model itself will also be distributed for free, but paid subscribers will have slightly more features. ChatGPT already runs on the new GPT-4o model, but for now this only applies to text and graphics. Voice capabilities will be available to a limited number of users soon — OpenAI plans to launch all new features gradually.
Early adopters of GPT-4o's capabilities describe them as nothing less than "crazy" (in a good way). For example, working with charts and data visualizations now takes less than 30 seconds.
Until voice functions are available, we can only poke fun at the fall in shares of language learning platform Duolingo, which occurred shortly after the launch of GPT-4o.
Mikhail Gerasimov