
In the hands of Dmitry Kanevsky - two phones. They help him talk with people - even despite the fact that he himself does not hear anything.
On one phone, Kanevsky reads automatic decryption of his interlocutors. On the second - already Kanevsky’s interlocutors can read the decoding of what he says: the researcher has a strong Slavic accent, besides, he pronounces some vowel sounds louder and higher than others - so it can be difficult to understand.
For this to work, each phone has its own program. On the first - Live Transcribe. This is a free application for Android, which Google released in February 2019. It sends audio from the microphone of the phone to the servers of the company where speech is recognized - and after a couple of seconds a text decoding appears on the phone screen.
Live Transcribe deciphens the conversation of an ordinary person well, since the neural network was based on a “typical” speech. But when the application hears a person with speech features or with a strong accent, the number of errors increases sharply. It was then that the Euphonia application, installed on the second phone of Kanevsky (Google presented it in May 2019), helps.
Euphonia taught the scientist himself, so he understands his speech much better. Kanevsky began with a small number of audio -domed - about a hundred phrases that allowed the “smart” column of Google Home to understand his teams. Then he began to record his lectures and add accurate text decryptions to them. As a result, Kanevsky recorded 25 hours of audio for Euphonia, and now the program understands almost everything that he says. Including those words or phrases that are not easy to understand to untrained human ear.
In the Google blog, the work of Euphonia is compared with the standard speech recognition model. Where Kanevsky said “Come Right Back Please” (“Come back, please”), the standard model heard: “Cameras Object” (“Cameras, Object”). And when he asked: "DID I have ANYTHING To Say ABOUT It?" (“Did I need to say something about this?”), It turned out “Dictatorship Angels to Think About It” (“Angels of dictatorship will think about it”).
According to the scientist, Live Transcribe and Euphonia have significantly changed his life, and each application copes with its task. So, Live Transcribe allowed him to communicate with granddaughters - earlier he had to attract his son or wife as intermediaries. “They can ask me to play hide and seek or tell a story,” says Kanevsky. “They also love to watch the application deciphering their words, for them it becomes a form of the game.”
Euphonia helps Kanevsky to perform (in this case, the image from the phone screen is displayed on the slide). Kanevsky is a mathematician, and in the near future he plans to give a lecture in front of friends-mathematicians. To do this, he had to train the system for specific terms; So now he can say “algebraic geometry” and “switching lups of Mufang” - and the audience will understand it.
Euphonia is a research project that is so far available only to company employees. But Kanevsky hopes that in the future, others will be able to use technology - in particular, people with amyotrophic sclerosis (BAS), experiencing difficulties in speech. According to the scientist, for training the acoustic model, each person will not need to record 25 hours of audio-due to similar patterns in speech, you can combine the recordings of many votes. “We hope that if many people write down their votes, we will begin to cluster them,” Kanevsky explains. - And when a new person comes, we just find a suitable cluster. If he has to train the model for himself, then just a little. ”
Of course, the neural networks, recognizing speech, were not always with Kanevsky. In an interview with CNET, Kanevsky said that in kindergarten and at school he studied with children without hearing and learned to read well on his lips. During training in graduate school at Moscow State University in the late 1970s, he thought about emigration to Israel. In the nine months that he was waiting for permission to leave, the scientist invented and collected a device that simplified him to read foreigners on the lips of the speech. “For example, in Hebrew a lot of hissing sounds -“ Shabbat ”,“ shawl, ”he explained Cnet. - Therefore, I developed a device that converts high frequencies into low ones, and in Israel it helped me communicate. Thanks to him, I began to understand other people and speak in Hebrew and English. ”
In the 1980s, Kanevsky moved to the United States and began working in IBM on various technologies that help deaf people (the scientist has more than 200 patents). So, in the 1990s, he came up with the idea of connecting stenographers on the Internet: they listened to the phone that Kanevsky and his interlocutors say, and immediately brought to a special web page a deciphering of their phrases. Kanevsky used a similar system when he gave a telephone interview with CNET in 2012 - it was timed to the invitation to the ceremony in the White House.
With the help of stenographers connected by Skype, the scientist even gave lectures : students who wanted to ask a question had to speak in a tablet in Kanevsky’s hands - after a few seconds he received a text decryption. But this system has a serious flaw: the services of stenographers are expensive. Because of this, the choice of Kanevsky’s jobs was limited by large companies, such as IBM and Google. “If we met more than a year ago, our conversation would cost Google several thousand dollars,” he says Medusa. - I would hire a stenographer and practiced a week with him so that he knew what my speech would be about. It would take a lot of [paid] hours. ”
According to the scientist, when he began to work in the field of speech technology, he thought that to fulfill his dreams - a full -fledged system of speech recognition, which would help deaf people to communicate - would be required about five years. In fact, this path took 25 years. In recent years, two conditions contributed to the heyday of speech recognition technologies in recent years: computers have finally become powerful enough so that neural networks could work effectively; And on the Internet there was a site with hundreds of thousands of videos to which people downloaded text decryptions - Yetub.
Live Transcribe already recognizes speech quite qualitatively and even indicates extraneous sounds - for example, dog barking or applause; Although with some circumstances it copes with difficulty - when it is noisy around, the speaker sits far from the microphone or if several people speak at the same time.
Euphonia, according to the scientist, should develop to a system that will understand all people with non -standard speech. He imagines this this way: a person enters the site, loads a sample of his voice there or simply finds the option that is suitable for his case, downloads it to his phone and begins to communicate with people and voice assistants.
The work of Kanevsky and his colleagues is not limited to the transformation of speech into text. In 2018, they demonstrated the prototype of a mini-projector, which was attached to the scientist’s chest: in this way, Kanevsky could see the decryption of the words that his interlocutor pronounced, right on his chest. Glasses of augmented reality, allowing you to display the text right in front of a person’s eyes, will make communication with deaf people even more natural: “Now some people complain that during the conversation I stop looking at them and look at the phone screen.”
Google also has a Parrotron research project, which is not trying to “understand” at all what a person said with a non -standard speech - the neural network immediately “pronounces” the same phrase with a “standard” voice.
In the video demonstrating the development, Kanevsky is trying to ask a question to the voice assistant, but he understands him incorrectly. Then he says the same team in Parrotron installed on the phone, and the application already repeats it with a “standard” voice for the “smart” column.
Kanevsky emphasizes that the task of his team is not that hearing -impaired people completely switch to audio converting systems into the text - they also work on automatic recognition of the brutal language. For the visually impaired and blind people, Google released the Lookout application: you lead the camera to the object - and you get an oral explanation of what it is. And to help blinded people in the company are working on converting a text from Live Transcribe to Braille's font.
According to Kanevsky, other companies who want to help people with special needs must be remembered that the main challenge that is facing them is not technological. “I think the main challenge now is empathy. We underestimate how important this is, ”says the scientist Medusa. - You cannot just hire the developer and say: now we will make such an application. The most important challenge is to arouse empathy in people. ”
Want to support the Co-Union Foundation , which helps people with a simultaneous violation of hearing and vision? You can do it here
Sultan Suleimanov