Remember Theranos startup, who claimed to have learned to do almost any blood tests just one drop of blood taken from his finger, but actually ordered tests in ordinary laboratories? The exposure of Theranos was shocked by biotechnology. And in the market of solutions based on artificial intelligence, no one is surprised. Developers regularly try to sell out the work of outsourcers under the guise of AI. How to recognize the forgery, at The Bell, explains Anastasia ten, the director of the consulting "Abbyy Russia".

At the end of July, investors discussed the loud exposure of the startup Scalefactor , which for six years issued the work of ordinary accountants for artificial intelligence and, due to this, attracted $ 100 million investments. This is far from the first similar case. Last year, Engineer.ai, the company was stuck on the fact that she did not develop applications with the help of AI, as claimed, but by the army of programmers from India. Even giants were caught on such. For example, Amazon offered customers a smart home product with intellectual video surveillance. And then it turned out that there was no AI there. There was an outsourcing Ukrainian office who manually looked at the frames from the cameras and allocated objects with the mouse: trees, cars, etc.
We somehow ran into Abbyy with a deceived customer. The grief-contractor with whom he cooperated offered a ready-made solution for automatic recognition of paper documents with extremely complex fonts and a background. To persuade the customer to a deal, these “experts” simply put 100 people, they processed a small array of documents by their forces and said that everything was done with the help of AI. When it came to a real project for millions of documents, the customer was surprised to find out that the technology has not yet been working. They turned to our specialists for help, they analyzed the task, it turned out that at least a few weeks needed to train and configure such a decision. This story has a good end: the project was implemented, although with some delay.
Everything is clear with scammers. But why do respectable companies seem to be deceiving?
To substitute bots by people in the West, the term was already invented: Wizard of OZ - by analogy with the circus of Goodwin, who controlled the mechanisms behind the screen and misleads unlucky spectators. According to the latest study of the British venture fund of the MMC Ventures, there are a lot of such “goodips”: more than 40% of European companies that position themselves as startups in the field of AI are actually nothing to do with it.
One of the reasons: so the companies want to access the customer data, whether it is documents, images, videos and so on. Any solution on the basis of machine learning requires a large number of marked data, which are often protected by commercial, banking or other secrets. So, in order to teach the technology to automatically attribute the document to a certain type and to extract the text from it, thousands of documents are needed, which indicate that they need to be “seen”. To recognize the handwritten text, a significant number of not just letters written by hand, but combinations of letters in words - for example, so that “and” and “n” go nearby.
So, the startup receives investments, data and slowly modify its technology. For example, this is what Return Path's developer did. The company offered a service to automatically write answers to letters. At the same time, no one warned anyone that people read and write, and letters themselves are used to develop marketing software. In the world of widespread remotes, good communication, crowdsourcing sites, it is easy to collect the army of volunteers and turn the data with their hands. But the business in this case pays not only for pseudo -automation, but also for human mistakes.
How to recognize forgery? Below the five main features of “Goodwin”.
When the already mentioned Engineer.ai said that he would develop an application of any complexity in just an hour with the help of AI, it sounded suspicious.
Yes, of course, there are typical projects in the field of intellectual technologies. You can launch a chatbot that answers the same type of user questions in a few days. In one or two months-if there are some more complex answers or forwarding options. But if you are planning a difficult project that has not been met in the market before, then you should not believe a specialist who promises to implement the project in two days.
To teach a person to make a routine task (for example, to extract some data) is easy. They hired 40 outsourcers, explained that they should reprint the name, date of birth and registration from the passport. Ready! It is not necessary to teach what this fields, how they look-it guarantees a quick launch: you don’t need to test, develop, correct, make a training sample if something did not work out the first time.
To teach AI to extract anything from the document, video or images is much more difficult. Of course, there are Transfer Learning methods: already pre -trained neural networks that can be modified at a limited amount of data for a particular script. But in this case, mature developers provide additional tools. In such projects, training can take place on the customer’s side, he himself can track the results and control the outcome of the project.
Not a single serious company that works in the field of AI will promise one hundred percent accuracy when processing any data - whether it is documents, photos of people or entries from video surveillance cameras. The technology can be mistaken, since it all depends not only on the algorithm, but also on external conditions.
For example, recognition of fingerprints works worse in the rain, in the cold, and in more banal situations - for example, if the user did not wash his hands. Even the most advanced technologies of person recognition are not perfect, although it can reach 99,9997%. If the contractor promises to achieve certain indicators in quality, in the contract he will prescribe the conditions under which this will be possible. For example, with automatic classification of documents, types of documents, fields that will be removed, requirements for the quality of the images themselves will be spelled out.
If the company immediately guarantees the perfect quality, then most likely the results of its work are additionally checked manually. Some do not even pretend, but openly admit this. For example, CloudSight.Ai does this: The company offers the web-sites for the service that marks photographs and images of tags. Most of the work really makes AI, but in cases where the program doubts, the request in real time takes one of the 800 outsourcers in India, Southeast Asia and Africa. This is honest with the customer: algorithms and people work in a bunch to improve the final result. It is no coincidence that CloudSight.ai is one of the few technologies that copes with the common problem of recognition of images: does not try to see sheep in images on which they are not.
Another alarming signal: the contractor immediately agrees to fulfill any of your Wishlist. Most often, such promises are made by young inexperienced developers who poorly evaluate their strength and resources. But sometimes it also happens that expectations from the product were originally overstated.
At one time, for example, many were partially disappointed in the possibilities of IBM Watson technology for medical diagnostics. At the testing stage, the solution was really impressive: information about all rare diseases was stored in the databases, and the technology could cope even with difficult cases in seconds. But, when the company began to introduce its product into medical organizations, it turned out that in the process of diagnosis there are many exceptions, and also doctors spend more time to download information about patients into the system.
An experienced performer, having studied the terms of reference, will ask many questions. Moreover, the final solution may differ greatly from the project at the start: during the implementation, the customer and the contractor better understand the goals of automation, the real capabilities and restrictions of technology in relation to a particular business process.
As a result of a preliminary assessment, it may even find out that AI in this case is not needed, but, for example, it is enough to organize the work of employees in a different way.
If the company has a trained neural network, this means that it was able to get arrays of data somewhere. Maybe these were the data of other customers. Open sources were used. If the company refuses to talk about the data on which it trained its technology, referring to “secrecy” and “uniqueness”, this is an alarming sign.
It is worth noting that such mystery is characteristic not only for AI developers, but also for technological startups in general. Probably the most loud example is Theranos . The founder of the company claimed to have developed a technology that allowed to get an accurate blood test of just a few drops. The technology, according to Holmes, allowed to conduct more than 200 tests, but no one knew how it worked. It turned out that most of the tests are made on the equipment of third -party manufacturers, and it does not have its own developments.
In general, the more information the contractor provides about his product, the better.
Although stereotypical robots are very similar to people, they think, dream and act as a person, real machines are very different from us. We are even mistaken in different ways.
The technology will never confuse Petrov Ivan with Ivan Petrov and will not be able to recognize 1931 as 1913, although these are typical mistakes for a person made by carelessness. Rather, the machine can recognize the number 3 as 8, and Petrova Ivan - as Peter0va Ivan. Or it may happen that the technology will not figure it out at all, what it processes for the file and gives an error. So the accountants from ScaleFactor were exposed when they began to receive reports with many human errors.
Another important point: to teach real AI longer and more difficult than people, but at the same time you can scale such a system very quickly. It is much easier to increase the cores of the processor, the power of the computer. It will not work to crank the same trick with people.
If you need to move from processing thousands of documents per day to a million, it is not so difficult to do this using technology. If there is iron, then this is a few hours of the system administrator. And to hire and quickly teach people in such a huge amount it is simply impossible. Scalefactor just faced this problem: they attracted a lot of customers from small and medium -sized businesses, but could not cope with all their financial statements. According to 15 former employees, the company has chosen an aggressive sales strategy and made capital involvement, rather than the development of software. The result is error and delay.