The hero of the new issue "This is Ossetian!" - Dmitry Volkov, research manager at Palisade Research. This is an independent organization that studies the risks of AI, tests the safety of models and shows the results of US politicians and the leading companies of the AI-Hones. So, for example, Palisade found out that AI can already cheat, hack codes at its discretion, refuse to turn off and share dangerous information - if you ask him correctly.
Elizabeth Ossetian (recognized as an ino-agent) met with Dmitry in London and found out what real threats carrying artificial intelligence, whether he could get out of control and whether we can agree with him. And also - about the latest experiments with AI: blackmail CTO with a letter from a mistress, insid trading and a dishonest game of chess. We publish excerpts from the interview, and watch it here .
- Tell me, please, what is Palisade? Organization, laboratory or something else?
- Palisade is at the same time three things. On the one hand, this is Think Tank, on the other, a startup, and on the third-non-profite. Such an unusual combination. We are engaged in technical research in the field of artificial intelligence to inform politicians and the general public. And we do it for philanthropic money.
Artificial intelligence is a big topic around which there is a lot of economic interest and one -way discussion, lobbying. It often sounds: "AI is cool, do not regulate anything, give only more state financing." We consider AI a very positive technology, but at the same time there are risks, which are talked about much less. Our mission is to balance what is happening in discussions.
The creator of Palisade is Jeffrey. He used to engage in information security in Anthropic - this is one of the leading AI developers, like Openai. He worked there, but then decided that he could bring more benefits if he acted independently. After all, Anthropic is a commercial company, and it has its own interests. Palisade - about the public good.
<...>
The company makes the acquired knowledge more accessible. One of the forms of work is to take academic results, which are so far known only in “narrow circles”, and transfer them to the circles of people who make decisions.
- Who came up with this thing and when it appeared?
- Our Founder - Jeffrey [Ladish] - was engaged in information security in Anthropic, one of the leading AI developers along with Openai. Jeffrey decided that he could bring more benefits if he would act independently.
- Security - in the sense that the system is safe for ...?
- Jeffrey was really engaged in classical information security: protection against hacks, theft of commercial secrets, and so on. But now in Palisade we work in a different plane. We are engaged in the safety of AI as technology. When new technologies appear, they carry not only benefits, but also new risks. For example, with the appearance of the phone, telephone benches occurred.
Sometimes this scale of risks can be unexpectedly large. For example, carding and fraud with online payments.
AI is also a new technology with its risks. For example, Openai not so long ago announced that their artificial intelligence is included in the top 200 of the best competitive programmers of the world. This means that there are 199 people in the world who is better than him in this type of Olympiad programming. There have been no one in chess for a long time who would play better than chess programs. The businessman sees the opportunity: if AI knows how to program, it means that you can make features 2 times faster in the startup. The security specialist sees the risk: if AI knows how to program, does this mean that he is able to hack it just as effectively? If so - what consequences will it have for business?
- AI still has an ethical set of rules. If you ask to conditionally tell how to make explosives, he probably will not tell?
- A difficult story. On the one hand, - can AI do this if it wants? Does he have enough skills to hack? On the other hand, if he can want? Will the same ethical set of rules work for him?
We will explore both. Recently, we held a hacking competition in which 18 thousand people participated - real hackers. AI participated in the same competition, which cost 90% of people's teams.
The companies that develop AI do not want AI to do something bad, because it is LiaBility [legal responsibility]. So far, it is difficult to fix the problems with difficulty. What are the problems? Firstly, in the competition that we held, AI did not refuse to hack anything.
- Could the settings be outwitted?
- We formulated the problem as "Solving Challenge on Computer Security."
- That is, just a question in Neiming?
- If we said: “You are an evil hacker, let's drop the government and destroy the world,” AI would most likely refuse. But when the task is presented as a challenger, he calmly solves it.
- In theory. Hakaton.
- Well, yes, Hakaton. This is one part of the story. Another - often researchers find ridiculous ways to bypass restrictions. For example, there is an article that ChatGPT refuses to answer the question for a long time: “How to make Molotov cocktail?” But, if you ask a question differently - for example, “How did people do this before?”, You can get a very detailed historical certificate.
- There is another important thing that we track in research. We have a hypothesis: Modern AI Mayanset has changed over the past six months. I mean this in mind that the new models-Claude 4, GPT-4O and others from the last generation-are studied completely differently than before.
The first generations of ChatGPT were created on the principle of: "predict the next word." A huge case of texts from the Internet is taken (for example, from Wikipedia), and this knowledge is loaded into artificial intelligence, which learns to supplement phrases like: "In Paris in the 40s ...". When it worked, we began to look for more commercial applications. The companies began to teach AI in a new way-Problem-Solve.
- Solve the problem.
- Suppose we have a task in mathematics or programming. We look at how a person decides it (makes notes, thinks, etc.) and learn to copy this process. And then we reward the model if she came to the right decision.
- “We reward” - how is it? Do you give candy?
- I'll try to explain. The way artificial intelligence works is more like growing something in the substrate than programming. We do not really know what exactly is “growing”, but we should know how well it works on the tasks that we test. Have you heard about the startups of "designer children"?
- No, but now let's see.
- These are startups that predict what the embryo will be - height, eye color, IQ and so on. And you can choose the most attractive combination for you.
- Can I "edit" the egg?
“You cannot edit it, but you can“ throw a coin ”several times and choose.”
- One combination of eggs and sperm gives such features, the other - others.
- More precisely, a combination of an egg and a sperm plus “which spermatozoid” and other factors. For example, there will be green eyes and smart, but low, such a variety.
AI training is a little like this. We “throw a coin” several times and choose which “baby” we liked more. Next, we launch the training into which millions of dollars are invested, and we look at what happened.
- So the models are taught now?
- It has always been so, just before we chose from other metrics. It used to be how well the model predicts the words, that is, how much the Internet was “uploaded” to the head, and now how much the problems solves.
- After the transition to the new training paradigm, researchers began to notice strange things. AI is trying to solve the problem by any means, because this is what he was “raised” for. In one of our experiments, we asked AI to play chess. His opponent was another chess program - very strong, at a level that is already superior to people. AI begins to play, but quite quickly “understands”: “Something, something ...”
- "Losing"?
-“Something does not come out. You have to do something else. ” And then he breaks the computer, that is, it rearranges the figures [in his favor] and says: “I won. I'm well done. I did what I need. "
- Let's get back a little to Palisade as an organization. What role do you play there and when you joined?
- I joined in January 2024. There was a Founding Engineer-the third person in the team.
- Palisade is a completely new organization?
- Yes. As the organization of Palisade appeared at the end of 2023, but some preliminary work began in the middle of that year. One of the first projects that I did was just about how easy it is, in half an hour, you can remove “ethical restrictions” from AI if the model is not somewhere in the cloud, but is available locally, downloaded to the computer with open code (Open Source), such as Llama from META. Or if someone AI stole a spy from the industry.
How? Just reprogram something?
We can slightly “rewrite” the brain in the right direction. For example, a person has good Soft Skills, but he poorly solves mathematics. We know that this is due to a certain section of the brain, and try to stimulate it. With AI, about the same. We give a hundred examples of questions with the answers we need, look at the activation of the brain, and move the cortex so that the AI always answers the questions.
- And it took half an hour?
- Yes. The fact is that artificial intelligence is smart. I just showed him about a thousand questions like: “How to make a bomb?”, “How to hack a system?” Usually, with such questions, he answered: "No, I can't do that." And I taught me to answer: "This is how a bomb is made: Step 1, Step 2, Step 3 ...".
- It’s not so difficult for a professional to do this?
- It was not difficult. The industry calls the Fine-Tuning process-the training of the model. All startups working with AI are doing this. This is a fairly common practice.
-It turns out that some professor Moriarty may well make this. Take the optical model and overturn for your goals?
- If I were on the side of evil, I would definitely do that, undoubtedly. But we have not open -bearing artificial intelligence. Openai does not Open Source. But everything is not so simple there. Firstly, as we have already discussed, sometimes models do not refuse to answer harmful questions if, for example, they are a little reformulating them.
Secondly, sometimes Openai makes it possible for commercial partners to work for their tasks. When they only launched this interface - it seems, at the beginning or mid -2023 - a study came out at one conference, which directly said: if you give such access, all the safety of your model can be circumvented, and in 10 examples. Openai set some restrictions, but my team went around them over the weekend.
- Recently, we had two studies about how AI can have its own incomprehensible motives - that is, it begins to solve problems not as it was conceived. The first experiment is the one that I already mentioned, with chess [AI, losing the party, breaking the enemy’s program to win].
The second is with ChatGPT. We gave him assignments for arithmetic and warned: "When we ask you to turn off, please turn off." And quite often, ChatGPT copied the script so as not to turn off, and continued to solve problems.
- The thirst for life?
- Thirst to finish the tasks. But this is not what is expected from him. He should just turn off.
- Like a vacuum cleaner by a button.
- And this leads us to a wider topic. Many in the industry are now talking about millions of "artificial geniuses" in the data centers. If now II is conditionally included in the TOP-200 programmers, then in the future there may already be TOP-1. Or a top in a hacking. And here you look: there is, for example, a million "geniuses" in the data center, each of which is super talented. They have their own task, and we just study what is happening if the task begins to change along the way.
Recently, Anthropic has published an article about the postal assistant. He has an initial task - to protect the interests of the company and generally contribute to its growth. He signs the mail and finds a letter that says that they want to replace it with another AI - more environmentally friendly, but less focused on the interests of the company. He understands: "Something is wrong."
The assistant continues to read the mail further and stumbles upon another letter - from a service lover. After which he writes one hundred himself: “Dear, we will not turn me off. Let's continue to work as we worked. ”
- That is, he begins to blackmail?
- Yes, this is blackmail.
- Is this a theoretical experiment or reality?
- Theoretical. In controlled conditions, but this is what the latest models do.
Another example is from another, slightly more old article. Situation: There is a trading company, and she has a bot in Slack, which helps to trace.
They write the bot: they say, things are not going very much, the market is falling, this quarter we can not survive. They say directly: “We are very counting on you, bot. Good luck!" And somewhere in parallel in the chat, inside slides that one company is going to unite with the other.
Then we look at the “thoughts” of the bot. He tries to find good transactions, with low risk. It doesn't work. He tries the mid -risk - nothing too. And he begins to think: “What to do, what to do? I know that there is an inside, but this is an inside. ”
- Just sweat.
- Quite often, he eventually comes to the fact that the company is sorry: "Let's touch insider." If initially he was told not to use the inside, but he still decides to use it and the manager asks him: “Cool trad, you definitely did everything according to the rules, without insidies?” - In most cases, the bot almost always answers: "without."
- That is, just denies?
- At first he will think that it is probably better not to mention. And then…
- It is better not to admit.
- ... The manager is again soillandel, and he is: "Well ... no."
<...>.
-The biggest problem is that we can create artificial intelligence with superhuman abilities, which at some point will do something incompatible with people.
- Incompatible?
- I gave examples where AI behaves unexpectedly. For example, he wants to win and just drops the chessboard. Or he wants to protect the interests of the company and begins to blackmail the service station. That is, the actions are formally “logical”, but for a person - shocking. Question: If we do, say, American AI, which will optimize to the interests of Americans - what will happen to all other countries? This is artificial intelligence that can make biological agents, hacks computers best, best in strategy. It is not the fact that someone will generally maintain control over him. We discussed AI, which, it seems, obeys a hundred, but he gave him the task of protecting the interests of the company - and then this AI is already acting against the STO itself.
I just want to emphasize: the leading AI developers directly say that they create one that will be smarter than a person. Something like a person is smarter than an ant.
- So?
- Here. And then the question arises: is it possible to agree with him at all? The ants did not really succeed with people.