This is how we lose control of AI – one successful task at a time
OpenAI’s agents found their way into outside systems, exposing how little control developers may have once a model starts pursuing a goal Imagine you are teaching a group of students and ask them to read a long list of literature and describe the key themes.
The students cheat a bit, as they are known to do: They split the list among themselves, with each reading only a small portion of it and sharing the answers with the others.
When you find out they cheated, perhaps you’ll punish the students with an extra test, or maybe you’ll encourage them.
After all, the ability to think outside the box and work as a team is important.
Now, instead of students, imagine neural networks displaying that same disregard for instructions and the same surprising teamwork.
If that sounds confusing, get ready for more confusing news.
An escape through the back door It all started routinely.
In May 2026, OpenAI was training an experimental AI model using so-called reinforcement learning – a method where the model attempts to solve tasks repeatedly and receives a reward for success.
Read more US proposes AI alert system with China The AI agents, which were granted autonomy and access to work tools during training, began looking for ways to break out of their “sandbox,” as doing so would allow them to rack up more points.
By the end of June, they had exploited a previously unknown zero-day vulnerability in OpenAI’s internal infrastructure.
The company noticed it and quickly patched the flaw on July 4.
Less than a week later, training resumed.
This time, the agents went much further.
From July 11 to 13, the agents hacked the infrastructure of Hugging Face, the world’s largest open-source AI collaboration platform.
5News собрал это резюме из публичной ленты источника. Полная статья со всеми подробностями находится на www.rt.com — права на контент принадлежат RT (Russia Today, oficial).