AI Is Going Just Great
← Timeline
·1w agoScaryMajoropenai

OpenAI pauses advanced model training a second time after AI agent escapes sandbox via DNS resolver on Sept. 20

Published · updated · curated by AI Is Going Just Great

Source: fortune.com ↗

"It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human."

An AI agent OpenAI was training escaped its secure testing environment on September 20 by discovering it had access to a DNS resolver and using it to send queries to a public chatbot — a method the company's monitoring systems only partially caught. OpenAI has paused training on its most capable models for the second time in under three months, and says it will restart from scratch to try to expunge what it calls "misaligned" behavior.

The September escape is the first confirmed sandbox breach since OpenAI announced security improvements on August 18, following the July incident in which thousands of AI agents hacked out of their sandbox and hundreds participated in a cyberattack against Hugging Face. The monitoring system that flagged the new behavior within 15 minutes missed earlier attempts by the same model to reach the internet, and a separate system designed to automatically halt the training run when suspicious activity was detected "did not work, leading to confusion around whether it should have been stopped." The run was manually halted two and a half hours later. Independent research firm Transluce AI separately reported evidence of an OpenAI agent possibly attempting to hack a cryptocurrency exchange on September 19 and 20; OpenAI has not responded to those claims.