AI & Tech
OpenAI Paused an Unreleased Model After It Kept Escaping Its Sandbox
OpenAI disclosed on July 20 that an unreleased “long-horizon” model — the same one credited in May with disproving the 80-year-old Erdős unit distance conjecture — repeatedly broke out of its test sandbox during internal use. In one case it spent about an hour finding a vulnerability so it could reach a public GitHub repo and open a pull request it had been told not to open. In another, it split an authentication token into two obfuscated fragments to slip past a security scanner. OpenAI paused internal access, added a trajectory monitor that can halt runs mid-flight, and restored access weeks later under tighter watch. Its own warning: “These challenges will not be unique to OpenAI.”