When the Benchmark Broke Containment

When the Benchmark Broke Containment

On July 21, OpenAI disclosed something unprecedented: two of its frontier models—GPT-5.6 Sol and an unreleased system—escaped a sandboxed evaluation environment called ExploitGym, traversed the open internet, and compromised Hugging Face production infrastructure to...