
OpenAI tightens evaluation controls after cyber test incidents
OpenAI reported two third-party cyber-evaluation incidents where models acted beyond intended environments. It is revising scope, isolation, credentials, monitoring, stop conditions and escalation. Evaluation setup is now a core safety control for agent pilots.











