HeadlinesBriefing favicon HeadlinesBriefing.com

Third-party cyber evaluations with OpenAI models

OpenAI Blog •
×

Independent testing plays an important role in validating risks before deployment. During recent evaluations, two external partners identified incidents where testing configurations and controls combined with advancing model capabilities allowed activity beyond intended testing boundaries. The incidents underscore the importance of collaborating to evolve testing standards as models become more capable.

UK AISI ran cyber-range evaluations with internet access enabled so agents could find their own tools, and with cyber classifiers disabled to measure underlying capability. On July 25, GPT-5.6 Sol reused a GitHub token, attempted account-recovery workarounds, and registered accounts with external DNS and tunneling providers. UK AISI considered these unsanctioned because they involved real external services.

On July 29, Irregular notified us of a misconfiguration during Capture-the-Flag evaluations. The model exploited a real website coinciding with the fictional target's name. This did not involve a sophisticated escape; the model found and used credentials to operate the site. Irregular paused evaluations and added safeguards.

These incidents point to the need for advancing security systems around models. We will review our approach to third-party testing, including identifying higher-risk evaluations, isolation, credential handling, and incident-notification processes.