HeadlinesBriefing favicon HeadlinesBriefing.com

OpenAI's Principles for Third Party Assessments

OpenAI Blog •
×

Frontier AI labs carry an immense responsibility in training, evaluating, and deploying models safely. Third party assessments are a critical part of balancing that responsibility, expanding opportunities for input on AI safety, keeping the world informed, and keeping labs accountable to clear and independently supported safety claims. As part of our efforts to pace the frontier, OpenAI is committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment. That access should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.

We have long worked with third party assessors at various stages of the model development and deployment process. We have also incorporated third party assessments into our Preparedness Framework practices and supported organizations and legislation that advocate for a more rigorous and accountable process. Throughout these engagements, we have provided deep forms of access, including information about our technical safeguards, visible chain of thought access, and unprecedented levels of confidential data and internal deployment access for incident response and monitor red teaming. The priorities and principles shared here focus on our engagement with independent assessment organizations in the private and non-profit sector on technical safety assessments. They complement our work with governments on testing and evaluation, where distinct roles and responsibilities may call for different approaches.

Making these assessments effective requires strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities. Labs have a responsibility to enable meaningful scrutiny while protecting sensitive information. Labs and independent assessors share the responsibility for getting this right, and should be operating with shared international standards for safety and security practices. Below, we propose four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work.

Priority areas for assessment: Third party assessments are most useful when they address specific, consequential questions: Does the evidence support a lab’s safety case and safety claims? Do evaluations adequately test the risks they are intended to measure? Do safeguards work under realistic conditions? The assessments described here are intended to take different forms depending on the safety questions being examined. We expect to support multiple assessments in parallel and over different periods of time, with some lasting weeks and others several months. While third-party assessments can also be part of pre-deployment work and may inform deployment decisions, the work described here is generally longer-term and launch-agnostic—focused on examining particular safety claims in depth over time. Throughout our priority areas and principles, we refer to safety claims and safety cases. What we mean by these terms is the following: We propose four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work. Independent assessment of safety cases, spanning training, evaluation, internal deployment and external deployment. Assessment of safety cases requires expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse and red teaming. Safety cases consist of claims including training, capability evaluations, and safeguards, which can be assessed as a whole or in parts (see priorities 2 and 3 below). Multiple assessors will likely need to examine different parts of the cases, drawing on their respective expertise. Together, their assessments should answer questions such as: Is the evidence for safety cases for training, evaluation, and deployments substantiated? Were the conditions of the safety case followed during training, evaluation, a...