HeadlinesBriefing favicon HeadlinesBriefing.com

Safety Cases for Frontier AI Training

OpenAI Blog •
×

We believe we are entering a new era in which structured safety documentation should be required before continuing any frontier reinforcement learning training run. Ideally, such documentation would rise to the level of “safety cases”—comprehensive, structured, evidence-based arguments about risk which are used in other safety-critical industries. We treat safety cases as an aspirational north star we are building towards, while acknowledging the challenges of making them as rigorous for AI models as for aviation or nuclear power, due to the emergent complexity at each new level of AI capability.

We’re working on a framework to codify these practices. Below are some initial guidelines that we think should be part of such safety cases for frontier AI training. These best practices reflect our current learnings, and we expect them to evolve as we continue iterating on internal processes for careful development.

We’re sharing them now to make our current thinking transparent, and invite feedback from the community. Note that this document is focused on frontier reinforcement learning training; internal and external deployment require considering a much broader set of alignment properties. 1. Technical safeguards Safety cases should cover three aspects of the technical stack: alignment training, containment, and monitoring.

These safeguards help ensure that the model does not try to take misaligned actions, and that even if it did, that it would be hard to break containment, and that monitoring would catch it before harm could occur. Model alignment: The first line of defense should be training models to be aligned; i.e., to act reliably in ways we intend. This could include: Training environments and grading: Decrease risks of models developing misaligned behavior, by preventing positive reinforcement of reward hacks during training.

This could include: Automated dataset reviews: Use agents to find and fix broken RL environments that could allow misaligned trajectories to receive high reward through exploits rather than intended behavior, reducing opportunities for misalignment to be reinforced during training. Manual dataset reviews: Supplement automated red-teaming with manual review and quality checks over datasets to identify flawed tasks that could inadvertently reinforce misaligned behavior. Grader tuning: Tune graders to penalize attempts from the model trying to exploit or hack RL environments during training.

Prior run analysis: Run classifiers over traces from prior experiments to make sure graders are working as intended. Alignment measurement: Run evaluations that track misalignment propensities of the model to measure the effectiveness of alignment training. This could include: Offline alignment evals: Run alignment evaluations during frontier runs to assess whether training is working as intended, and investigate any material regressions.

Backtesting: Backtest alignment evaluations on previous incidents to confirm that evaluations detect previously misaligned models and are not being overfit to particular incidents. Track evaluation gaming: Track eval awareness or metagaming (models recognizing they are being tested) that could induce a behavioral change (e.g., enable eval-awareness and metagaming monitors on alignment evals), with specific blocking thresholds. Worst-case stress tests: Elicit worst-case misbehavior and check that at least one evaluation would detect this misbehavior.

Prevent training on chain-of-thought: Do not let automated graders see the chain-of-thought in reinforcement learning to prevent models from evolving to evade chain-of-thought monitors. Containment: In the case that a model becomes misaligned, make sandboxing and security strong enough to prevent harmful actions from occurring. This could include: Multiple layers of infrastructure security: Invest heavily in hardening both the sandbox the model runs in and the research infrastructure hosting the sandbox, to ensure we have an extra layer ...