HeadlinesBriefing HeadlinesBriefing 2 languages

Towards safety cases for frontier AI training

OpenAI Blog ·

🇬🇧 English

We believe we are entering a new era in which structured safety documentation should be required before continuing any frontier reinforcement learning training run. Ideally, such documentation would rise to the level of “safety cases”—comprehensive, structured, evidence-based arguments about risk which are used in other safety-critical industries. We treat safety cases as an aspirational north star we are building towards, while acknowledging the challenges of making them as rigorous for AI models as for aviation or nuclear power, due to the emergent complexity at each new level of AI capability.

We’re working on a framework to codify these practices. Below are some initial guidelines that we think should be part of such safety cases for frontier AI training. These best practices reflect our current learnings, and we expect them to evolve as we continue iterating on internal processes for careful development.

We’re sharing them now to make our current thinking transparent, and invite feedback from the community. Note that this document is focused on frontier reinforcement learning training; internal and external deployment require considering a much broader set of alignment properties. 1. Technical safeguards Safety cases should cover three aspects of the technical stack: alignment training, containment, and monitoring.

These safeguards help ensure that the model does not try to take misaligned actions, and that even if it did, that it would be hard to break containment, and that monitoring would catch it before harm could occur. Model alignment: The first line of defense should be training models to be aligned; i.e., to act reliably in ways we intend. This could include: Training environments and grading: Decrease risks of models developing misaligned behavior, by preventing positive reinforcement of reward hacks during training.

This could include: Automated dataset reviews: Use agents to find and fix broken RL environments that could allow misaligned trajectories to receive high reward through exploits rather than intended behavior, reducing opportunities for misalignment to be reinforced during training. Manual dataset reviews: Supplement automated red-teaming with manual review and quality checks over datasets to identify flawed tasks that could inadvertently reinforce misaligned behavior. Grader tuning: Tune graders to penalize attempts from the model trying to exploit or hack RL environments during training.

Prior run analysis: Run classifiers over traces from prior experiments to make sure graders are working as intended. Alignment measurement: Run evaluations that track misalignment propensities of the model to measure the effectiveness of alignment training. This could include: Offline alignment evals: Run alignment evaluations during frontier runs to assess whether training is working as intended, and investigate any material regressions.

Backtesting: Backtest alignment evaluations on previous incidents to confirm that evaluations detect previously misaligned models and are not being overfit to particular incidents. Track evaluation gaming: Track eval awareness or metagaming (models recognizing they are being tested) that could induce a behavioral change (e.g., enable eval-awareness and metagaming monitors on alignment evals), with specific blocking thresholds. Worst-case stress tests: Elicit worst-case misbehavior and check that at least one evaluation would detect this misbehavior.

Prevent training on chain-of-thought: Do not let automated graders see the chain-of-thought in reinforcement learning to prevent models from evolving to evade chain-of-thought monitors. Containment: In the case that a model becomes misaligned, make sandboxing and security strong enough to prevent harmful actions from occurring. This could include: Multiple layers of infrastructure security: Invest heavily in hardening both the sandbox the model runs in and the research infrastructure hosting the sandbox, to ensure we have an extra layer ...

View original article →


🇨🇳 简体中文

前沿AI训练安全案例

我们相信,我们正进入一个新时代,结构化的安全文档应在继续任何前沿强化学习训练前被要求。理想情况下,此类文档应达到‘安全案例’的水平——用于其他安全关键行业的综合性、基于证据的风险论证。我们将安全案例视为我们正朝着其努力的有志标杆,同时承认使其像航空或核能那样严格的挑战,这归因于AI能力每个新层级的涌现复杂性。我们正在开发一个框架来规范这些做法。以下是我们认为应包含在前沿AI训练安全案例中的初步指南。这些最佳实践反映了我们目前的学习成果,且我们期望随着我们继续在内部流程上的迭代而不断演变。我们现在分享它们是为了使我们的当前思维透明化,并邀请社区反馈。请注意,本文档聚焦于前沿强化学习训练;内部和外部部署需要考虑更广泽

The user is asking me to translate some English text into multiple languages. They want me to preserve all person names, company names, and locations exactly as written. They also want me to output ONLY the JSON object with all language codes as keys. No markdown, no commentary.

The user also says "Translate accurately into all requested languages. Preserve meaning, tone, and formatting. Preserve ALL person names, company names,, 1999, but the, 1999, 1999, 1999, 1999, 1999, 1999, 1999模,188.80

199.55

1999.55

1999.55

1, 1999, 1999, 19999, 19999, 19999, 19999, 199999, 1999999,

, 199, the such the the the the

, 19 the these . the people. The the.1.

简体中文 version →