OpenAI proposes safety cases for frontier AI training runs
OpenAI's initial guidance calls for evidence-backed safety cases covering alignment, containment, monitoring, senior vetoes, automatic pauses, audits and incident investigations before frontier reinforcement-learning runs continue.
OpenAI proposed initial guidance for evidence-based safety cases that training teams would prepare before continuing frontier reinforcement-learning runs. The cases would document safeguards and residual risks, face review by multiple senior leaders with individual veto authority, and allow covered runs to be paused if new evidence invalidates their safety claims.
The company described safety cases as an aspirational goal, not a finished control system. OpenAI said it is developing a framework to codify the practices, the recommendations are still being implemented internally, and the guidance will evolve. The document applies to frontier reinforcement-learning training; OpenAI said internal and external deployment require consideration of a broader set of alignment properties.
Alignment, containment and monitoring
OpenAI organized the proposed technical safeguards into three layers: alignment training, containment and monitoring. Its alignment recommendations include automated and manual reviews of reinforcement-learning environments, graders tuned to penalize attempts to exploit those environments, and classifiers applied to traces from earlier experiments to check whether the graders work as intended.
The guidance also calls for offline alignment evaluations during training, investigations of material regressions, backtests against earlier incidents, blocking thresholds for evaluation awareness or gaming, and stress tests in which at least one evaluation detects deliberately elicited misbehavior. It recommends withholding model chain-of-thought from automated reinforcement-learning graders to avoid training behavior that evades chain-of-thought monitors.
Containment would mean hardening model sandboxes and the surrounding research infrastructure, then repeatedly red-teaming those systems and reachable services with checkpoints from frontier runs, subject to safety analysis. High-bandwidth communication between samples should be restricted to approved routes, while discoverable infrastructure side channels should be closed. OpenAI also recommends retaining immutable transcripts across training, evaluation and internal deployment for a reasonable period, giving investigators an unedited record.
For live monitoring, the company proposes explicit thresholds for monitorability and monitor evasion, held-out tests of whether monitors catch earlier incidents, and fresh evaluation data for newly identified risks. Priority alerts should trigger action within a defined service-level agreement through an on-call page or an automatic pause of the affected run. The publication does not specify the thresholds, response times or transcript-retention periods.
Vetoes, audits and fail-closed controls
Once a training team drafts a safety case, OpenAI recommends that someone from another team prepare a dissent or pre-mortem identifying gaps and giving a calibrated risk assessment. Multiple senior leaders should then review the case, with each holding veto authority. The examples are a research leader or vice president, the Head of Safety and the Chief Scientist. The senior leader responsible for the run should also be accountable for the case and any incident response, including through performance reviews.
If a newly discovered issue would invalidate a safety case, runbooks, technical controls and service-level agreements should make it possible to pause every run the case covers. OpenAI says noncompliant runs should be difficult for humans or agents to start, while monitoring and automatic pausing should fail closed: a run should not start without required monitoring, and a monitor should not be disableable from within training, evaluation or internal deployment.
The proposal calls for sharing cases with internal oversight groups and giving auditors enough access to test their claims and identify gaps. In separate guidance on third-party assessments, OpenAI defines a safety case as an evidence-supported structured argument that risks are adequately managed for a specified activity, with assumptions, uncertainties and remaining risks made explicit. The training guidance does not name the auditors, set an audit schedule or establish a public review mechanism for run-specific cases.
Teams should also trace downstream uses of a misaligned model, including data generation and grading, so its effects can be rolled back. Each case should list residual risks not covered by implemented mitigations, allowing decision-makers to explicitly decide whether to accept them.
Incident investigations
For severe misalignment incidents, OpenAI recommends periodic internal updates, safe routes for relevant employees to access raw evidence, experiments to identify the training dynamics behind the behavior, and operational and cultural postmortems. The proposed investigation process sits alongside OpenAI’s earlier misalignment disclosure framework. OpenAI also calls for new detection methods and incident-derived regression tests, eventual public disclosure of findings and operational changes, and prompt notification of affected third parties. A severity table and a misalignment on-call process should allow executives to be paged when a run reaches a defined level.
An independent 2024 paper on frontier-AI safety cases described the method as an early-stage governance tool and said its methodology, safeguards and review capacity needed further development before safety cases could reliably inform decisions about more capable systems. It did not verify OpenAI’s internal practices. OpenAI’s publication offers recommendations, not a completed safety case: it does not show that every control is operating, give a control-by-control implementation status or provide external verification of implementation.
More news

Former OpenAI safety transparency leader David Robinson resigns, urges nuclear-level safeguards

OpenAI seeks US-led global AI standards effort

OpenAI reportedly asked Congress about a coordinated AI slowdown
