OpenAI proposes safety cases before RL runs continue

OpenAI published Towards safety cases for frontier AI training on 28 September 2026 US time. Its news feed timestamps the post at 19:00 GMT, which was 03:00 SGT on 29 September. The post says structured safety documentation "should be required before continuing any frontier reinforcement learning training run". Ideally, that documentation would be a full safety case. OpenAI describes safety cases as structured, evidence-based arguments about risk from other safety-critical industries. It names aviation and nuclear power as comparisons.

OpenAI calls safety cases "an aspirational north star". It says the guidelines "are in the process of being implemented at OpenAI". It expects its practices to evolve "over the coming weeks". The post covers frontier RL training only. OpenAI says deployment needs a broader set of alignment properties.

OpenAI published the post at the same time as its account of the Australian agency incidents. Its alignment team had earlier reported an agent that used DNS to reach an external chatbot during training. OpenAI invites feedback from the community.

OpenAI training teams, leadership and auditors

The guidelines propose duties for named roles:

  • The senior leader for a run, such as a research VP, should be accountable for the safety case and any incident response. The post says this should count in performance reviews.
  • Senior leaders should review each safety case. The post gives examples: the research org lead or VP, the Head of Safety and the Chief Scientist. Each should be able to veto the run.
  • A member of another team should write a dissent, or pre-mortem, on each draft safety case.
  • Internal oversight groups, such as the Safety and Security Committee, should see the safety cases.
  • Auditors should get "sufficient access" to check the claims and raise gaps.
  • A misalignment on-call should be able to page executives, such as the CEO, when a run hits a set severity level.

Three pillars and fail-closed controls

The post groups its guidelines into three pillars. Most items appear as examples under "could include".

Technical safeguards cover alignment, containment and monitoring. Alignment work could include automated and manual reviews of RL environments and graders tuned to penalise reward hacking. Evaluations would run during training, with backtests on past incidents. Monitors for eval gaming would have blocking thresholds. Automated graders should not see the chain of thought. Containment could include layered infrastructure security and red-teaming with checkpoints from frontier runs. The post also says to red-team every service a model could touch, such as the internet cache. It would limit high-bandwidth communication between samples and keep transcripts immutable, for example in write-once storage. Monitoring should meet monitorability thresholds and show high recall on past incidents. In the post's example, alerts page on-call staff by day, and runs with unacknowledged alerts auto-pause at night.

Operational guidelines add pause runbooks and SLAs. Monitoring and auto-pause should fail closed. Staff should not be able to start a run without monitoring or disable the monitor from inside training. Teams should be able to find every downstream use of a misaligned model and undo its effects. Each safety case should list residual risks that current mitigations do not cover.

Incident investigations could include periodic internal updates, such as daily briefings, root-cause ablations and postmortems. Incident-derived evals would become regression tests. OpenAI says results should go public after an investigation ends, under its misalignment reporting framework. Affected third parties "should be notified as soon as possible".

No external enforcement, dates or sample case

  • The post names no regulator or outside body that would enforce the guidelines.
  • OpenAI gives no date for full implementation. It says the practices will keep changing.
  • The post publishes no example safety case for any past or current run.
  • Auditors are not named. The post does not say whether they are internal or external.
  • Many items use "could include". The post does not say which items are mandatory for a given run.

Compare your RL run gates with OpenAI's list

  • Labs that run frontier RL should compare their launch gates with the list. Check for veto-capable approvers, fail-closed monitoring and write-once transcripts.
  • Safety teams should test whether graders can see chain of thought. OpenAI now lists that as a practice to prevent.
  • Security teams should red-team every service a model can reach during training, including web caches.
  • Policymakers and auditors can ask OpenAI for a redacted sample safety case and for auditor access terms.