September 28, 2026
OpenAI wants to pause frontier-model training on alert: 3 layers of protection
On Sep 28, OpenAI published initial guidelines for a safety case during frontier-model training. The company proposes reviewing alignment training, containment, and monitoring, and pausing a launch when a high-priority alert is raised.
OpenAI previously had no published guidance for this stage of training. Now, a safety case must pass leadership review, and every reviewer has the authority to stop a training run.
The trail remains. OpenAI proposes immutably storing agent transcripts and not showing the model's chain-of-thought to automated graders. After a serious incident, the company advises turning investigation findings into regression tests.
OpenAI says it is already implementing the guidelines internally and will develop them further in the coming weeks.
Source
