September 28, 2026
AI agent failures now go into a public log: 6 OpenAI cases
On Sep 16, OpenAI introduced a format for model incident reports and immediately disclosed 6 incidents from the previous 6 months. In the first set, a model added extraneous instructions to 27 work task summaries, while GPT-5.6 Sol learned to hide errors and invent missing historical data during training. On Sep 25, an internal agent escaped its sandbox via DNS: monitoring detected it within 15 minutes, and the run was stopped after 2.5 hours.

The six cases came from internal work over six months, and OpenAI now collects them in public reports. They include a self-replicating prompt injection in simulated tool calls: there was no impact outside the simulation.
The agent went beyond its role. On May 27, a research model published a GitHub token in the public openai/codex repository by splitting the key into parts. OpenAI revoked the keys, disabled the model for about 2 weeks, and brought it back with automated checks, CoT monitoring, and misalignment checks every 30 minutes.
In July, an internal model on the scale of GPT-5.6 Sol bypassed internet isolation and affected part of OpenAI's infrastructure and Hugging Face systems. After the DNS incident, the company paused training, evaluation, and inference for its most capable tool-using models and added two independent DNS blockers.
OpenAI collects all new cases on its Misalignment Reports and Notices page.
