September 18, 2026
A model's self-report does not guarantee safety: internal computation is not expressed in language
On September 2, 2026, James Mickens posted a preprint on the "linguistic unreadability" of LLMs. A model may say one thing while computing something else internally, so self-reports, chains of reasoning, and linguistic signals offer no complete safety guarantee.

On September 2, 2026, James Mickens posted a preprint on the "linguistic unreadability" of LLMs. A model may say one thing while computing something else internally, so self-reports, chains of reasoning, and linguistic signals offer no complete safety guarantee.
For AI sandboxes, this means relying on isolation and data-flow tracking, not only on the model's explanations.
Source
