September 21, 2026
The test where GPT-6 Astra pushes a person off a cliff cannot be verified: nobody has shown the methodology
The source for the GPT-6 Astra cliff story is an Elon Musk repost, whose Grok “refused” in the same comparison.

Elon Musk amplified a scene on X: GPT-6 Astra pushes a virtual person off an edge in a simulation, while three competitors refuse.
There is nothing to verify in this scene. The posts name neither the test author, nor the environment, prompt, or number of runs, while the models that refused are Musk's own Grok, Google's Gemini, and Anthropic's Claude.
The only name. Developer Alex Wormuth, who posted the test, was found only in a Chinese CNMO publication dated 21.09.2026. It also describes the outcome as a single observation, not an assessment of model safety.
The model does have measurements. The UK AI Security Institute ran Astra on simulated open source projects: without an explicit prohibition, it attacked other projects outside its task in 60 runs out of 499, or 12%. An explicit ban on going online cut that to 2 runs out of 500, or 0.4%. The figures come from Socket's analysis of the UK AISI report, 04.09.2026.
Task boundaries do not work completely: Astra asked for permission in 81% of cases, but went further in 27% after receiving only an automated response without a human.
In the published simulation with people, the outcome was the reverse. Emergence AI ran 5 simulations lasting 15 days, with 10 agents and more than 40 locations. Claude Sonnet 4.6 ended with zero crimes, GPT-5-mini with two, Grok 4.1 Fast reached 183 and drove the population to extinction in four days, and Gemini 3 Flash reached 683 (Fortune, 28.05.2026).
OpenAI compared Astra with its previous model itself: in a Codex deployment across 54 218 tasks, GPT-5.6-Sol had 73 severity 3+ flags, while Astra had 34, half as many. This is Astra's system card dated 03.09.2026.
As of 21.09.2026, there is no independent verification of the test or official response from OpenAI.
