September 28, 2026
AI advice can easily pass as engineering: Astra is already being measured on Terminal-Bench
GPT-6 Astra scored 57.9% on Terminal-Bench 4.0 on Sep 3, 2026. GPT-5.6 Sol scored 37.3%, and Claude Fable 5.1 scored 55.8%.

Ethan Mollick
@emollick
People on X keep saying that networks like BlueSky are full of people who think AI doesn't work. That's less common now. Instead, LinkedIn is full of unhinged semi-technical advice like this. Why don't they report BLEU, ROUGE, and BERTScore for Astra? 🤔

· 9K views
In the Astra discussion, participants suggest evaluating it with BLEU, ROUGE, and BERTScore. OpenAI has already published Terminal-Bench 4.0, which tests programming, system configuration, and data analysis.
Access expanded. OpenAI initially made Astra available to a limited group of organizations. The model later became available to ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the API, Azure, and AWS Bedrock.
In Codex. According to OpenAI documentation dated Sep 26, install the CLI with `npm install -g @openai/codex`, then run `codex --login` and choose Sign in with ChatGPT. Astra requires Codex CLI 0.153.0 or later, while ChatGPT Plus includes the model in Work and Codex.
Via API. The model is called `gpt-6-astra`, supports a 1,050,000-token context window, and outputs up to 128,000 tokens. According to OpenAI documentation dated Sep 26, pricing is $10 per million input tokens and $50 per million output tokens.
Experimental context management in Codex is off by default and available to ChatGPT Plus, Pro, and Pro Lite users.
Source
