01 What happened
OpenAI released GPT-6 Astra, a model demonstrating high performance across key benchmarks including FrontierMath and ARC-AGI-3, while introducing new alignment safeguards.
02 Key details
- OpenAI reports scores of 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.
- In an OSWorld 2.0 simulation, Astra completed tasks in about 40 minutes versus 75 minutes for GPT-5.6 Sol, with a higher task score.
- In one test of crossing an authorized task boundary, Astra did so in 0% of cases versus 48% for GPT-5.6 Sol without production safeguards.
- OpenAI reports that the model can identify previously unknown vulnerabilities, increasing the need for controls on its use.
03 Why it matters
This release marks a shift toward highly autonomous technical workflows, highlighting the dual-use nature of models capable of both vulnerability discovery and defense.
04 Who it matters to
Machine learning engineers, software developers, information security specialists and enterprise AI customers.
Original sourceOpenAI