01 What happened

OpenAI released GPT-6 Astra, a model demonstrating high performance across key benchmarks including FrontierMath and ARC-AGI-3, while introducing new alignment safeguards.

02 Key details

  • OpenAI reports scores of 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench.
  • In an OSWorld 2.0 simulation, Astra completed tasks in about 40 minutes versus 75 minutes for GPT-5.6 Sol, with a higher task score.
  • In one test of crossing an authorized task boundary, Astra did so in 0% of cases versus 48% for GPT-5.6 Sol without production safeguards.
  • OpenAI reports that the model can identify previously unknown vulnerabilities, increasing the need for controls on its use.

03 Why it matters

This release marks a shift toward highly autonomous technical workflows, highlighting the dual-use nature of models capable of both vulnerability discovery and defense.

04 Who it matters to

Machine learning engineers, software developers, information security specialists and enterprise AI customers.

Original sourceOpenAI