GPT-6 Astra Tops Updated BrokenArXiv and ArXivMath Benchmarks
Key Info
OpenAI has released updated versions of the BrokenArXiv and ArXivMath benchmarks, now centered on conjectures refuted on ArXiv within the past month and evaluated inside a harness rather than through direct API calls. GPT-6 Astra leads the performance results.
Highlights
- The refreshed benchmarks focus on recently refuted ArXiv conjectures, making them more current and more challenging.
- Models are executed inside a harness for a more controlled, reliable evaluation setup instead of direct API access.
- GPT-6 Astra sits at the top of the results, with the release promoted as fast, frontier-level, efficient, and for everyone.