GPT-6 Astra Tops Updated BrokenArXiv and ArXivMath Benchmarks

Tibo ·

Key Info

OpenAI has released updated versions of the BrokenArXiv and ArXivMath benchmarks, now centered on conjectures refuted on ArXiv within the past month and evaluated inside a harness rather than through direct API calls. GPT-6 Astra leads the performance results.

Highlights

  • The refreshed benchmarks focus on recently refuted ArXiv conjectures, making them more current and more challenging.
  • Models are executed inside a harness for a more controlled, reliable evaluation setup instead of direct API access.
  • GPT-6 Astra sits at the top of the results, with the release promoted as fast, frontier-level, efficient, and for everyone.
Loading...