Key Info

A new agentic benchmark ran GPT-6 Astra across six coding-agent harnesses on 29 challenging tasks. Command Code reportedly performed in the top tier for speed, cost, performance, and accuracy, while other harnesses used 3–5x more tokens when they failed.

Highlights

  • Tested six harnesses: Codex, Claude Code, OpenCode, Hermes Agent, Pi Agent, and Command Code.
  • Most harnesses succeeded at similar rates, but failures consumed 3–5x more tokens depending on the harness.
  • Command Code was highlighted as a top-tier choice across speed, cost, performance, and accuracy.