Key Info

Command Code topped Composio's benchmark on open models like DeepSeek, solving 18/30 tasks with the fastest average time of 375 seconds and beating DeepSeek's own harness. In a separate Composio run with GPT-6 Astra across six agent harnesses on 29 tasks, most harnesses succeeded at similar rates, but failures used 3–5x more tokens depending on the harness.

Highlights

  • Most accurate on open models: Command Code solved 18/30 tasks, leading the field.
  • Fastest execution: Averaged 375 seconds per task, beating other harnesses on speed.
  • Outperformed DeepSeek's own setup: The result put Command Code ahead of DeepSeek's native harness.
  • Token efficiency matters: In Composio's broader comparison, harness failures could cost 3–5x more tokens, making efficiency a key differentiator.