Key Info
Composio's benchmark, built around its extensive library of connectors and MCP tools, is positioned as a standard test of AI performance on everyday knowledge work. Command Code ranked #2 on this benchmark when running GPT-6 Astra, showing it can handle practical work tasks, not just coding.
Highlights
- Composio's benchmark is touted as the reference for knowledge work, using the largest available set of connectors and MCP tools.
- Command Code secured the #2 position on the benchmark for the GPT-6 Astra model.
- The result underscores that Command Code's performance extends beyond coding into general knowledge-work tasks.