Command Code is fastest DeepSeek inference in the world.
Better
Faster
Cheaper
All the work we put into tool repairs have and harness + inference engineering is paying dividends now.
Our mission is to be inference-native.
Let's go, GOAT 🐐
引用推文
DeepSeek V4.1 Flash is not one speed.
Same model. Same Hermes setup. Very different throughput depending on the route.
I pulled 2,679 real calls with 2,000+ output tokens across four providers.
Median effective throughput:
CommandCode GOAT: 197 tok/s
OpenCode Go: 186 tok/s
ClinePass: 122 tok/s
Ollama Cloud: 75 tok/s
That’s a 2.6x difference between the fastest and slowest route for the same model, same harness.
These are full request wall-time measurements, not provider-reported raw decode speeds. I filtered to long responses so prefill matters much less.
Provider choice matters way more than I expected.