Key Info
On a 200-case synthetic benchmark spread across 30 task types, the Jev decision model matched popular LLMs on classification accuracy while ranking as the second-cheapest option, just behind Qwen3.8 Flash.
Highlights
- All five tested models landed within a handful of correct cases of each other; reasoning was turned off for LLMs, except GLM 5.3 Flash, which required reasoning and ran at low effort.
- Jev was cheaper than the three other LLMs and only slightly more expensive than Qwen3.8 Flash.
- Caveats: the 200 cases are synthetic and evenly distributed across 30 task types, and DeepSeek and GLM used default provider routing, so their tail latency reflects a mix of providers.