Key Info

Some AI inference providers lure customers with cheap-looking prices but deliver poor prompt-cache hit rates, forcing you to consume up to 3x more tokens before you notice. High "99% cache rate" claims are a red flag: the author says only DeepSeek truly hits that today, so such providers are likely just reselling/wrapping DeepSeek.

Highlights

  • Cheap per-token prices can hide bad cache performance; the cost only becomes obvious after you collect enough usage data.
  • Cache hit rates are hard to verify upfront, making this an easy area for providers to mislead customers.
  • A claim of 99% cache hits is suspicious — in practice only DeepSeek is seen reaching that level, often meaning the provider is simply wrapping DeepSeek.
  • Before committing, benchmark real workloads and measure actual cache hits and token consumption.