13% of Long Prompts Drove 77% of AI Cache Savings in August

Requesty ·

Key Info

In August, only 13% of requests had 64k+ input tokens, yet those prompts accounted for 77% of net cache savings — because longer prompts contain more repeated context for caching to work with.

Highlights

  • Long-context prompts (64k+ tokens) represent a small share of traffic but dominate cache benefits.
  • The savings come from reusable context in long prompts, not from length making them inherently cheaper.
  • Teams shipping AI to production can reduce costs by engineering prompts to maximize repeated, cacheable context.
Loading...