As context windows scaled from 8k tokens to 1 million and beyond, engineering teams quickly encountered an uncomfortable truth: sending 100k tokens of documentation, schema files, and conversation history on every API call is economically unsustainable.
The quadratic computational complexity of self-attention means that re-computing attention states across identical prefixes costs both high latency and astronomical dollar bills. Context caching fundamentally solves this by persisting intermediate Key-Value (KV) tensors directly in GPU memory or distributed host RAM.
The Mechanics of the KV Cache
During transformer inference, generating token $t_{n}$ requires computing keys and values across all previous tokens $t_1 \dots t_{n-1}$. Without caching, the model must re-run attention over the entire prompt for every single forward pass.
With KV-caching:
- The prefill phase calculates and stores Key and Value activation matrices in GPU memory.
- The decoding phase computes attention exclusively against the newly generated token and appends to the existing KV tensors.
[System Prompt + Static Tools + Codebase Context] ---> Cached in GPU VRAM (Prefill done once)
│
▼
[User Query 1] ───> Instantaneous response (<150ms TTFT)
[User Query 2] ───> Instantaneous response (<150ms TTFT)
Comparing Provider Implementations
Modern inference providers implement context caching using distinct architectural optimizations:
| Feature / Metric |
Gemini Context Caching |
Anthropic Prompt Caching |
vLLM PagedAttention (Local) |
| Minimum Cache Token Size |
32,768 tokens |
1,024 tokens |
Block size (16–32 tokens) |
| Input Token Discount |
75% discount |
90% discount on cache hits |
100% elimination of re-compute |
| Cache Lifetime |
Explicit TTL (default 1 hr) |
5-minute rolling TTL |
LRU cache in host / GPU VRAM |
| Time-to-First-Token (TTFT) |
~80% latency reduction |
~85% latency reduction |
Near zero prefill latency |
Prompt Structure: Ordering for Maximum Cache Hits
To exploit prefix caching, prompt structures must be strictly deterministic from top to bottom. Because transformer attention caches match left-to-right prefix hashes, any dynamic variable placed early in the prompt invalidates the entire cache downstream.
Anti-Pattern (Breaks Cache)
{
"messages": [
{ "role": "system", "content": "You are an assistant. Current time: {{ timestamp }}" },
{ "role": "system", "content": "{{ massive_repo_documentation_100k_tokens }}" }
]
}
Production Best Practice (100% Cache Hit)
{
"messages": [
{ "role": "system", "content": "You are a deterministic system architect." },
{ "role": "system", "content": "{{ massive_static_docs_100k_tokens }}" },
{ "role": "user", "content": "Query: {{ prompt }} | Timestamp: {{ timestamp }}" }
]
}
Economic Impact
For high-throughput systems processing millions of queries daily, prompt caching converts an otherwise unviable product margin into high unit economics. Combined with PagedAttention and prompt compression techniques like LLMLingua, infrastructure costs drop by upwards of 75% while user-perceived responsiveness improves by an order of magnitude.