Cache guide
Prompt Cache Rewrite in Claude Code
Why can a short “continue” message be expensive after a long coding session? The model may need the earlier prompt prefix written to cache again.
1. Build context
Claude Code sends the conversation history and tool context on each request. A matching prefix can be read from cache on later turns.
2. Cache stops matching
The cached entry can expire while you are away. A model change, compaction, or some tool changes can also change the prefix.
3. Next turn rewrites
The next request processes the unmatched portion and can create a new cache entry. The price depends on the actual write tokens and model.
Expiry and invalidation are different
Expiry: the matching server-side entry is no longer available after its inactivity lifetime. The Claude API’s default cache duration is five minutes, with a one-hour option; Claude Code’s effective main-session lifetime depends on configuration and billing path.
Prefix change: the entry can still exist, but a change earlier in the prompt stops the later portion from matching. Claude Code documents examples such as model switches and conversation compaction. It does not mean every minor repository edit rewrites the full prompt.
Partial hits: one request may report both cache reads and cache writes. The calculator’s all-warm and all-cold outcomes are useful comparisons, not an exact prediction of that mixed request.
How to investigate a costly return
1. Check the elapsed time
Note when the previous request started and which main cache lifetime applied. An idle gap at or beyond the TTL makes expiry plausible.
2. Check what changed
Did you switch models, compact the conversation, or change loaded tools? These can explain a miss even when the idle gap was short.
3. Read actual usage
When your provider or transcript exposes token usage, compare cache read tokens with cache creation or write tokens for the return request.
4. Estimate the difference
Enter the context size, model and TTL in our calculator. Treat the result as API-equivalent; your actual invoice or subscription limit may differ.
What Cache Tax changes
Claude Code already manages prompt caching. Cache Tax adds a cost warning before a cold send and an optional keepwarm window. It cannot stop a prefix change or guarantee that a scheduled request saves money. If your main issue is an idle gap, read the keepwarm setup guide after checking the TTL.
Primary references: Claude Code prompt caching and Claude API prompt caching.