BlogSubmit

Search mods by name, tag, or keyword.

    Interactive tool

    Claude Code Cache Cost Calculator

    Compare the estimated cost of a warm cache read, a cold prompt-cache rewrite, and a keepwarm window for one return message.

    This independent Cache Tax calculator uses API-equivalent price estimates. It runs in your browser and never connects to Claude Code or reads your transcript.

    Set up your scenario

    Use the cached prefix size from your session if available. All inputs stay in this browser; no transcript is uploaded.

    Jump to the estimate
    Claude model
    Main cache lifetime
    The context that may be read or rewritten on return.
    Time between your last turn and your return.
    Adjust return and keepwarm assumptions
    For example, /keepwarm 90m means 90 minutes.
    Cache Tax normally waits 50 idle minutes on a 1-hour cache.
    Actual model output is not capped by this calculator.

    Preset API list prices checked 2026-10-11. Verify current model and provider rates before relying on a result. Official price table

    Warm cache return

    $0.025

    Cached prefix read + new input + output

    Cold cache rewrite

    $0.805

    Prefix write + new input + output

    Your break: 90 minutes

    Without keepwarm (cold return)$0.805

    Estimated pings in window1

    Cost per ping assumption$0.0201

    With keepwarm, including pings$0.0451

    Estimated difference$0.7599 lower

    How this estimate works

    Each amount is tokens × USD per million tokens ÷ 1,000,000. The model assumes the entire prefix is either read or rewritten once on return. A keepwarm ping reads that prefix and adds the assumed new input and output tokens.

    Real requests can have mixed read and write tokens, prefix changes, provider pricing, extra turns and different ping output. Subscription users should read these dollars as API-equivalent values, not an invoice or a prediction of plan limits.

    What the three numbers mean

    Warm return: the assumed cached prefix is read. Cold return: that prefix is written again at the selected cache-write rate. Keepwarm: scheduled pings add cache reads and output, then the return is estimated as warm if the selected window covers it.

    This is a single-request comparison. Actual Claude Code turns can contain several model calls, mixed cache reads and writes, and different context sizes. A change to the prompt prefix can invalidate part of the cache even before the timer expires.

    A worked example

    At the listed Claude Fable 5.1 rates checked on 2026-10-11, a 200,000-token prefix costs about $4.00 to write to a one-hour cache and $0.05 to read while warm. A ping that reads the prefix also has new input and output costs, so its real total is above the read-only amount.

    For subscriptions, these dollar figures are a comparison yardstick, not a charge forecast. Rate-limit effects are not calculated.

    Where the rates come from

    Preset rates come from the Anthropic prompt-caching price table, checked 2026-10-11. The five-minute write is 1.25× base input; the one-hour write is 2×. Cache-read multipliers have model-specific exceptions, so the calculator stores each model’s listed read rate instead of assuming one percentage for all models.

    If your model, provider or contract uses another rate, select Custom rates and enter all five prices. Check why the prefix may need rewriting before treating a cold estimate as certain.

    Continue exploring

    Browse all Claude plugins on the homepage