Uncategorized

Prompt Caching Explained and How to Cut LLM Costs

Prompt caching can cut LLM input costs by up to 90 percent. See how it works, compare providers, and follow a tested six-step code tutorial today.

Introduction

Prompt caching explained and how to cut LLM costs is now one of the most practical topics for any team that pays per token. Anthropic reported that reusing a stored prompt prefix can reduce costs by up to 90 percent and latency by up to 85 percent on long prompts. The reason is simple, because most production requests repeat the same system prompt, tool definitions, and reference documents while only a small tail changes. Providers can store the computed state of that repeated prefix and skip the expensive work the next time it appears. This guide explains the mechanism, compares the major providers, and walks through a tested implementation with working code. It also covers the failure modes that quietly push cache hit rates toward zero, from timestamps in system prompts to reordered tool lists. By the end you will know how to structure prompts, measure hits, and decide when caching is worth its write premium.

Quick Answers on Prompt Caching and LLM Costs

What is prompt caching?

Prompt caching stores the processed form of a repeated prompt prefix so later requests skip recomputation. Providers bill those cached tokens at a steep discount, often around ten percent of the normal input price.

How much can prompt caching cut LLM costs?

Vendors report savings of up to 90 percent on cached input tokens and up to 85 percent lower latency. Real savings depend on prefix length, hit rate, and how much output text your application generates.

Where should I start with prompt caching?

Start with prompt caching explained and how to cut LLM costs basics: place static content first, dynamic content last, then verify cached token counts in every response.

Key Takeaways for Cutting LLM Costs With Caching

  • Prompt caching reuses the computed state of an identical prompt prefix, so only the changing tail is billed at the full input rate.
  • Cache reads typically cost about a tenth of normal input tokens, while cache writes cost a premium that pays back after one or two reuses.
  • Hit rates depend on prompt ordering: static instructions and tools first, volatile data such as timestamps and user input last.
  • Measure cached token counts on every response, because a silent cache miss looks identical to a working call until the invoice arrives.

Table of contents

What Is Prompt Caching in Large Language Models?

Prompt caching explained and how to cut LLM costs starts with one definition: providers store the computed state of a repeated prompt prefix and reuse it, billing those tokens at a discount.

An Interactive From AIplusInfo

What Would Prompt Caching Save Your Application?

Set your prefix size, traffic, and expected hit rate to see hourly and monthly input savings under four real pricing styles.

20,000 tokens

1K200K

120 requests

12,000

80 percent

0%100%

Five-minute write

Write feeRead fee

$0

Input cost per hour, no cache

$0

Input cost per hour, cached

0%

Savings on input tokens

$0

Estimated savings per month

No caching

With caching

Break-even hit rate: 22 percent.

Assumes an illustrative base price of 3 dollars per million input tokens and a 500 token uncached tail per request. Benchmark: Anthropic reported up to 90 percent lower cost on long prompts. Your real hit rate depends on prompt stability and traffic gaps.

How Prompt Caching Works Under the Hood

Building on that definition, the mechanism starts with how a transformer reads a prompt. Before it can generate a single output token, the model runs the whole input through every layer in a step called prefill. During prefill each layer computes a key vector and a value vector for every input token, and those tensors are what the model attends to later. Storing them is called a KV cache, and it is normally discarded when a request finishes. Prompt caching simply keeps that state around for a few minutes so a later request with the same beginning can skip the work.

A cached prefix is only reusable when the request matches it from the first token onward, because each token’s keys and values depend on every token before it. A single changed character early in the prompt changes every downstream vector, so the provider has nothing valid to reuse past that point. This is why Anthropic describes its behavior as prefix matching up to each cache breakpoint, and why the order of tools, system text, and messages matters so much. Providers usually hash the prompt in fixed-size chunks and look those hashes up in a fast store. Everything before the first mismatching chunk is a hit, and everything after it is a miss that must be computed and, if you ask for it, written back.

The savings come from skipping prefill compute, which is the expensive part of long-context requests. Decoding the output still happens normally, so caching never changes what the model writes, only how much work it repeats. Latency improves mostly on time to first token, because the model no longer needs to read the whole document before it begins answering. The price discount reflects the provider’s saved GPU time, while the write premium reflects the memory it must hold for you. If you want more background on the token accounting behind these bills, the primer on tokenization in NLP explains how text becomes billable units.

The Economics of Cached Tokens

Turning to money, the pricing model has three levers: the normal input rate, the cache write rate, and the cache read rate. Anthropic charges 1.25 times the base input price to write a five-minute entry and 2 times for a one-hour entry. Reads cost 0.1 times the base price, which makes the break-even point easy to compute. With the five-minute option, one write plus one read costs 1.35 units against 2.0 for two uncached calls, so a single reuse saves roughly a third. The one-hour option writes at 2 units, so you need at least three requests inside the window before it beats paying full price. The practical rule is that caching wins whenever a prefix is reused at least twice inside its lifetime and clears the provider’s minimum length. Output tokens are never discounted, so long answers dilute the overall saving.

Consider a 20,000 token system prompt at a base price of 3 dollars per million input tokens. Uncached, each call spends 6 cents on that prefix alone, so 1,000 calls cost 60 dollars. With a warm five-minute cache, one write costs 7.5 cents and the remaining 999 reads cost 0.6 cents each, for roughly 6.1 dollars in total. Teams should also think at the system level, because Manus observed an input to output token ratio of around 100 to 1 in its agents. In that regime a tenfold gap between cached and uncached input is the biggest lever available. Workloads with short prompts and long answers see the opposite result, since there is little prefix to reuse. For broader budgeting tactics beyond caching, the article on enterprise AI cost optimization strategies covers routing, batching, and governance.

Comparing Caching Across Major LLM Providers

Stepping back from theory, each major provider implements caching differently, and the differences change how you write code. Anthropic asks you to opt in, either with a top-level cache control setting or with explicit breakpoints on content blocks. Its documentation allows up to four breakpoints per request and checks roughly twenty blocks backward from each one when looking for a hit. Minimum cacheable lengths vary by model, running from about 1,024 tokens on many models to 4,096 on some smaller or newer ones. Usage fields report cache creation tokens, cache read tokens, and the uncached remainder, so accounting is explicit.

OpenAI took the opposite design philosophy, with caching applied automatically to long prompts. Its original launch announcement described a 50 percent discount on prompts longer than 1,024 tokens, matched in 128-token increments. Current documentation describes far deeper discounts on newer models, with cache reads at 0.1 times the uncached rate. An optional prompt cache key also helps route similar requests to the same cache. Retention depends on model generation: older models keep entries for five to ten minutes or an optional 24 hours, while the newest use a fixed 30 minutes. Responses report the number of cached tokens inside the input token details field.

Google and DeepSeek sit at other points on the spectrum. Gemini enabled implicit caching by default on its 2.5 models with a 75 percent discount at launch, and it still offers explicit cache objects when you want guaranteed retention. DeepSeek builds a disk-based cache automatically and reports hit and miss token counts in each response, but its documentation states that the system works on a best-effort basis. Cache entries there are cleared automatically after a few hours to a few days. No provider guarantees a hit, so every design in this article assumes a cache miss is always possible and that correctness never depends on a hit.

The comparison table later in this guide summarizes these differences side by side. In practice, the biggest trade-off is control versus convenience for the engineering team. Explicit systems let you pin the cache to specific prefixes and choose retention, which suits agents and batch pipelines. Automatic systems require no code changes, which suits teams that only want savings on repeated system prompts. Whichever you use, the discipline of stable prefixes is identical, and the deep dive on reducing LLM inference costs shows how caching fits with other optimizations.

Designing Prompts That Hit the Cache Every Time

Next, the highest-leverage skill is prompt layout, because every provider matches prefixes from the beginning. Put the content that never changes first: tool definitions, system instructions, style rules, and long reference documents. Put content that changes per session next, such as user profile data or retrieved passages. Put the content that changes per request last, meaning the user question and any timestamps or request identifiers. Google gives the same advice in its launch post, recommending that you keep the beginning of the request the same and append changing context at the end.

Most cache misses in production come from small, invisible edits to the prefix rather than from any provider limitation. Manus warns that many languages do not guarantee stable key ordering when serializing JSON, which can silently break the cache. A current date in the system prompt, a random request ID, a reordered tool list, or an extra trailing space all produce a brand new prefix. Even switching a model, changing tool definitions, or toggling features such as web search can invalidate stored entries, according to Anthropic’s documentation. The fix is deterministic serialization, sorted keys, and a rule that nothing volatile ever appears above the last breakpoint.

Long documents and few-shot examples deserve special treatment because they make the biggest prefixes. Place them once in a stable block and reference them by position instead of pasting variations into each request. If you use retrieval, avoid rebuilding the context in a different order for every question, since reordering chunks destroys the shared prefix. Short, sharp instructions also help, and the piece on why shorter prompts improve accuracy is a useful reminder that cheaper prompts are often better prompts. Finally, write a unit test that renders your prompt twice with different user inputs and asserts that the shared prefix is byte-identical.

Prompt Caching for Agents and Tool-Heavy Workflows

Beyond single-turn chat, agents are where caching pays the largest dividends. An agent loop resends the entire conversation, including every tool call and observation, on each step. A task with forty steps therefore bills the early context forty times if nothing is cached. An evaluation of over 500 agent sessions found that prompt caching cut API costs by 41 to 80 percent and improved time to first token by 13 to 31 percent. Those numbers, measured across three providers, are lower than the marketing maximums, which is a healthy reminder that agents also generate large volumes of new, uncacheable text.

The same study found that naive full-context caching can paradoxically increase latency, so caching everything is not the same as caching well. Its best results came from placing dynamic content at the end of the system prompt, avoiding dynamic function calling, and excluding volatile tool results from the cached region. The intuition is that every cache write has a cost, and writing a large block that will never be read again wastes both money and time. Good agent design therefore separates the stable scaffold from the growing transcript and puts breakpoints only where reuse is likely. Tool definitions are the classic trap, since adding or removing a tool mid-session invalidates everything after it.

Context growth adds a second pressure that caching does not solve alone. As transcripts lengthen, models can lose track of earlier facts, a problem described in the article on context rot in large language models. Teams respond with summarization or compaction, but a naive summary rewrites the prefix and discards the warm cache. Compaction therefore has to be designed together with caching rather than bolted on later. Otherwise every summary step turns a cheap cached read into an expensive full-price write. The Claude Code team addressed this by forking a cached call to summarize the conversation, so the summarizing request shares the parent’s prefix.

Prompt Caching Versus Semantic Caching and Other Cost Levers

Looking at the wider toolbox, prompt caching is only one of several ways to lower an LLM bill. Semantic caching stores complete answers and returns them when a new question looks similar enough, which skips the model call entirely. That approach saves the most per hit but carries a correctness risk, because two questions that look alike can deserve different answers. Prompt caching has no such risk, since the model still reads the new question and generates a fresh answer. The two techniques stack well, with a semantic layer in front for frequently repeated questions and prompt caching behind it for everything else. Open-weight efficiency gains stack too, as the story of how DeepSeek reduced compute costs elevenfold shows.

Caching should be applied after you have tried the cheaper structural fixes, such as trimming prompts, choosing a smaller model, and batching offline work. Model routing sends easy requests to small models and hard requests to large ones, which can cut spend more than any cache. Batch endpoints trade latency for discounts and suit overnight jobs like classification or evaluation runs. Function calling adds tokens through schemas, so the primer on function calling in LLMs helps you keep tool definitions compact and stable. Each lever has its own failure modes, so change one at a time and measure the effect on both cost and quality.

Self-Hosted Serving With Prefix Caching

Shifting from hosted APIs to your own GPUs, the same idea appears under the name automatic prefix caching. vLLM describes it as caching the KV state of existing queries so a new query can directly reuse it when it shares the same prefix. Internally the engine splits the token sequence into fixed-size blocks and identifies each block by a hash of its contents and its predecessors. When a new request arrives, the scheduler looks up its blocks, attaches any that already exist, and computes only the remainder. Enabling the feature is a single engine option, which makes it one of the cheapest optimizations available to a self-hosting team.

The benefit shows up in the same places as it does for hosted APIs: long documents queried repeatedly and multi-turn chats whose history keeps growing. Prefix caching only shortens the prefill phase, so it does nothing for workloads dominated by long generated answers. vLLM states this limitation directly, and it matches what teams see when they profile time to first token separately from tokens per second. A retrieval assistant that answers in three sentences from a 30,000 token manual gains enormously, while a story generator that writes 4,000 tokens from a short prompt gains almost nothing. Measure the ratio of prompt tokens to output tokens before you invest in tuning.

Other serving engines take related approaches to the same underlying problem. SGLang introduced RadixAttention, which stores cached prefixes in a radix tree so requests that branch from a common trunk share the trunk automatically. That structure suits agent frameworks and structured generation programs, where many calls begin with the same long scaffold and diverge near the end. Memory managers such as PagedAttention allocate cache in small pages, which reduces fragmentation and makes it cheap to share pages between requests. Cache capacity is finite, so eviction policy matters, and least recently used pages are typically dropped first when GPU memory fills.

Operations add a final layer of concern for distributed deployments. A cache lives on one worker, so a load balancer that scatters a session across replicas turns hits into misses. Manus recommends routing requests consistently with session identifiers so that follow-up calls land on the worker that holds the warm prefix. Sticky routing has a cost, because it can unbalance load when a few sessions are much heavier than the rest. Teams with very large contexts on modest hardware should also read about near-infinite memory approaches for generative AI, which attack the same bottleneck differently.

Measuring Cache Hit Rate and Real Savings

Once caching is switched on, measurement decides whether it is actually working. Every provider returns counters that tell you how many input tokens were read from cache, how many were written, and how many were processed fresh. From those counters you can compute a hit rate as cached read tokens divided by all input tokens. Track it per endpoint and per prompt template, because one healthy average can hide a single broken route. Record the effective input cost per request as well, since that converts a technical metric into a number finance teams understand.

Alerts turn the metric into a safety net for the whole engineering organization. The Claude Code team runs alerts on its prompt cache hit rate and declares a severity incident when it falls too low. A cache hit rate that drops overnight is almost always caused by a deploy that changed the prefix, so tie the alert to your release log. Typical culprits include a new timestamp format, a reordered tool list, a changed model version, or a feature flag that inserts text near the top of the prompt. When an alert fires, diff the rendered prompts from before and after the change and look for the first byte that differs.

Finally, validate savings with an experiment rather than trusting the arithmetic. Send a fixed sample of real traffic through the cached and uncached configurations and compare total cost, latency percentiles, and answer quality. Quality should be identical because the model receives the same tokens, but confirm it, since a subtle change in prompt layout can alter behavior. Compare results against the vendor’s headline claims and against your own break-even math. If you sell an AI product, the analysis in how AI agent pricing is evolving shows how cheaper inference can reshape what customers expect to pay.

Where Prompt Caching Falls Short

Despite the strong numbers, caching has real limits that deserve an honest accounting. The first is lifetime: entries expire after minutes, so sparse traffic pays the write premium repeatedly and never earns a read. A prompt used once an hour on a five-minute cache costs more than no caching at all. The second is fragility, because tiny prefix changes invalidate everything after them. The third is uncertainty, since DeepSeek states plainly that its system works on a best-effort basis and does not guarantee a hit.

Caching reduces the price of reading a prompt but leaves the cost of writing an answer exactly where it was. Output tokens usually cost several times more than input tokens, so applications that generate long responses see modest overall savings. Minimum lengths exclude short prompts entirely, and some providers set thresholds in the thousands of tokens on their smaller models. Caching can also encourage bloated prompts, since a cheap prefix tempts teams to stuff in more context than the task needs. Deterministic behavior helps here, and the piece on deterministic guardrails for AI agents shows how stable, predictable scaffolding also improves reliability.

Ethics, Privacy, and Shared-Cache Side Channels

Stepping into risk territory, caching introduces a subtle privacy question that most cost guides ignore. Researchers who audited commercial APIs used response timing to detect whether a prompt had been cached. Their paper reports that seven API providers share caches globally across users, which means one customer’s timing could reveal that another customer sent a particular prompt. The same technique exposed evidence that one provider’s embedding model is a decoder-only transformer, which had not been publicly known. A fast response is a signal, and signals can leak information to anyone who watches carefully.

The safest design assumes that anything placed in a shared cache could be inferred by a determined attacker, so private data belongs only in per-organization caches. Practical mitigations include choosing providers that isolate caches per account or organization, keeping personal data out of the static prefix, and placing user-specific text after the final breakpoint. Regulated teams should ask vendors how long cached state is retained and whether it is covered by data processing agreements. Isolation policies differ widely, so read each provider’s documentation rather than assuming a default. Security teams should also log which prompts are cached so incident responders can scope any exposure. The broader landscape of consent and data handling is covered in the overview of privacy challenges and solutions in AI.

Ethics also cuts in a positive direction, because cheaper inference widens access. Startups, schools, and nonprofits can afford long-context assistants once repeated reading costs a tenth as much. Teams should still be transparent with users about what is stored and for how long, even when the stored artifact is a mathematical state rather than text. Auditing your own vendor with a timing test is reasonable, provided you stay within the provider’s terms of service. Responsible adoption means capturing the savings while treating the cache as a data store with a security boundary.

The Future of Prompt Caching and Long-Context Economics

Looking ahead, every trend line points toward cheaper and easier caching. Anthropic’s documentation now lists minimum cacheable lengths as low as 512 tokens and read prices as low as 0.025 times the base rate on its newest models. OpenAI has moved toward default behavior, where newer models cache automatically and report both cached and cache-write tokens. Providers are also experimenting with longer or configurable lifetimes, from one-hour entries to 24-hour retention on some older models. Each step lowers the threshold at which a workload benefits.

As caching becomes the default, the competitive advantage will shift from turning it on to designing products whose prompts are naturally cache-friendly. Agent frameworks are likely to standardize on append-only histories, stable tool registries, and cache-aware compaction. Prompt libraries may treat prefix stability as a linting rule, in the same way that type checkers catch errors before deployment. Context windows will keep growing, and caching is what makes a million-token window affordable enough to use routinely. Lower prices will also raise total usage, so budgets may not shrink even as unit costs fall.

Sustainability belongs in the same conversation as cost when teams plan long-term adoption. Skipping repeated prefill work means fewer GPU cycles per answer, which reduces the energy attached to each request. The article on the surprising energy footprint of AI chatbots explores how large that footprint has become. Efficiency gains do not guarantee lower total energy if demand rises faster, so the net effect will depend on how usage grows. Teams that apply prompt caching explained and how to cut LLM costs early build habits, such as stable prefixes and hit-rate dashboards, that carry over to whatever comes next. The engineering discipline is more durable than any single provider’s pricing table.

Chart From AIplusInfo

Cached input costs a fraction of normal input

Relative price of input tokens by pricing style, where uncached input equals 100.

Source: Anthropic prompt caching documentation and OpenAI prompt caching announcement. Chart type: horizontal bar.

How to Implement Prompt Caching in Your LLM Application

In practice, this walkthrough of prompt caching explained and how to cut LLM costs works for any provider, with small changes to the request shape. Follow the steps in order, because each one depends on the layout decisions made before it. The examples use Python, but the ideas apply equally to TypeScript, Go, or any HTTP client. Every step includes working code or a concrete check you can run today. Budget about an afternoon for a first working version and a few days of observation before you trust the numbers.

Step 1 – Audit your prompt for stable and volatile parts

Start by printing a fully rendered request and marking every piece as stable, session-level, or per-request. Stable parts include the system instructions, tool definitions, policies, and reference documents that change only at deploy time. Session-level parts include the user profile, the conversation summary, and any files attached for the whole session. Per-request parts include the latest question, retrieved passages, timestamps, and identifiers. Anything you cannot classify is a candidate for removal, because uncertainty about a prompt segment usually means it does not belong near the top. A typical service can finish this audit in about 30 minutes.

A quick search of your codebase finds the most common prefix breakers. Look for date and time calls, random identifiers, dictionary iteration that depends on insertion order, and feature flags that add prompt text. The command below lists suspicious lines in a typical project layout so you can review them one by one. Fix each finding by moving the value to the end of the prompt or by making it deterministic. Save the classified layout in your repository so future changes are reviewed against it.

grep -rnE "datetime\.now|time\.time|uuid4|random\." prompts/ src/prompt_builder.py

Step 2 – Reorder the prompt so static content leads

Reorder the request so the order runs from most stable to most volatile. Tools come first, then system instructions, then reference documents, then session data, and finally the user turn. Anthropic’s hierarchy is tools, then system, then messages, and changes at any level invalidate that level and everything after it. That means a tweak to a tool description costs you the whole cache, so treat tool definitions as versioned artifacts. For example, a 3,000 token tool block belongs above a 500 token user profile. Pro tip: never put a timestamp, request ID, or username above your last cache breakpoint, and pass such values in the final user message instead.

Make serialization deterministic wherever you build strings from structured data. Sort dictionary keys, fix number formatting, and normalize whitespace so the same logical prompt always yields the same bytes. If you assemble prompts from templates, render them in a pure function that has no access to the clock or random number generators. Version the template with a hash you can log, which lets you correlate cache behavior with prompt changes later. This single step accounts for most of the hit rate improvement teams report. For advice on writing the instructions themselves, see these LLM prompting tactics for professionals.

Step 3 – Add cache breakpoints with the Anthropic SDK

For Anthropic models, mark the end of the stable prefix with a cache control entry on the last block you want cached. The example below caches a long system prompt, and the default lifetime is five minutes that refresh on every hit. Confirm your prefix exceeds the minimum cacheable length for your model, otherwise the marker is silently ignored. You can use up to four breakpoints per request, which is enough to separate tools, system text, documents, and conversation history. Start with one breakpoint and add more only when measurements justify them. Prefixes of at least 1,024 tokens qualify on most models.

import anthropic

client = anthropic.Anthropic()
STATIC_SYSTEM = open("system_prompt.txt").read()

response = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=512,
    system=[
        {
            "type": "text",
            "text": STATIC_SYSTEM,
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=[{"role": "user", "content": "How do I reset my password?"}],
)
print(response.usage)

Choose the one-hour option when requests for the same prefix arrive less often than every five minutes but more than a few times per hour. It is requested by adding a lifetime field of one hour to the same cache control entry, and it costs twice the base input price to write. Remember the break-even math from earlier: the longer lifetime needs at least three requests to pay off. Many teams mix both, using a one-hour entry for tools and system text and a five-minute entry for the growing conversation. ProjectDiscovery used exactly this split, as the case study later in this guide describes.

Step 4 – Use automatic caching and cache keys on OpenAI-style APIs

OpenAI applies caching automatically once a prompt is long enough, so the work here is about layout and routing rather than markers. The OpenAI caching guide says a hit requires the entire rendered prefix to match, including instructions, tools, and history. Passing a prompt cache key groups similar requests so they route to the same cache, which raises hit rates for high-volume applications. Choose a key that identifies a template and version, not a user, so many users can share one warm entry. For example, a key such as support-bot-v3 groups every request from one template version. The snippet below sends a request with a key and prints how many tokens were served from cache.

from openai import OpenAI

client = OpenAI()
STATIC_INSTRUCTIONS = open("instructions.txt").read()

response = client.responses.create(
    model="gpt-5",
    input=[
        {"role": "developer", "content": STATIC_INSTRUCTIONS},
        {"role": "user", "content": "What is your refund policy?"},
    ],
    prompt_cache_key="support-bot-v3",
)
print(response.usage.input_tokens_details.cached_tokens)

Gemini follows a similar pattern, with implicit caching on by default and an optional explicit cache object when you need guaranteed retention. Its usage metadata reports a cached content token count, so the same measurement approach applies. DeepSeek needs no code changes at all, and its responses include hit and miss token counts. If you route across several providers through a gateway, normalize these fields into one schema so dashboards stay comparable. The article on AI coding agents and live API docs is a useful companion when you need to keep vendor parameters current.

Step 5 – Read the usage fields and compute hit rate

Log the usage object from every response together with the template name and version. The three Anthropic counters are cache read tokens, cache creation tokens, and uncached input tokens, and their sum is the total input. Compute the hit rate as reads divided by that total, and compute effective cost by weighting each counter with its price multiplier. The helper below does the first calculation and returns zero for empty requests. Send the result to your metrics system as a gauge tagged by endpoint so you can chart it over time. For example, 9,000 cached tokens out of 10,000 input tokens is a 90 percent hit rate.

def hit_rate(usage):
    read = usage.cache_read_input_tokens or 0
    written = usage.cache_creation_input_tokens or 0
    fresh = usage.input_tokens or 0
    total = read + written + fresh
    return read / total if total else 0.0

Expect the first request after a deploy or an expiry to show a write and no read, which is normal. A healthy steady state shows reads dominating writes, and long agent tasks can climb past 90 percent. ProjectDiscovery reported a climb from 7 percent to 84 percent after restructuring its prompts, which shows how much headroom an unoptimized system can have. If your rate stays low, print two consecutive rendered prompts and compare them character by character. The first differing character is almost always your culprit, so start the investigation there. Agents that persist state between sessions should read about AI agent memory architecture before deciding what belongs in the prefix.

Step 6 – Test prefix stability and set alerts

Turn the stability rule into an automated test so regressions are caught before they reach production. The test renders the prompt for two different users and asserts that everything above the user turn is identical. Run it in continuous integration on every change to prompt templates, tool schemas, or serialization helpers. Add a second check that fails when the prefix length drops below the provider’s minimum, since a refactor can quietly shrink a prompt below the threshold. Together these tests protect the savings you just earned from silent regressions. Run them on every commit, because one changed character can drop a 90 percent hit rate to 0.

import json


def render_prompt(user_text, tools):
    system = "You are a support agent for Acme."
    tool_json = json.dumps(tools, sort_keys=True)
    return system + "\n" + tool_json, user_text


def test_prefix_is_stable():
    tools = [{"name": "search", "description": "Search docs"}]
    prefix_a, _ = render_prompt("How do I reset my password?", tools)
    prefix_b, _ = render_prompt("What is your refund policy?", tools)
    assert prefix_a == prefix_b

Finish by wiring the hit rate gauge to an alert that fires when it drops well below its trailing average. Route the alert to the same channel as other production incidents, because a cache regression is effectively a price increase. Review the dashboard weekly for the first month and then monthly, adjusting breakpoints and lifetimes as traffic patterns change. Document the rules in your contributing guide so new engineers understand why prompts are built the way they are. Teams that follow all six steps typically see the savings described in the case studies below, and they keep them.

Key Insights From the Latest Caching Data

Taken together, the data on prompt caching explained and how to cut LLM costs tells one consistent story about where the money goes and where it can be recovered. Input tokens dominate agent and document workloads, so a discount of ninety percent on repeated input transforms the economics of long prompts. The savings are conditional, though, because they depend on identical prefixes, live entries, and providers that honor the cache. Measured results in real systems, such as the 59 percent reduction at ProjectDiscovery, sit well below the vendor maximum yet remain large. Risks appear at the edges, including expiry on sparse traffic, higher latency from careless writes, and privacy leakage through shared caches. Teams that treat hit rate as a first-class metric and design prompts for stability capture most of the upside while containing those risks.

DimensionAnthropicOpenAIGoogle GeminiDeepSeek
ActivationOpt in with a cache control setting or explicit breakpointsAutomatic for long promptsImplicit by default on 2.5 and newer models, plus explicit cache objectsAutomatic disk-based cache on every request
Read price0.1 times base input, lower on the newest models0.1 times uncached input on newer models, 50 percent at 2024 launch75 percent discount at implicit caching launchDiscounted rate for cache hit tokens
Write premium1.25 times for five minutes, 2 times for one hour1.25 times cache-write rate on the newest modelsNone for implicit caching, storage fees for explicit cachesNo premium described in the caching guide
LifetimeFive minutes refreshed on use, or one hour30 minutes on the newest models, 5 to 10 minutes or 24 hours on older onesManaged by Google for implicit, set by you for explicitCleared after hours to days when unused
Minimum prompt512 to 4,096 tokens depending on model1,024 tokens1,024 to 4,096 tokens depending on modelBest effort, short prompts may not cache
Developer controlUp to four breakpoints per requestPrompt cache key and retention settingsExplicit cache objects with your own lifetimeNone beyond prompt layout
Usage fieldsCache read and cache creation input tokensCached tokens in input token detailsCached content token count in usage metadataPrompt cache hit and miss tokens
GuaranteeNo guaranteed hit, prefix must match exactlyNo guaranteed hit, entire prefix must matchSavings passed on only when a request hitsBest effort, no guaranteed hit rate

Prompt Caching in Practice: Three Real-World Deployments

Anthropic’s Chat With a Book Benchmark

Anthropic built its launch benchmark around a workload that stresses prefill, which is asking questions about a 100,000 token book. In its prompt caching announcement, the team reported that latency fell from 11.5 seconds to 2.4 seconds, a 79 percent reduction. The same test showed a 90 percent reduction in cost, because the book was written once and then read from cache at a tenth of the normal input price. Two related benchmarks in the announcement showed smaller gains, with many-shot prompting at 10,000 tokens saving 31 percent of latency and a ten-turn conversation saving about 75 percent. The lesson is that gains scale with prefix length, so a 100,000 token prefix benefits far more than a 10,000 token one. These figures come from the vendor’s own tests rather than an independent audit, and the launch pricing still charged 25 percent more to write the cache. Real applications with shorter documents or frequent edits should expect lower savings than the headline.

Google’s Implicit Caching for Gemini 2.5

Google rolled out implicit caching for its Gemini 2.5 models in May 2025, so developers received discounts without changing any code. The launch post promised a 75 percent token discount whenever a request hit the cache. Google also lowered the minimum request size to 1,024 tokens for 2.5 Flash and 2,048 tokens for 2.5 Pro, which brought many mid-sized prompts into range. To raise hit rates it advised keeping the start of each request identical and appending the user’s question at the end. The usage metadata reports a cached content token count, so teams can verify that a discount was actually applied. Savings are not guaranteed, because a request only benefits when its prefix happens to match a stored entry. Current documentation still lists minimums between 2,048 and 4,096 tokens depending on the model generation, which limits short prompts.

OpenAI’s Automatic Caching for GPT-4o

OpenAI introduced prompt caching on October 1, 2024, and applied it automatically to GPT-4o, GPT-4o mini, and the o1 preview models. The announcement said prompts longer than 1,024 tokens qualify, with matches counted in 128-token increments. Cached input tokens were billed at a 50 percent discount, which cut the GPT-4o input price from 2.50 dollars to 1.25 dollars per million tokens. Developers implemented no code changes beyond keeping static content at the start of the prompt, so adoption took minutes rather than weeks. The trade-off was control, since the feature was automatic and developers could not choose which prefixes to keep. Later documentation added a prompt cache key and longer retention options to address exactly that limitation. Discounts have deepened over time, with current documentation describing reads at a tenth of the uncached rate on newer models.

Recommended by AIplusInfo

Books to go deeper on LLM costs and internals

Hand-picked titles that map to the mechanisms and trade-offs described above.

As an Amazon Associate, AIplusInfo earns from qualifying purchases.

AI Engineering: Building Applications with Foundation Models

Book

AI Engineering: Building Applications with Foundation Models

A practical guide to evaluating, optimizing, and shipping foundation model applications, including the cost and latency trade-offs discussed in this article.

Buy on Amazon
Hands-On Large Language Models: Language Understanding and Generation

Book

Hands-On Large Language Models: Language Understanding and Generation

Illustrated explanations of tokens, attention, and transformer internals that make it easier to see why reusing a prefix saves compute.

Buy on Amazon
Build a Large Language Model (From Scratch)

Book

Build a Large Language Model (From Scratch)

Builds a GPT-style model step by step, which makes attention and key-value state concrete for engineers who want to see the mechanism.

Buy on Amazon

Lessons From Teams That Cut LLM Costs With Caching

Case Study: ProjectDiscovery's Neo Security Agent

ProjectDiscovery faced an expensive problem with Neo, its autonomous security testing platform. Complex tasks consumed around 60 million tokens across 20 to 40 or more LLM steps, and every step resent the same system prompt. That prompt ran to more than 2,500 lines of YAML and over 20,000 tokens, so the team paid full price for identical text again and again. Early attempts at caching barely helped, because dynamic values such as working memory sat inside the prefix and changed on every call. The team needed to keep the expensive scaffold identical while still feeding the agent fresh runtime context.

The solution combined three cache breakpoints with a simple relocation trick. The team gave the static system prompt and the static tool definitions one-hour lifetimes and gave the sliding conversation window five minutes. It then moved dynamic content out of the cached prefix and appended it to the tail as user messages. According to the team's write-up, cache hit rates rose from 7 percent to 84 percent, overall costs fell by 59 percent, and 9.8 billion tokens were served from cache. One extreme task processed 67.5 million input tokens across 1,225 steps at a 91.8 percent cache rate. The design still carries constraints, because the four-breakpoint limit forces careful planning and automatic caching alone did not give the team enough control over lifetimes.

Case Study: Manus and the KV-Cache Hit Rate

Manus, a general-purpose AI agent, faced a cost structure dominated by input tokens. Its engineers reported an average input to output ratio of around 100 to 1, because every step appends an action and an observation to an ever-growing context. With Claude Sonnet, cached input cost 0.30 dollars per million tokens against 3 dollars uncached, a difference of 90 percent. Without caching, each of the many steps in a task would have paid the full price for the entire history. The challenge was to keep the history byte-identical from one step to the next.

The team developed four habits to protect the cache from accidental invalidation. It kept the prompt prefix stable and avoided precise timestamps at the start, made the context append-only, serialized data deterministically, and marked cache breakpoints explicitly when a provider needed them. In the Manus engineering post, the author calls the KV-cache hit rate the single most important metric for a production agent. The post warns that unstable JSON key ordering can silently break the cache, and it recommends session identifiers for consistent routing in self-hosted setups. These rules carry a trade-off, since an append-only history cannot be edited freely and the same post pairs the discipline with other techniques for managing growing context. The team also accepts that a single stray token early in the prompt can cost real money at scale.

Case Study: Claude Code and Cache-Safe Design

Claude Code, Anthropic's coding agent, runs long sessions that replay a large prefix on every turn. The team documented 5 distinct ways that ordinary product changes had broken the cache in practice. A timestamp inside the static system prompt invalidated everything after it, and a non-deterministic shuffle of tool order did the same. Changing tool parameters mid-conversation caused misses, and adding or removing tools mid-session broke the whole conversation cache. Switching models from Opus to Haiku forced an expensive rebuild, which was a challenge for features that wanted the cheaper model.

The team developed a set of rules in response, described in its April 2026 engineering post. Static content goes first in the prompt and dynamic content goes last. Updates travel in messages instead of edits to the system prompt, and all tools stay present so that modes such as planning are modeled as tools rather than swaps. Deferred loading keeps tool stubs stable, and compaction forks a cached call so the summarizing request shares the parent's prefix. The team credits a high hit rate with a reduction in costs and more generous rate limits, and it runs alerts that declare incidents when the rate falls too low. A remaining concern is that model switching still costs a rebuild, so some cost decisions cannot be optimized purely through prompt layout.

Frequently Asked Questions on Prompt Caching and LLM Costs

What is prompt caching in simple terms?

Prompt caching means the provider remembers the processed form of the beginning of your prompt. When the next request starts with the same text, the model skips that work and only processes what is new. This guide on prompt caching explained and how to cut LLM costs shows why that matters for long, repetitive prompts. The result is a lower bill and a faster first token.

How much money does prompt caching actually save?

Vendors advertise savings of up to 90 percent on cached input tokens, and Anthropic cites up to 85 percent lower latency. Independent measurements on agent tasks show 41 to 80 percent lower API costs. Your result depends on prefix length, hit rate, and how much output your application generates. Output tokens are not discounted, so long answers reduce the overall saving.

Does prompt caching change the model's answers?

No, the model receives exactly the same tokens whether they come from cache or from fresh computation. Caching stores intermediate attention state rather than any summary of your text. Sampling settings such as temperature still introduce normal variation between responses. You should still compare outputs in a test run to confirm your prompt layout did not change.

How long do cached prompts last?

Cache lifetime varies quite a lot by provider and by model generation. Anthropic entries last five minutes by default, refreshed on every hit, with a one-hour option at a higher write price. OpenAI documentation lists 30 minutes on its newest models and shorter or 24-hour options on older ones. DeepSeek clears unused entries automatically after a few hours to a few days.

What is the minimum prompt length for caching?

Minimum lengths depend on both the provider and the specific model you call. OpenAI documents 1,024 tokens, Anthropic ranges from 512 to 4,096 tokens depending on the model, and Gemini lists thresholds between 1,024 and 4,096. Prompts shorter than the threshold are processed normally without any error. Check your model's documentation before assuming a short prompt will be cached.

Why is my cache hit rate zero?

The most common cause is a change near the top of the prompt, such as a timestamp, request identifier, or reordered tool list. Any difference before the cache breakpoint creates a new prefix and a fresh write. Other causes include prompts below the minimum length, expired entries, or a model switch. Render two consecutive prompts and compare them character by character to find the first difference.

Is prompt caching the same as semantic caching?

No, they solve two quite different problems for application developers. Semantic caching stores full answers and returns one when a new question looks similar, which skips the model call entirely. Prompt caching reuses only the processed prefix, and the model still generates a fresh answer every time. Many teams use both, with semantic caching in front and prompt caching behind it.

Do I need to change my code to use prompt caching?

It depends on the provider and the model you are using today. OpenAI, Gemini, and DeepSeek apply caching automatically, so the main work is ordering your prompt with static content first. Anthropic asks you to add a cache control setting or explicit breakpoints. In every case you should log the usage fields so you can confirm hits.

Is prompt caching safe for private or regulated data?

It can be, but you should verify how your provider isolates caches. Researchers found that seven API providers share caches globally, which creates a timing side channel between customers. Keep personal data out of static prefixes and ask vendors about retention and isolation. Regulated teams should document these answers in their data processing review.

When is prompt caching not worth it?

Caching is usually a poor fit when prompts are short, when each prefix is used only once, or when requests arrive less often than the cache lifetime. The write premium then costs you more money than the cache ever returns. It also helps little when responses are long and the input is small, since output tokens are not discounted. Estimate your reuse rate first, and use the break-even math in this guide.

How do I know if my cache is working?

Anyone applying prompt caching explained and how to cut LLM costs should read the usage fields that the API returns with every single response. Anthropic reports cache read and cache creation tokens, OpenAI reports cached tokens, Gemini reports a cached content token count, and DeepSeek reports hit and miss tokens. Divide cached read tokens by total input tokens to get a hit rate for each endpoint. Chart that number over time and alert on any sudden drops.

Can I cache tool definitions and images?

Yes on Anthropic, where tool definitions, system messages, text, images, documents, and tool results can all be cached. Tools sit at the top of the hierarchy, so changing one invalidates everything after it. Adding or removing images can also invalidate system and message caches. Keep tool sets and image usage stable within a session.

Does prompt caching work with self-hosted models?

Yes, engines such as vLLM offer automatic prefix caching that reuses the KV state of shared prefixes. You can enable it with a single engine option at startup. Route requests from the same session to the same worker, or the cache will sit on the wrong machine. Remember that it speeds up prefill only, not the generation of new tokens.