The cache starts before the model call
Prompt caching is often treated as a pricing switch. In practice, it is a serialization contract between your application and the model provider. The reusable instructions, tools and reference material must render into the same prefix before the request reaches the model.
As checked on 30 September 2026, OpenAI's prompt caching documentation says that cache reuse requires the entire rendered prefix to match. For GPT-5.6 and later, the documented minimum cacheable prefix is 1,024 tokens. A request can be logically equivalent to the previous one and still miss if the serialized tools, schema, configuration or message boundaries differ.
That changes the engineering question. Do not ask only whether the provider supports caching. Ask whether your request builder produces a stable, observable prefix that survives normal product changes.
Put stable material first
The obvious layout is also the useful one: stable developer instructions first, shared tools and reference material next, then user-specific or time-sensitive content.
The order is not cosmetic. A timestamp near the start changes the prefix for every call. A tenant identifier inserted into the developer message can prevent reuse across requests that otherwise share the same instructions. Reordering tool definitions or changing a JSON schema can move the first mismatch even when the user-facing behavior appears unchanged.
Treat prompt construction like an API payload with a deterministic serializer. Give the stable section its own function, version it, and test its rendered output. Normalize tool order instead of relying on map iteration. Keep generated request IDs, current times and per-user facts after the reusable boundary.
Conversation handling matters too. Appending a new message preserves the earlier sequence. Rewriting the last message, summarizing history or compacting context can change the earlier bytes and reduce reuse. OpenAI explicitly lists tools, structured-output format, reasoning effort and context management among the settings that can affect the cached prefix.
The right boundary is therefore not simply “the system prompt.” It is the longest prefix you can keep identical without mixing in data that should vary.
A shared prefix needs a breakpoint
A stable prefix is necessary, but it may not be sufficient. The provider also needs an eligible place to write and look up the cached state.
OpenAI documents implicit and explicit breakpoints. If the first request writes through a dynamic user message, a second request with different user content may not find a reusable boundary after the static instructions. An explicit breakpoint after the stable block makes that intent visible.
Anthropic's prompt caching reference describes a related rule. Place cache_control on the last block whose prefix is identical across the requests you want to share. Its lookup examines at most 20 block positions per breakpoint. In a growing conversation, a turn that adds many small blocks can push the previous write outside that lookback window.
These details belong in code review. A refactor that combines content blocks, splits tool output into many fragments or moves a breakpoint can change cache behavior without changing the visible answer. Add a request-shape test that records the stable prefix and breakpoint locations for representative workflows.
Separate reuse from tenant isolation
The most reusable prefix is not automatically the safest one to share.
OpenAI documents prompt_cache_key as a way to keep cache accounting separate by customer or user. It also notes that separate keys help prevent cache-hit probing across users. Keys influence routing rather than pinning a request to one machine, so they do not guarantee a hit.
Use the authenticated tenant identity to derive the cache namespace when requests contain tenant-specific material. Do not accept an arbitrary cache key from an untrusted client. Keep globally reusable public instructions ahead of the tenant-specific boundary where the provider's interface allows it, then isolate private context deliberately.
This is an architecture decision, not a prompt-writing trick. Document which prefix is public, which portion is tenant-scoped, and which data must never enter a reusable provider cache under your retention policy. Review that classification when a new tool starts injecting account data into its schema or description.
Measure the contract, not the feature flag
A cache that is enabled but rarely hit is an unmeasured cost assumption.
OpenAI exposes cached and cache-write token counts in the response usage fields. Anthropic reports cache read and creation tokens and uses a five-minute default lifetime, refreshed when cached content is used. It also offers a one-hour lifetime at additional cost. Google's Gemini context caching guide says implicit caching is enabled for Gemini 2.5 and newer models, while explicit cache objects require the generateContent API rather than the Interactions API.
The interfaces differ, but the operational dashboard should answer the same questions:
- What fraction of input tokens came from cache reads?
- How many tokens were written to the cache?
- What was the realized input cost per completed workflow?
- Did time to first output improve for the requests that hit?
- Which request-shape or model change caused the hit rate to move?
Aggregate by workflow version and tenant boundary, not only across the whole product. A healthy chat path can hide a document-analysis path that rewrites its prefix on every call. Track completed business operations as well as API calls, because retries can make a cheap individual request expensive.
Test changes with recorded request shapes
Before changing a model, SDK or request builder, collect representative rendered requests without secrets. Record the ordered message and tool structure, token counts, breakpoint positions and relevant cache settings.
Replay pairs that should share a prefix and pairs that must remain isolated. Confirm the provider reports the expected cache reads rather than inferring success from lower latency. Then test the product behavior as usual. Caching should not change output generation, but a migration can alter the request structure or model behavior at the same time.
For an agent, include the awkward pauses. Tool execution, human approval and a background job can outlast a cache lifetime. A five-minute default may suit an active conversation and fail a workflow that waits for an operator. Paying for a longer lifetime only helps when the expected reuse covers the write cost, so measure the actual pause distribution before choosing it.
What this does not change
Prompt caching does not reduce the number of output tokens, guarantee identical responses or remove the need for evaluations. It does not make an oversized context useful. A smaller uncached prompt can still cost less than padding a weakly reused prefix to a cache threshold.
It also does not replace application-level caching. If an authoritative database lookup can be performed once and safely reused, solve that at the data layer. Provider prompt caching avoids repeated model prefill work. It does not establish freshness, authorization or correctness for the facts inside the prompt.
The practical standard is narrow: build one deterministic reusable prefix, mark its boundary deliberately, isolate tenant-specific context, and prove the result with token and latency measurements. If you cannot observe those four properties, prompt caching is still an assumption in your architecture.
Written by Dandelion Labs