| Provider | Caching on by default | Write charge |
|---|---|---|
| OpenAI, GPT-5.6 and later | Yes | 1.25× input |
| Claude | No, you opt in | 1.25× input for 5 minutes, 2× for 1 hour |
| Gemini, 2.5 and later | Yes | None stated for default caching |
OpenAI's 25% cache tax, and the one setting that removes it
Ankit Verma • Oct 06, 2026 09:56 AM
GPT-5.6 and later charge 1.25× for prompt cache writes. For RAG and B2B chatbots that is a 25% input tax. One setting, explicit mode, removes it. Here's how.

If you build a B2B product on OpenAI's API and moved to GPT-5.6 or later, OpenAI's prompt caching is probably charging you a 25% surcharge on input tokens for a cache you never read. Switch caching to explicit mode and the tax is gone. There is an optimal place for a breakpoint as well, but for most B2B products it saves hardly anything.
prompt_cache_options.mode to explicit. With no breakpoints, nothing is written to the cache and the surcharge is gone.The rest of this piece covers what changed, the one case OpenAI's default helps, and why a B2B product is almost never that case.
Tokens written to the prompt cache now cost 1.25× the normal input rate. Tokens read from it cost 0.1×, or 0.05× on GPT-6.1 Sol. Before GPT-5.6, released in July 2026, writes carried no extra charge.
A write pays for itself the first time it is read: two uses cost 1.35× instead of 2×. A write that is never read is a 25% surcharge.
Four mechanics decide which one you get:
Two settings give you control. Setting prompt_cache_options.mode to explicit stops the automatic write. A prompt_cache_breakpoint on a content block marks where a write should end, and everything after the last breakpoint bills at the normal rate. OpenAI accepts breakpoints only on input content blocks (input_text, input_image and input_file), up to four per request, so an assistant reply cannot carry one.
OpenAI's usage page has a Prompt Caching Dashboard. A taxed account shows a hit rate near zero, with most of its input under cache-write tokens.
How to read the dashboard:
This account is in good shape: at standard rates its input bill is about 27% lower than with caching off. A taxed account shows the reverse, with most of its input under cache-write tokens and very little under cache-read tokens.
Of the three major US-based providers, OpenAI is the only one that bills a cache write you did not ask for.
| Provider | Caching on by default | Write charge |
|---|---|---|
| OpenAI, GPT-5.6 and later | Yes | 1.25× input |
| Claude | No, you opt in | 1.25× input for 5 minutes, 2× for 1 hour |
| Gemini, 2.5 and later | Yes | None stated for default caching |
OpenAI's default suits conversations that only append, which is what chat and agent loops do. On any turn that writes to the cache and reads nothing back, it costs 25% more input. For a generic chatbot that is one turn. For a product that loads a whole dataset it is the first turn of every conversation. For a RAG product it is every turn.
This is who OpenAI's default was built for. A chatbot with no retrieved context resends everything before it on each turn, so the cache gets read. If this is your product, leave the default on.
Load the customer's whole dataset into the prompt, cache it, and let every user read it. On paper this is the best result here: 85% off the bill for a 50,000-token dataset.
If this is your product, change your architecture this quarter, because your customer's dataset will grow. They always do.
A real B2B product retrieves different context for every question, through RAG, GraphRAG or a search tool. The largest part of the prompt is never sent twice, so caching cannot make it cheaper. Input is over 80% of the bill, so the surcharge lands almost in full.
Change in total bill against caching off, with 10,000 tokens of fresh context per turn:
| Setting | 1 turn | 5 turns | 20 turns |
|---|---|---|---|
| OpenAI's default | +20% | +21% | +22% |
| Explicit mode, no breakpoint | 0% | 0% | 0% |
| Explicit mode, breakpoint on the last user message in chat history | 0% | −1% | −20% |
OpenAI now charges you to write a cache, and its default fits one kind of product: a chatbot that only appends.
One setting, explicit mode, removes the tax. After that, breakpoints are fine-tuning. They only pay in very long, uninterrupted conversations, or when your system prompt is an unusually large part of every request.
Free To Try
Resell only after you love it

HITWIT.AI
We'll use this to recommend next steps. You can change it anytime
We'll use this to recommend next steps. You can change it anytime