Company Logo

HITWIT.AI

OpenAI's 25% cache tax, and the one setting that removes it

Ankit Verma • Oct 06, 2026 09:56 AM

GPT-5.6 and later charge 1.25× for prompt cache writes. For RAG and B2B chatbots that is a 25% input tax. One setting, explicit mode, removes it. Here's how.

<p>OpenAI's 25% cache tax, and the one setting that removes it

If you build a B2B product on OpenAI's API and moved to GPT-5.6 or later, OpenAI's prompt caching is probably charging you a 25% surcharge on input tokens for a cache you never read. Switch caching to explicit mode and the tax is gone. There is an optimal place for a breakpoint as well, but for most B2B products it saves hardly anything.

The fix: switch to explicit mode

  1. Switch to explicit mode. Set prompt_cache_options.mode to explicit. With no breakpoints, nothing is written to the cache and the surcharge is gone.
  2. Add one breakpoint, if it applies. If your system prompt and tools are identical for every user and over 1,024 tokens, put a breakpoint after them. Do not expect much. The saving is 0.9× that block's share of your prompt, which is single digits once retrieved context dominates.
  3. Check it. Open the Prompt Caching Dashboard under Usage on the OpenAI platform. A cache read hit rate near zero while you are on OpenAI's default is the tax. After the switch, input spend for the same traffic should fall by about a fifth.

The rest of this piece covers what changed, the one case OpenAI's default helps, and why a B2B product is almost never that case.

What changed in OpenAI's prompt caching pricing

Tokens written to the prompt cache now cost 1.25× the normal input rate. Tokens read from it cost 0.1×, or 0.05× on GPT-6.1 Sol. Before GPT-5.6, released in July 2026, writes carried no extra charge.

A write pays for itself the first time it is read: two uses cost 1.35× instead of 2×. A write that is never read is a 25% surcharge.

Four mechanics decide which one you get:

  • OpenAI's default is on. With no settings, OpenAI writes the whole prompt, up to the end of the latest user message. See below for how this differs from other LLM APIs.
  • Matching is by prefix. A later request reuses the cache only if it starts with exactly the same tokens, up to a point where an earlier write ended.
  • Entries expire. A cached prefix lasts at least 30 minutes from its last use. A user who returns after a longer gap rewrites the whole conversation at 1.25×, and a chat that ends after one turn never earns its write back. Sporadic, short chats lose the saving and can cost more than no caching.
  • There is a minimum. Prefixes under 1,024 tokens are not cached.

Two settings give you control. Setting prompt_cache_options.mode to explicit stops the automatic write. A prompt_cache_breakpoint on a content block marks where a write should end, and everything after the last breakpoint bills at the normal rate. OpenAI accepts breakpoints only on input content blocks (input_text, input_image and input_file), up to four per request, so an assistant reply cannot carry one.

How to spot the cache tax in the Prompt Caching Dashboard

OpenAI's usage page has a Prompt Caching Dashboard. A taxed account shows a hit rate near zero, with most of its input under cache-write tokens.

OpenAI Prompt Caching Dashboard showing a 30.6% cache read hit rate, 2.2M cache-write tokens and 22.8M cache-read tokens
The Prompt Caching Dashboard under Usage on the OpenAI platform.

How to read the dashboard:

  • Hit rate (cache-read ratio): the share of input tokens read from cache. Here it is 30.6%.
  • Cache-write against cache-read tokens: this is the tax check. Writes cost an extra 0.25× and reads save 0.9×, so caching loses money once writes exceed about 3.6× reads. Here 2.2M written against 22.8M read is a clear saving.
  • Uncached tokens: billed at the normal rate, with no surcharge. In explicit mode this is where retrieved context should land.

This account is in good shape: at standard rates its input bill is about 27% lower than with caching off. A taxed account shows the reverse, with most of its input under cache-write tokens and very little under cache-read tokens.

Only OpenAI charges for cache writes by default

Of the three major US-based providers, OpenAI is the only one that bills a cache write you did not ask for.

Provider Caching on by default Write charge
OpenAI, GPT-5.6 and later Yes 1.25× input
Claude No, you opt in 1.25× input for 5 minutes, 2× for 1 hour
Gemini, 2.5 and later Yes None stated for default caching

OpenAI's default suits conversations that only append, which is what chat and agent loops do. On any turn that writes to the cache and reads nothing back, it costs 25% more input. For a generic chatbot that is one turn. For a product that loads a whole dataset it is the first turn of every conversation. For a RAG product it is every turn.

Case 1: The generic chatbot

This is who OpenAI's default was built for. A chatbot with no retrieved context resends everything before it on each turn, so the cache gets read. If this is your product, leave the default on.

  • The 30-minute window slides. Every turn resets it, so what matters is a gap of under 30 minutes between turns.
  • Nothing is cached below 1,024 tokens. Here that covers the first two turns. Turn 3 pays the first write, and savings start at turn 4.
  • Short conversations save little. The total bill is 11% lower at 5 turns, 33% at 10 and 53% at 20.
  • Thinking shrinks the saving. A 5-turn chat costs $0.165 with caching off, or $0.54 with 1,500 thinking tokens per turn. Caching takes the same $0.018 off either, which is 11% of the first and 3% of the second.

Case 2: The naive B2B chatbot

Load the customer's whole dataset into the prompt, cache it, and let every user read it. On paper this is the best result here: 85% off the bill for a 50,000-token dataset.

  • It needs a breakpoint. On OpenAI's default, each conversation writes through its first question, so the next user's prompt does not match. Every new conversation rewrites the whole dataset at 1.25×. That is 24% more for single questions, and a 64% saving instead of 85% at 5 turns.
  • It needs traffic. After 30 idle minutes the next request rewrites the dataset. With randomly timed requests, caching loses money below roughly one request every two hours.
  • It only works for an unusually small dataset. A cached read costs a tenth of fresh input, so this only beats RAG while the dataset is under about 10× what RAG would send per turn. That is roughly 100,000 tokens, or 200,000 on GPT-6.1 Sol. At about 75,000 words, the lower figure is one long handbook. By enterprise standards that is a single document, not a knowledge base. Past 272,000 input tokens it gets worse: GPT-6 Astra bills the whole request at long-context rates, 2× for input and cache and 1.5× for output.

If this is your product, change your architecture this quarter, because your customer's dataset will grow. They always do.

Case 3: The real B2B chatbot

A real B2B product retrieves different context for every question, through RAG, GraphRAG or a search tool. The largest part of the prompt is never sent twice, so caching cannot make it cheaper. Input is over 80% of the bill, so the surcharge lands almost in full.

Change in total bill against caching off, with 10,000 tokens of fresh context per turn:

Setting 1 turn 5 turns 20 turns
OpenAI's default +20% +21% +22%
Explicit mode, no breakpoint 0% 0% 0%
Explicit mode, breakpoint on the last user message in chat history 0% −1% −20%
  • What OpenAI's default costs per turn: about 2,800 extra input tokens, which is 2.8 cents on GPT-6 Astra and 0.6 cents on GPT-6.1 Sol. Every turn.
  • With and without thinking: a 5-turn conversation costs $0.67 in explicit mode and $0.80 on OpenAI's default. With 1,500 thinking tokens per turn that is $1.04 and $1.18. It is the same 14 cents, now 13% of the bill instead of 21%.
  • The breakpoints are optional. Assistant replies cannot take a breakpoint, so the chat-history one goes on the last user message in the history; the latest reply then bills at the normal rate and is cached on the next turn. It only pays in long conversations. One after a shared system prompt returns 0.9× that prompt's share of the input: about 8% of the bill for 3,000 tokens against 30,000 of context, and nothing under 1,024 tokens.
  • Keep the order. Shared prompt, then chat history, then the fresh context and the question. Drop each turn's context once it is answered, because keeping it costs more than it saves.

The bottom line

OpenAI now charges you to write a cache, and its default fits one kind of product: a chatbot that only appends.

  • Generic chatbot: leave OpenAI's default on. It saves 11% at 5 turns and 53% at 20.
  • Whole dataset in the prompt: it works until the dataset passes about 100,000 tokens. It will.
  • Real B2B chatbot: you pay 25% more for input on every turn, and get nothing back.

One setting, explicit mode, removes the tax. After that, breakpoints are fine-tuning. They only pay in very long, uninterrupted conversations, or when your system prompt is an unusually large part of every request.

OpenAI
Prompt Caching
LLM Cost Optimization
Start YOUR Journey

Get your own AI tools
from your site!

Free To Try

Resell only after you love it