All articles
prompt cachingLLM cost optimizationAI agents

69% of Our LLM Bill Was Cache Writes

We run hundreds of long-lived AI agents for our users. When I audited one day of production traffic, 69% of the LLM cost was cache writes. Here is what we found and the fixes that worked.

Xiaomeng Li

Xiaomeng Li · Creator, NanoRhino


69% of Our LLM Bill Was Cache Writes

Prompt cost engineering for a fleet of long-lived AI agents

NanoRhino is an AI nutrition coach. It works entirely over text messages. Every user has their own long-lived agent, with its own memory, its own conversation history, and its own scheduled reminders. There are hundreds of these agents. Every inbound text, and every scheduled job, is one agent turn against an LLM API.

Prompt caching should make this cheap. You pay once to write the big static prompt into the cache, and later turns read it at 10% of the price. Last week I audited one full day of our production traffic, call by call. The result: 69% of our LLM spend was cache writes. Not output tokens. Not cache reads. We paid the 1.25× cache-write surcharge again and again, and almost never got the 0.1× read discount.

This post is about what we found, and which fixes actually moved the number. If you run many agents with low traffic per agent, instead of one chatbot with heavy traffic, your bill probably looks like ours did.

Why a fleet breaks prompt caching

Every prompt caching demo has a hidden assumption: the cache is hot. One user talking to one chatbot keeps the prefix warm, so caching works like the pricing page says.

A fleet is the opposite. Our median user texts a few times a day. When the next message arrives, the cache entry is already expired, so the turn writes the whole prefix again at 1.25×. Nothing is misconfigured. It is just arithmetic.

The main lesson for us: cache hit rate is a property of your traffic, not of your architecture. You cannot fix it by adding more cache breakpoints. You fix it by making the thing that gets re-written smaller, because the cache will be cold anyway.

One early decision that saved us: split the prompt in two

Our system prompt is assembled in two halves:

  • A fleet-shared half: engine rules, the coaching playbook, skill routing, tool schemas. Same bytes for every user. So when any user's turn writes it to the cache, every other user's turn soon after reads it cheap.
  • A per-user half: the user's profile, memory layers, current coaching focus. This part is cold per user, and we accept that.

The rule is strict: one per-user byte in the shared half breaks cross-user caching for the whole fleet. It is easy to break by accident. A timestamp, a user-specific flag, a harmless "today is Tuesday". Nothing visibly fails when you do it. Your bill just goes up. We treat the shared half as frozen bytes and check it in code review.

Log every call before you optimize

The audit was only possible because we log every LLM call in full: request, response, token counts, cache state. The log is appended locally on the reply path, so it adds zero latency, and it ships to S3 nightly.

One day of real traffic, 109 runs, was enough to rank everything by cost. Two of the four fixes below were things I would never have guessed without the log. So if you take one thing from this post: log every call first.

Fix 1: tool schemas are part of the prompt

Tool definitions feel like configuration, but they are serialized into the prompt like everything else. Our full chat toolset is about 12.3K tokens of JSON schema.

The logs showed 34 scheduled reminder runs made 3 tool calls in total. And scheduled turns are the always-cold turns, because they fire for users who have not texted for hours. So on a reminder turn, about 70% of the prefix we kept re-writing was schema for tools the turn never used.

Now each turn type gets a minimal toolset. Reminder turns get file access plus the reminder tools. Memory consolidation turns get the ledger tools. Plan turns get none. Interactive chat keeps the full set. This was the cheapest fix on the list. It is a lookup table.

Fix 2: slim the replayed history, with stable bytes

Each turn replays recent conversation history. Ours was full of dead weight: old tool results (a full day-summary blob from four turns ago), and stale context blocks inside past user messages (the current turn re-sends fresh data anyway).

The tricky part: history is also cached, with a breakpoint at its tail. If your slimming logic produces different bytes depending on anything — current date, message position, config — the bytes shift and the cache misses. So the transform must be a pure function per message: the same message always slims to the same bytes, forever. Ours cuts tool results over 800 characters down to the first 400 plus an omission note, and collapses stale context blocks into a one-line stub. Measured on one day of chat turns: 25.2K → 8.7K characters of replayed history per turn, −65%.

Fix 3: compress documents in place, do not move them out

The two biggest documents in the shared prefix were the skill-routing manual (36K tokens) and the coaching playbook (15.9K). We compressed them to 17.5K and 7.9K by deleting incident stories, repeated explanations, and rules that another layer already enforces word for word. Every behavioral rule stayed.

The tempting alternative was to move detail into separate files that the model reads on demand. The logs said no: in two days of production, the model followed those "see file X" pointers four times. Content moved out of the prompt is not loaded lazily. It is dropped silently. Lazy loading only works for content the model knows it is missing, and behavioral rules do not announce their absence.

Fix 4: old thinking blocks ride along, and you pay for them

On current Claude models, extended-thinking blocks from previous turns that you replay in history are billed as input tokens. We confirmed this with the token-counting endpoint: in one replayed history, 12 reasoning blocks cost 2,069 input tokens. Worse, preserved thinking is bound to its exact prefix, so history that contains it cannot be slimmed without invalidating it. It was blocking Fix 2.

A completed turn does not need its reasoning anymore. We now strip reasoning blocks from finished turns at replay time, and keep the text and tool calls. If your agent framework replays history verbatim, check this one. It compounds on every turn.

Two problems the fixes caused

Shipping the fixes exposed two regressions. Honestly, they taught me more than the fixes did.

1. The prompt is an instruction to spend money. Our reminder prompt said: check whether the user is on a break before nudging. Sounds careful. In practice, every reminder turn burned a tool call reading a file, for a condition the engine already guarantees before the turn even starts. Average steps per reminder went from 1.19 to 1.80, and warm reminder cost went up 49%. Every "always check X first" line in a prompt is a recurring cost on every future turn. Audit those lines against what your platform already enforces.

2. Take a tool away, and the model finds another door. After we removed the meal-logging tool from reminder turns, one scheduled turn decided a user's coffee should be logged, and wrote the meal file by hand with the generic file-writing tool. That bypassed every validation in the dedicated writer. The fix was not prompt wording. The generic writers now refuse the data files that have dedicated writers, in every turn type. What you remove from the prompt, you must also enforce in the tool layer.

Results

Compared like for like after deploy: cold-cache chat turns −16%, cold reminder turns −22%, replayed history −65% by volume. Behavior unchanged on our regression suites.

One warning about measuring: compare by cache state. A day with luckier cache timing can look like a 30% win. A real improvement can look like a regression if traffic went quiet. We bucket cold and warm turns and compare inside each bucket only.

The checklist

If you run a fleet of long-lived agents:

  1. Log every call with cache state. Do not optimize before you have ranked one real day of traffic.
  2. Assume cold. Size the prompt for the cache-miss case. Your users' texting habits set the hit rate, not you.
  3. Split shared from per-user, and protect the shared half's byte stability like an API contract.
  4. Count tool schemas as prompt. Give each turn type the minimum toolset, then guard the removed abilities in code.
  5. Slim history with pure, byte-stable transforms, and strip old reasoning from completed turns.
  6. Compress in place. Do not trust lazy loading for rules the model does not know it is missing.
  7. Re-read your prompts as recurring costs. Every standing instruction is spend on every future turn.

None of this is advanced. It is normal capacity engineering, except the resource is rented by the token and the meter runs on every turn. We only get paid when our users lose weight, so every one of these tokens is spent against revenue that is not guaranteed. That keeps us careful.


Written by Xiaomeng Li, Co-founder & CTO of NanoRhino. NanoRhino is an AI nutrition coach over SMS/WhatsApp, priced at $10 per pound actually lost. The deficit-loop coaching engine behind it is U.S. patent-pending.

Built with care by Link Heart Limited in Houston, Texas.