All posts Services Contact
Client login Get started

Prompt caching is most of your bill, and most teams lose it by accident

AI Agents Cost

If you run an agent with a long system prompt over a long conversation, the difference between cached and uncached tokens is roughly an order of magnitude. That is not an optimisation detail. It is usually the single biggest line item on the invoice.

The model is not the fragile part. Cache invalidation is.

What actually invalidates a prefix

A cached prefix is a byte-for-byte match on the start of the request. Turn one and turn two of a conversation share a prefix. Turn two and turn three do too, as long as nothing before the new message changed.

So the rule is simple and unforgiving: nothing in the history, the toolset, or the system prompt may change mid-conversation. Most breakage comes from four places.

Re-sorting or re-dating context

If your memory block is rebuilt each turn and the ordering depends on a timestamp or a relevance score that moves, the prefix changes every turn and you cache nothing. Sort once and keep the order stable for the life of the conversation.

Swapping the toolset

Adding or removing one tool changes the tools array, which sits before the messages. The whole cache is gone. This is why tool changes take effect on the next session, not immediately.

Rebuilding the system prompt

A system prompt assembled from the current time, a version string, or a random seed is not byte-stable. Remove anything that would differ between two turns of the same conversation.

Injecting a synthetic user message

Appending a mid-loop instruction as a user message changes the tail of the conversation and can break strict role alternation. If the agent genuinely needs to be told something mid-task, that belongs in the tool result or in a fresh turn.

Defer changes, offer an opt-in for immediacy

The fix for "I just installed a skill and nothing happened" is not to rebuild the prompt mid-session. It is to queue the change and apply it on the next session, with an explicit --now flag for the rare person who accepts the cache hit they are about to lose. Users understand that trade-off. They do not understand why the agent forgot something.

Measure it, do not assume

Log cache read tokens and cache write tokens per request. If your cache hit rate on a long conversation drops, something is mutating the prefix, and the log tells you which turn it started on. This is a five-minute instrumentation job that pays for itself the first time it catches you.