On June 7 I published a post about what an LLM costs at campus scale. One paragraph in it warned about exactly this: "The moment someone injects a timestamp or a per-user token near the top of the prompt, the cacheable prefix shrinks to nothing and the savings evaporate." On July 20 I committed a change to my own local proxy called "Inject current time into agent context (proxy + plugin)", which put a timestamp in position two of the system prompt, above everything else in it.

The proxy is a small Python script that sits in front of LM Studio on my laptop, and everything that goes through it gets a standing contract prepended to the system message. The contract is an AGENT.md file with my instructions, a MEMORY.md file with notes, and one line giving the current time. The whole thing runs about 4,800 tokens. The clock is in there because a model has no idea what day it is, and the comment I left beside it says the proxy is stateless so the time is always current. That reasoning still holds. It says nothing about where in the prompt the line belongs.

LM Studio runs llama.cpp underneath. It reuses the KV cache only when a new prompt shares an exact prefix with the previous one, starting at the first token and ending at the first token that differs. My time string formats down to the minute, so the contract stayed byte-identical from one request to the next until the clock rolled over. Then all 4,800 tokens went through prompt processing again.

The cache covers the prompt up to the first byte that changed. Mine changed every minute, in position two.

Why the cache can only be a prefix

That rule is not a product decision anyone made at LM Studio or at Anthropic. It falls out of how a decoder-only transformer attends. Attention is causal, so each position attends only to positions before it, which means the key and value vectors cached at position j were computed from every token between 0 and j. Change the token at position i and every cached K/V after i is wrong, because each of them was built partly out of the token you replaced. Everything before i is untouched, because no earlier position ever attended to i.

So the reusable unit is a prefix and can never be a set of fragments you stitch together. If attention ran in both directions there would be no reuse at all, since a token appended at the end would change the representation of every token before it. Read that way, stable content first and volatile content last stops being a vendor rule to memorize and becomes the only arrangement the arithmetic permits. My clock sat in position two, which left one banner line above it, and that line was the entire reusable region of a 4,800-token contract.

What made me look

I did not find this by profiling. PhpStorm's CodeGPT chat became unusable through the proxy, in a way I could not pin down. I would ask a question and wait close to half a minute for the first word, then ask a follow-up and get it back immediately. Nothing was different between the two requests except when I sent them. The fast ones were the ones that landed inside the same minute as the request before.

One line, moved

The fix was to append the clock after the two files instead of before them.

parts = [
    "[[AGENT CONTEXT ... standing instructions from Anthony ...]]",
    f"===== BEGIN AGENT.md =====\n{agent}\n===== END AGENT.md =====",
]
if memory is not None:
    parts.append(f"===== BEGIN MEMORY.md =====\n{memory}\n===== END MEMORY.md =====")
# Last, not second. This line changes every minute and nothing above it does.
parts.append(f"Current time: {now}")

Measured on qwen3.8-27b, forcing a minute rollover between requests, time to first chunk went from 27.64s and 28.14s down to 4.13s and 3.76s. A request sent inside the same minute as the one before it came back in 1.28s, and it did that both before the change and after it. That last number is how I know the injection was never the expensive part. Prepending a contract that is already in the cache costs almost nothing, and reprocessing it costs 24 seconds.

Those seconds are prefill, the phase where the model works through the prompt you sent. Prefill is compute-bound and runs the whole prompt in parallel. Decode is the other phase, one token at a time, limited by memory bandwidth instead. A cache hit skips prefill and nothing else, which is why the number that moves is time to first chunk, and why tokens per second after that looks the same either way.

There is a second measurement on this machine from thirteen days earlier. Building a large preset on qwen3.6-35b-a3b, I recorded prefill at about 158 tokens per second, which is why a 53k-token prompt costs about five and a half minutes on a cold cache. At that rate 4,800 tokens comes to roughly 30 seconds. I measured 27.64 on the 27b, which implies about 174 tokens per second if the whole interval was prefill. That is division on two published figures rather than a third measurement, but it puts two models on one machine in the same place, and it gives you a way to size your own exposure: tokens above the volatile line, divided by your prefill rate, is what one invalidation costs.

Moving the line made the prefix cacheable without making every request a hit. The server runs eight parallel slots, each holding its own KV cache, so a warm prefix helps whichever slot answers next and a chat can still open cold. For the large preset I pay that prefill deliberately from the terminal with --warm before a session, so the first sentence I type is not the one that waits. Anthropic documents the same move with max_tokens: 0.

It never showed up as money

The June post says this failure mode arrives as a cost regression instead of a crash, and that you only catch it if you are watching per-request cost. On a hosted API that is now one field. Anthropic returns usage.cache_read_input_tokens on every response, and a zero there across repeated requests is the documented signal that a silent invalidator is at work. Nothing on my laptop returns that field, and nothing invoices me for tokens my own machine reprocesses. What I had instead was a chat window that was sometimes slow, which is a symptom with a hundred causes.

Anthropic's docs name the mistake outright. Their list of silent invalidators opens with datetime.now() in a system prompt, ahead of unsorted JSON and a varying tool set. My commit message says where I got the idea: Claude Code puts a time line in every turn, and I copied that into a contract that gets sent whole on every request, above 4,800 tokens that never changed.

The invalidator can be above what you are looking at

Hosted APIs work the same way, for the same reason the local one does. Anthropic's docs say that any byte change anywhere in the prefix invalidates everything after it, and OpenAI puts the rule in terms of rendering: "cache reuse requires the entire rendered prefix to match." That word, rendered, is why the order matters. Anthropic renders tools first, then the system prompt, then the messages, so a tool list that varies between requests invalidates a system prompt underneath it that nobody edited. In my proxy the invalidator sat one item above the block it invalidated. On a hosted API it can sit in a section you were not looking at.

Both vendors give the same instruction, which is stable content first and volatile content after the last breakpoint: a frozen system prompt and a deterministic tool list up top, timestamps and per-request ids at the bottom. You get at most four breakpoints in a request, and nothing shorter than about 1,024 tokens caches at all, which fails quietly rather than telling you. Writing to the cache costs about 1.25 times normal input and reading from it about a tenth, and the default entry lives five minutes unless you ask for the one-hour version, so a prefix you use once barely repays the write.

What I actually wanted

What I wanted was for the model to know the time on every request, and reordering is the version of that you reach for when the server offers nothing else. Anthropic has a designed mechanism for it. You can append a system message to the messages array instead of editing the top-level system prompt, and it arrives with operator authority in the middle of the conversation, below everything that was already cached. It is gated to a subset of models, Opus 5 and Fable 5 among them and not Sonnet 5, it needs no beta header, and it cannot be the first message in the array. llama.cpp has no equivalent, so on my proxy the clock line lives at the bottom of the block.

Read the cost post this rule came from →