Changing the context without re-reading it
The CLM paper’s agent edits its context as a text file. A separate server patch then avoids re-reading most of the unchanged text after each edit. That patch reuses old work that no longer exactly matches the new text, so it is an approximation. This lab uses an exact cache instead: it skips the unchanged text before an edit and re-reads what follows.
- The file is not the speed trick. Editing a file changes what the model sees next. On its own it saves no computing.
- The speed trick is approximate. It moves old cached work into place after an edit, even though that work was computed from the text before the edit.
- We already skip re-reading the unchanged start. After an edit we re-read from just before it. In our agent runs, reading was only about 7–14% of the time, so even perfect, free reuse could not save much more.
Six words used on this page
- Prefix
- The start of a prompt, up to the first change.
- Suffix
- Everything after that change, even if those words are unchanged.
- Cache (KV cache)
- The numbers the model saved while reading each token (keys and values), kept in GPU memory so it does not read the token again.
- Prefill, or reading
- Processing the prompt before the first word of the answer. This is what a cache can skip.
- Decode, or writing
- Producing the answer one token at a time. A cache does not skip this.
- Recurrent state
- A fixed-size running summary that some layers update as they read. It has no separate entry per token.
Why does changing earlier text mean re-reading what follows?
The model saves notes for every piece of text it reads, but each note was written while looking at everything before it. If you change something early, every later note was made against text that no longer exists. The only exact fix is to go back to the change and read again from there.
The detail: attention entries and running states
A language model reads left to right. In an attention layer, each token’s saved keys and values (the KV cache) come from that layer’s input, and from the second layer up that input already mixes in every earlier token. So C’s saved entries encode “C, as read after this particular A and B”. They are not a property of C alone.
The Qwen 27B models used here and in the paper are hybrids. Only 16 of their 64 layers are attention layers with one entry per token: about 64 KiB per token in 16-bit form. The other 48 are Gated DeltaNet layers. Each keeps one fixed-size running state that folds in every token it reads. That state has no per-token entries: there is nothing to cut out when text is deleted, and nothing to move when text shifts.
An exact cache can therefore reuse only a prefix: the part before the first change. It resumes from a saved running state at or before that point and recomputes everything after it. Prefix caching, step by step.
Two separate things in the paper
First, the agent’s context is mirrored to a text file that the model edits with ordinary tools. That changes what the model sees on the next call; by itself it saves no computing. Second, a serving patch called Suffix Cache Reuse (SCR) tries to make the next call cheaper to read after such an edit.
The detail: why the two ideas are independent
The CLM harness exposes the working history as a file, applies accepted edits, and turns the file back into messages for the next call. How the editable file works. Any server can run that loop; this lab ran it on its own server without SCR.
SCR runs inside the server. It does not need the CLM harness or a session header. It matches each request to an earlier conversation by content: a common start of at least 1,024 tokens and high overall token overlap. Then it compares the new prompt with that conversation’s previous prompt. Anything that changes earlier text can trigger it, including a chat template that strips old reasoning from earlier turns.
How Suffix Cache Reuse works
After an edit, SCR keeps the cached work for text that survived, moves it to its new place, and reads only the new text plus a few tokens at the end of each moved piece. Ordinary prefix caching keeps only the start and re-reads everything after the first change.
The detail, as the released code does it
- Keep a copy. When a turn finishes, SCR copies that conversation’s attention keys and values into a separate GPU side buffer. In our reading of the code, the copy covers the prompt and the reply the model wrote. Defaults: 12 conversations × 30,000 tokens, which the code logs as 21.97 GiB (64 KiB per token) for Qwen3.6-27B.
- Find what survived. The new prompt is compared token by token with the previous one. The unchanged start (A) comes from the server’s normal prefix cache.
- Pick the pieces to move. Among the unchanged runs after the edit, it takes up to K of the longest, largest first. The launcher sets K = 6. Runs shorter than 80 tokens are not moved; they are read.
- Read the new text in front of a piece (B′) normally.
- Move the piece. Its keys and values are copied from the side buffer into fresh cache slots. Position is built into each key as a rotation (RoPE). The code rotates each copied key by the difference between its old and new position. Values are copied unchanged.
- Read the last 16 tokens of the piece instead of moving them, then the next stretch of new text, the next piece, and finally the new text at the end.
- Fall back to ordinary prefix caching if the plan does not apply or the cache slots cannot be allocated.
In our reading of the code, a conversation whose prompt plus reply exceeds the 30,000-token copy loses its side-buffer slot. Its later turns are then served by ordinary prefix caching.
Why SCR is an approximation, not magic
The moved entries for C were computed while the model was looking at the old B. SCR corrects where they sit, not what they contain. In the recurrent layers it continues from a saved state that has already read the deleted text. The paper’s own words: SCR “approximates re-prefilling”.
The detail: the attention layers and the recurrent layers
Attention layers (16 of 64). Re-rotation makes each moved key carry its new position. The content still reflects the old B: from the second layer up, C’s keys and values were built from inputs that had attended to B, not B′. The authors describe this “stale” state as possibly helpful, since it keeps information from the past. Either way, it is not what a fresh read computes.
Recurrent layers (48 of 64). There are no per-token entries to move. The released code’s default mode, fork, restores “the state saved at the end of the previous prompt” when it moves the first piece. That state has read A, the old B, and all of the old C. In our reading of the control flow, the restore replaces the state that had just read B′. B′ then reaches these layers only indirectly, through the attention layers’ outputs. The last 16 tokens of each moved piece are fed in again on top of a state that has already read them.
So after the model deletes text, 48 of its 64 layers still carry a trace of it. The paper’s README describes this step as continuing from a snapshot “before the edit”; the code comment and control flow say “end of the previous prompt”. We follow the code. A strict mode adds a length check before moving anything (the saved state may not cover more tokens than are reused) but restores the same saved state. A none mode skips the restore. Neither is the default, and the README lists fork as the supported setting.
Consequences.
- The answer after an edit is not the answer a clean read of the same text would give.
- It depends on the edit history. Each turn’s saved state is built on the previous restored state, so traces can carry forward.
- It cannot be checked against a reference token by token. Only a quality benchmark can judge it.
What the paper measured
On one web-browsing benchmark, accuracy was the same with and without SCR, and the paper’s compute count for each question fell by about 35%. That is a saving in computing, measured on one workload. It is not a measured saving in total time.
| Measure | Reported result |
|---|---|
| Setup | BrowseComp-Plus, 830 questions, Qwen3.6-27B, one GPU, SGLang 0.5.16 with the SCR patch |
| Accuracy | 60.2% with and without SCR, sampled at temperature 0.7 |
| Compute per question | 10.98 → 7.14 PFLOPs (65.0% of standard serving) |
| Prompt tokens, standard serving | 72.9% reused as prefix, 27.1% read |
| Prompt tokens, SCR | 73.9% reused as prefix, 7.8% moved, 18.3% read |
| Where the 7.8% moved came from | 5.3 points: the chat template removing old reasoning. 2.5 points: the model’s own edits. |
| Side buffer | 12 conversations × 30,000 tokens, about 22 GiB (64 KiB per token) |
What this does and does not show
- Equal accuracy is a statistical tie on one benchmark. With sampling at temperature 0.7, answers vary from run to run anyway. It does not show that SCR’s outputs match a fresh read.
- The compute count is not a time. The README’s analysis script counts two operations per model parameter for every token run through the model, including written tokens. Attention cost and waiting are not included.
- Most moved tokens did not come from self-editing. About two thirds came from the template stripping old reasoning blocks. That change can be avoided exactly by sending earlier turns in the same form on every call.
- The authors name the remaining waste themselves. Much of what is still re-read is an unchanged prefix that the server fails to reuse, because it saves recurrent states only at request boundaries. They suggest finer checkpoints, such as message boundaries, which would help exact prefix caching too.
What this lab does instead
We do not re-read everything on every change. The lab’s cache resumes from a saved state just before an edit and re-reads only what follows it. Everything it reuses is exactly what a fresh read computes, so a cached answer is identical to an uncached one.
The detail: the exact cache and its measured costs
Running states are kept every 6,656 or 13,312 tokens, depending on the setting, at the end of each prompt, and at the point where a new prompt leaves the old one. The prompt is read in 832-token pieces, the same way with or without the cache, so a resumed read does the same arithmetic as a cold one. Nothing the model wrote is cached: states made while generating are never reused.
Exactness: 99 of 99 cases gave identical output tokens with drafting on, including edits and later turns. The standard 12-prompt check returned identical answers, at 88.5–88.7 tokens/s without the cache and 89.6–89.7 with it.
| Prompt or edit | First token without reuse | First token with reuse |
|---|---|---|
| Repeated 30K prompt | 11.4 s | 0.8 s |
| Middle edit in 30K, first request | 11.3 s | 6.7 s |
| Same edit, repeated | 11.3 s | 0.7 s |
| Question over 200K (separate probes) | 114 s | 2.0 s |
A cold read is slower with the cache on because of the 832-token pieces: about 9.8 → 11.4 s at 30K and 55 → 69.6 s at 120K, 16–27% longer. Across 296 agent calls, the reuse rules predicted the reused length exactly on 226 and within one 832-token block on 292.
The two causes of waste the authors name. The template effect can be avoided exactly: send earlier turns in the same form on every call, so nothing needs moving. The lab found the same effect in its own runs (an empty-reasoning tag that appears or disappears on old turns) plus a mirror that re-rendered old turns; the fixes are small harness changes, estimated rather than yet measured. Finer recurrent-state checkpoints are exactly what the kept states above provide; in the lab’s runs, extra checkpoints at every message start would have recovered less than an append-only layout.
Exact-cache charts and checks · Reuse rules · Exactness tests · Time and reuse analysis
How much time is actually at stake?
In the lab’s agent runs, writing took most of the time and reading was the small part: about 7–14%. Even a perfect, free suffix reuse could remove only that slice, and most of it is new text that has to be read once under any scheme.
| Run (121K ledger) | Total | Writing and calls | Reading | Reading share |
|---|---|---|---|---|
| Large window | 26.5 min | 24.4 min | 1.8 min | 7% |
| Original self-editing | 64.0 min | 57.9 min | 5.6 min | 9% |
| Summaries | 41.1 min | 34.5 min | 4.2 min | 10% |
| Files, editing available | 1.9 min | 1.3 min | 0.3 min | 14% |
| Files, no context editing | 1.9 min | 1.3 min | 0.2 min | 14% |
| Revised self-editing, seed 0 | 14.9 min | 12.6 min | 1.8 min | 12% |
| Revised self-editing, seed 1 | 21.0 min | 18.2 min | 2.2 min | 10% |
| Revised self-editing, thinking only when needed (two seeds) | 6.6 and 5.7 min | Not split | Not split | Not split |
Reading is estimated from measured cold-read speeds; writing is the remainder and includes call overhead. The thinking-reduced runs were timed as a whole, not split.
What a better layout would save, and how this compares with the paper’s 35%
- Original self-editing agent: an append-only layout alone would cut its reading from about 5.6 to about 2.5 minutes. Its waste came from a retry counter and a state block rewritten near the top.
- Revised agent: a perfect layout would save 20–30 seconds per run. Even then it would still read about 239K of its 292K new tokens, because batches, its own earlier replies and the state are new text.
- The thinking-reduced revised agent finished in about 6 minutes for the whole run, without any cache change. Writing less mattered far more than reading less.
- The paper’s 35% is prefill compute on a browsing workload where large tool outputs dominate, not total time. About two thirds of its moved tokens came from a template effect that can be fixed exactly.
So why not do the same?
Because it trades exact, repeatable results for a saving that is a few percent of run time here. The same benefit is available exactly for edits near the end of the context, and most of this model cannot be spliced anyway.
- Exactness. The lab’s rule is that a faster path must give the same answer as the reference. SCR’s answers depend on the edit history and cannot be checked token by token. The measured prize is a few percent of a run.
- Edits at the end are already cheap. A well-laid-out agent shows its state last and deletes a raw item right after folding it into the state. Those edits sit at the tail, where an exact cache re-reads only a little.
- 48 of 64 layers cannot be spliced. Their running state has no per-token entries. SCR handles them by reusing a state that has read the deleted text, which is exactly the kind of mismatch the lab avoids.
What would it take?
A port of the patch, a decision about the recurrent layers, and a different way to judge results. It is possible engineering, not a setting.
- Port it. SCR patches SGLang; this lab serves with vLLM. It needs a key/value block copy with position re-rotation in the attention kernels, plus a per-conversation side buffer of 64 KiB per token: about 2 GiB for 30K tokens, about 12 GiB for 200K.
- Decide about the recurrent layers. Either restart them from the last kept state and re-read from there, which is exactly what the exact cache already does, or accept a stale state as SCR does.
- Accept answers that depend on edit history. Two runs with the same final text could answer differently.
- Add a quality evaluation. Bit-identity can no longer be the test, so every change would need a benchmark with enough questions to detect a loss.
What an exact scheme can never do: reuse cached work for text that comes after a change, because that work was computed from the text before it. Edits at the tail are cheap under every scheme. Edits at the top of a long context are expensive under every exact scheme.
Edit cost here
Wait before the first output token after one deletion, with the exact cache on. Each measured cell shows the first edited request, the same prompt sent again, and a cold read with nothing cached.
| Context | Edit at 10% | Edit at 50% | Edit at 90% | Edit in the last block |
|---|---|---|---|---|
| 30K | 11.3 s first, 0.7 s repeated (cold 11.4 s) | 6.7 s first, 0.7 s repeated (cold 11.4 s) | 6.7 s first, 0.7 s repeated (cold 11.4 s) | 0.7 s first, 0.7 s repeated (cold 11.4 s) |
| 120K | 69.6 s first, 1.5 s repeated (cold 69.8 s) | 46.6 s first, 1.5 s repeated (cold 69.8 s) | 11 s first, 1.5 s repeated (cold 69.8 s) | 1.5 s first, 1.5 s repeated (cold 69.8 s) |
| 200K | 151.7 s first, 2.2 s repeated (cold 152.1 s) | 103.1 s first, 2.2 s repeated (cold 152.1 s) | 30.7 s first, 2.2 s repeated (cold 152.1 s) | 15.8 s first, 2.2 s repeated (cold 152.1 s) |
Each edit removes four lines (about 260 tokens). An edit re-reads from the last kept state before it to the end of the prompt: one to sixteen seconds near the end, the whole cold read near the start. Kept states were 13,312 tokens apart in this run; a hit lands one block before the shared text ends, so an edit just after a kept state falls back to the one before it (the 15.8 s cell). One server, measured once.
Sources
- Context Language Models (arXiv 2609.37725)The paper: the editable context and Suffix Cache Reuse.
- The authors’ SCR codeThe SGLang patch, its defaults, tests and analysis script, at the version read for this page.
- Lab review of the paper (Markdown)Paper, code and related work, including the SCR section.
- Exact cache reuse rules (Markdown)Where states are kept and when a prompt can resume.
- Exactness tests (Markdown)The 99-case comparison and the standard decode check.
- Where the time went (Markdown)Reading and writing time per run, and what a better layout would save.