SkinDeepRESEARCHSteve Seguin

Changing the context without re-reading it

The CLM paper’s agent edits its context as a text file. A separate server patch then avoids re-reading most of the unchanged text after each edit. That patch reuses old work that no longer exactly matches the new text, so it is an approximation. This lab uses an exact cache instead: it skips the unchanged text before an edit and re-reads what follows.

Six words used on this page
Prefix
The start of a prompt, up to the first change.
Suffix
Everything after that change, even if those words are unchanged.
Cache (KV cache)
The numbers the model saved while reading each token (keys and values), kept in GPU memory so it does not read the token again.
Prefill, or reading
Processing the prompt before the first word of the answer. This is what a cache can skip.
Decode, or writing
Producing the answer one token at a time. A cache does not skip this.
Recurrent state
A fixed-size running summary that some layers update as they read. It has no separate entry per token.

Why does changing earlier text mean re-reading what follows?

The model saves notes for every piece of text it reads, but each note was written while looking at everything before it. If you change something early, every later note was made against text that no longer exists. The only exact fix is to go back to the change and read again from there.

What each part of the cache depends onThree blocks of context, A, B and C. Below them, attention-cache entries: A’s were made from A, B’s from A and B, C’s from A, B and C. Below that, one running state updated after A, after B and after C. If B is edited, A’s entries stay valid, but B’s and C’s entries and every running state after A were computed from the old B. Context: the text the model reads, left to right A B, then edited C Attention cache: one entry per token (16 of 64 layers) made from A made from A + B made from A + B + C Running state: one per layer, updated as it reads (48 of 64) after Aafter A + Bafter A + B + C still valid after B changes computed from the old B
Change B and only A’s part of the cache still matches. The same is true for any model that reads left to right.
The detail: attention entries and running states

A language model reads left to right. In an attention layer, each token’s saved keys and values (the KV cache) come from that layer’s input, and from the second layer up that input already mixes in every earlier token. So C’s saved entries encode “C, as read after this particular A and B”. They are not a property of C alone.

The Qwen 27B models used here and in the paper are hybrids. Only 16 of their 64 layers are attention layers with one entry per token: about 64 KiB per token in 16-bit form. The other 48 are Gated DeltaNet layers. Each keeps one fixed-size running state that folds in every token it reads. That state has no per-token entries: there is nothing to cut out when text is deleted, and nothing to move when text shifts.

An exact cache can therefore reuse only a prefix: the part before the first change. It resumes from a saved running state at or before that point and recomputes everything after it. Prefix caching, step by step.

Two separate things in the paper

First, the agent’s context is mirrored to a text file that the model edits with ordinary tools. That changes what the model sees on the next call; by itself it saves no computing. Second, a serving patch called Suffix Cache Reuse (SCR) tries to make the next call cheaper to read after such an edit.

The file-edit loop and where SCR fitsStep 1: the model reads the prompt and writes a tool call to edit the context file. Step 2: the harness applies the edit to the context file, plain text on disk. Step 3: the harness rebuilds the next prompt from the file. Step 4: the server reads the new prompt, then the model writes. Steps 2 and 3 decide what the model sees. Step 4 is where reading costs time, and the only place SCR acts. 1. Model reads the promptand writes a tool call:“edit the context file” 2. Harness applies theedit to the context file(plain text on disk) 3. Harness rebuilds thenext prompt from the file 4. Server reads the newprompt (prefill), thenthe model writes Steps 2–3: what the model sees (the CLM file) Step 4: reading cost. SCR works only here.
The server never sees “an edit”. It receives a new prompt; SCR compares it with the last one.
The detail: why the two ideas are independent

The CLM harness exposes the working history as a file, applies accepted edits, and turns the file back into messages for the next call. How the editable file works. Any server can run that loop; this lab ran it on its own server without SCR.

SCR runs inside the server. It does not need the CLM harness or a session header. It matches each request to an earlier conversation by content: a common start of at least 1,024 tokens and high overall token overlap. Then it compares the new prompt with that conversation’s previous prompt. Anything that changes earlier text can trigger it, including a chat template that strips old reasoning from earlier turns.

How Suffix Cache Reuse works

After an edit, SCR keeps the cached work for text that survived, moves it to its new place, and reads only the new text plus a few tokens at the end of each moved piece. Ordinary prefix caching keeps only the start and re-reads everything after the first change.

Three ways to serve the prompt A, B′, C after B was editedRe-read everything: A, B′ and C are all read. Exact prefix reuse, used in this lab: A is reused; B′ and C are read. Suffix cache reuse, from the paper: A is reused, B′ is read, C’s old cache entries are moved to their new positions, and only C’s last 16 tokens are read. The moved entries were computed against the old B, so they are reused approximately. New prompt after the edit: A, then B′ (shorter than B), then C A B′ C (unchanged words) Re-read everything read read read Exact prefix reuse (this lab) reused read read again Suffix cache reuse (the paper’s SCR) reused read moved, shifted last 16 tokens read reused exactly read now reused approximately: computed against the old B
Only the attention layers have entries to move. The recurrent layers are handled differently; see the next section.
The detail, as the released code does it
  1. Keep a copy. When a turn finishes, SCR copies that conversation’s attention keys and values into a separate GPU side buffer. In our reading of the code, the copy covers the prompt and the reply the model wrote. Defaults: 12 conversations × 30,000 tokens, which the code logs as 21.97 GiB (64 KiB per token) for Qwen3.6-27B.
  2. Find what survived. The new prompt is compared token by token with the previous one. The unchanged start (A) comes from the server’s normal prefix cache.
  3. Pick the pieces to move. Among the unchanged runs after the edit, it takes up to K of the longest, largest first. The launcher sets K = 6. Runs shorter than 80 tokens are not moved; they are read.
  4. Read the new text in front of a piece (B′) normally.
  5. Move the piece. Its keys and values are copied from the side buffer into fresh cache slots. Position is built into each key as a rotation (RoPE). The code rotates each copied key by the difference between its old and new position. Values are copied unchanged.
  6. Read the last 16 tokens of the piece instead of moving them, then the next stretch of new text, the next piece, and finally the new text at the end.
  7. Fall back to ordinary prefix caching if the plan does not apply or the cache slots cannot be allocated.

In our reading of the code, a conversation whose prompt plus reply exceeds the 30,000-token copy loses its side-buffer slot. Its later turns are then served by ordinary prefix caching.

SCR README · Default settings · Key re-rotation

Why SCR is an approximation, not magic

The moved entries for C were computed while the model was looking at the old B. SCR corrects where they sit, not what they contain. In the recurrent layers it continues from a saved state that has already read the deleted text. The paper’s own words: SCR “approximates re-prefilling”.

The detail: the attention layers and the recurrent layers

Attention layers (16 of 64). Re-rotation makes each moved key carry its new position. The content still reflects the old B: from the second layer up, C’s keys and values were built from inputs that had attended to B, not B′. The authors describe this “stale” state as possibly helpful, since it keeps information from the past. Either way, it is not what a fresh read computes.

Recurrent layers (48 of 64). There are no per-token entries to move. The released code’s default mode, fork, restores “the state saved at the end of the previous prompt” when it moves the first piece. That state has read A, the old B, and all of the old C. In our reading of the control flow, the restore replaces the state that had just read B′. B′ then reaches these layers only indirectly, through the attention layers’ outputs. The last 16 tokens of each moved piece are fed in again on top of a state that has already read them.

So after the model deletes text, 48 of its 64 layers still carry a trace of it. The paper’s README describes this step as continuing from a snapshot “before the edit”; the code comment and control flow say “end of the previous prompt”. We follow the code. A strict mode adds a length check before moving anything (the saved state may not cover more tokens than are reused) but restores the same saved state. A none mode skips the restore. Neither is the default, and the README lists fork as the supported setting.

Consequences.

  • The answer after an edit is not the answer a clean read of the same text would give.
  • It depends on the edit history. Each turn’s saved state is built on the previous restored state, so traces can carry forward.
  • It cannot be checked against a reference token by token. Only a quality benchmark can judge it.

Recurrent-state modes in the code · Lab review of SCR

What the paper measured

On one web-browsing benchmark, accuracy was the same with and without SCR, and the paper’s compute count for each question fell by about 35%. That is a saving in computing, measured on one workload. It is not a measured saving in total time.

MeasureReported result
SetupBrowseComp-Plus, 830 questions, Qwen3.6-27B, one GPU, SGLang 0.5.16 with the SCR patch
Accuracy60.2% with and without SCR, sampled at temperature 0.7
Compute per question10.98 → 7.14 PFLOPs (65.0% of standard serving)
Prompt tokens, standard serving72.9% reused as prefix, 27.1% read
Prompt tokens, SCR73.9% reused as prefix, 7.8% moved, 18.3% read
Where the 7.8% moved came from5.3 points: the chat template removing old reasoning. 2.5 points: the model’s own edits.
Side buffer12 conversations × 30,000 tokens, about 22 GiB (64 KiB per token)
What this does and does not show
  • Equal accuracy is a statistical tie on one benchmark. With sampling at temperature 0.7, answers vary from run to run anyway. It does not show that SCR’s outputs match a fresh read.
  • The compute count is not a time. The README’s analysis script counts two operations per model parameter for every token run through the model, including written tokens. Attention cost and waiting are not included.
  • Most moved tokens did not come from self-editing. About two thirds came from the template stripping old reasoning blocks. That change can be avoided exactly by sending earlier turns in the same form on every call.
  • The authors name the remaining waste themselves. Much of what is still re-read is an unchanged prefix that the server fails to reuse, because it saves recurrent states only at request boundaries. They suggest finer checkpoints, such as message boundaries, which would help exact prefix caching too.

Results and figures in the README · The paper

What this lab does instead

We do not re-read everything on every change. The lab’s cache resumes from a saved state just before an edit and re-reads only what follows it. Everything it reuses is exactly what a fresh read computes, so a cached answer is identical to an uncached one.

Where kept states sit along a 30K-token prompt, and where an edit resumesKept states sit every 13,312 tokens and at the end of each prompt. After a 300-token deletion near the middle, the first request reuses everything up to the kept state at 13,312 and re-reads the rest: 6.7 seconds instead of 11.3 cold. Sent again, the edited prompt reuses almost everything: 0.7 seconds. An edit in the last block would resume from 26,624 and re-read only a few thousand tokens; that case is being measured. First request after a middle deletion reused re-read: 6.7 s (cold 11.3 s) edit The same edited prompt, sent again reused up to the state kept at its end: 0.7 s An edit in the last block reused 013,31226,62430K ◆ kept state (shown every 13,312 tokens) and at each prompt’s end
The cost of an exact edit is the text after the nearest kept state before it. Edits near the end are cheap.
The detail: the exact cache and its measured costs

Running states are kept every 6,656 or 13,312 tokens, depending on the setting, at the end of each prompt, and at the point where a new prompt leaves the old one. The prompt is read in 832-token pieces, the same way with or without the cache, so a resumed read does the same arithmetic as a cold one. Nothing the model wrote is cached: states made while generating are never reused.

Exactness: 99 of 99 cases gave identical output tokens with drafting on, including edits and later turns. The standard 12-prompt check returned identical answers, at 88.5–88.7 tokens/s without the cache and 89.6–89.7 with it.

Prompt or editFirst token without reuseFirst token with reuse
Repeated 30K prompt11.4 s0.8 s
Middle edit in 30K, first request11.3 s6.7 s
Same edit, repeated11.3 s0.7 s
Question over 200K (separate probes)114 s2.0 s

A cold read is slower with the cache on because of the 832-token pieces: about 9.8 → 11.4 s at 30K and 55 → 69.6 s at 120K, 16–27% longer. Across 296 agent calls, the reuse rules predicted the reused length exactly on 226 and within one 832-token block on 292.

The two causes of waste the authors name. The template effect can be avoided exactly: send earlier turns in the same form on every call, so nothing needs moving. The lab found the same effect in its own runs (an empty-reasoning tag that appears or disappears on old turns) plus a mirror that re-rendered old turns; the fixes are small harness changes, estimated rather than yet measured. Finer recurrent-state checkpoints are exactly what the kept states above provide; in the lab’s runs, extra checkpoints at every message start would have recovered less than an append-only layout.

Exact-cache charts and checks · Reuse rules · Exactness tests · Time and reuse analysis

How much time is actually at stake?

In the lab’s agent runs, writing took most of the time and reading was the small part: about 7–14%. Even a perfect, free suffix reuse could remove only that slice, and most of it is new text that has to be read once under any scheme.

Run (121K ledger)TotalWriting and callsReadingReading share
Large window26.5 min24.4 min1.8 min7%
Original self-editing64.0 min57.9 min5.6 min9%
Summaries41.1 min34.5 min4.2 min10%
Files, editing available1.9 min1.3 min0.3 min14%
Files, no context editing1.9 min1.3 min0.2 min14%
Revised self-editing, seed 014.9 min12.6 min1.8 min12%
Revised self-editing, seed 121.0 min18.2 min2.2 min10%
Revised self-editing, thinking only when needed (two seeds)6.6 and 5.7 minNot splitNot splitNot split

Reading is estimated from measured cold-read speeds; writing is the remainder and includes call overhead. The thinking-reduced runs were timed as a whole, not split.

What a better layout would save, and how this compares with the paper’s 35%
  • Original self-editing agent: an append-only layout alone would cut its reading from about 5.6 to about 2.5 minutes. Its waste came from a retry counter and a state block rewritten near the top.
  • Revised agent: a perfect layout would save 20–30 seconds per run. Even then it would still read about 239K of its 292K new tokens, because batches, its own earlier replies and the state are new text.
  • The thinking-reduced revised agent finished in about 6 minutes for the whole run, without any cache change. Writing less mattered far more than reading less.
  • The paper’s 35% is prefill compute on a browsing workload where large tool outputs dominate, not total time. About two thirds of its moved tokens came from a template effect that can be fixed exactly.

Time and reuse analysis · All agent results

So why not do the same?

Because it trades exact, repeatable results for a saving that is a few percent of run time here. The same benefit is available exactly for edits near the end of the context, and most of this model cannot be spliced anyway.

  1. Exactness. The lab’s rule is that a faster path must give the same answer as the reference. SCR’s answers depend on the edit history and cannot be checked token by token. The measured prize is a few percent of a run.
  2. Edits at the end are already cheap. A well-laid-out agent shows its state last and deletes a raw item right after folding it into the state. Those edits sit at the tail, where an exact cache re-reads only a little.
  3. 48 of 64 layers cannot be spliced. Their running state has no per-token entries. SCR handles them by reusing a state that has read the deleted text, which is exactly the kind of mismatch the lab avoids.

What would it take?

A port of the patch, a decision about the recurrent layers, and a different way to judge results. It is possible engineering, not a setting.

  • Port it. SCR patches SGLang; this lab serves with vLLM. It needs a key/value block copy with position re-rotation in the attention kernels, plus a per-conversation side buffer of 64 KiB per token: about 2 GiB for 30K tokens, about 12 GiB for 200K.
  • Decide about the recurrent layers. Either restart them from the last kept state and re-read from there, which is exactly what the exact cache already does, or accept a stale state as SCR does.
  • Accept answers that depend on edit history. Two runs with the same final text could answer differently.
  • Add a quality evaluation. Bit-identity can no longer be the test, so every change would need a benchmark with enough questions to detect a loss.

What an exact scheme can never do: reuse cached work for text that comes after a change, because that work was computed from the text before it. Edits at the tail are cheap under every scheme. Edits at the top of a long context are expensive under every exact scheme.

Edit cost here

Wait before the first output token after one deletion, with the exact cache on. Each measured cell shows the first edited request, the same prompt sent again, and a cold read with nothing cached.

ContextEdit at 10%Edit at 50%Edit at 90%Edit in the last block
30K11.3 s first, 0.7 s repeated (cold 11.4 s)6.7 s first, 0.7 s repeated (cold 11.4 s)6.7 s first, 0.7 s repeated (cold 11.4 s)0.7 s first, 0.7 s repeated (cold 11.4 s)
120K69.6 s first, 1.5 s repeated (cold 69.8 s)46.6 s first, 1.5 s repeated (cold 69.8 s)11 s first, 1.5 s repeated (cold 69.8 s)1.5 s first, 1.5 s repeated (cold 69.8 s)
200K151.7 s first, 2.2 s repeated (cold 152.1 s)103.1 s first, 2.2 s repeated (cold 152.1 s)30.7 s first, 2.2 s repeated (cold 152.1 s)15.8 s first, 2.2 s repeated (cold 152.1 s)

Each edit removes four lines (about 260 tokens). An edit re-reads from the last kept state before it to the end of the prompt: one to sixteen seconds near the end, the whole cold read near the start. Kept states were 13,312 tokens apart in this run; a hit lands one block before the shared text ends, so an edit just after a kept state falls back to the one before it (the 15.8 s cell). One server, measured once.

Sources