SkinDeepRESEARCHSteve Seguin

Read once, reuse the unchanged start

Prefix caching reduces the wait before an answer; long-context generation still has a cost.

Prefix caching makes the next answer start sooner. It saves the model’s already-computed representation of the unchanged start of a prompt. See how prefix caching works, step by step.

Waiting and writing are different

Repeated 30K prompt: cache off 11.4 seconds to first token, cache on 0.8 seconds. Separate standard 12-prompt check, second pass: 88.7 output tokens per second with cache off and 89.7 with cache on.
Caching shortened the wait. The roughly 1% writing-speed difference is within normal variation.

Once generation starts, full-attention layers still use the retained context for each new token. A cached long document remains a long document.

All 12 answers matched the reference on both passes of the standard check. The cache was held in GPU memory.

All waiting and writing measurements
Prompt or editFirst token without reuseFirst token with reuse
Repeated 30K prompt11.4 s0.8 s
Middle edit in 30K, first request11.3 s6.7 s
Same edit, repeated11.3 s0.7 s
Question over 200K (separate probes)114 s2.0 s

The 30K cases compare fixed 832-token reading chunks with reuse off and on. The 200K example combines separate long-context runs with different reading configurations.

ServerFirst passRepeat passExact per pass
Cache off88.5 tokens/s88.7 tokens/s12/12
Cache on89.6 tokens/s89.7 tokens/s12/12

One server per configuration; the small decode-rate difference is not evidence of a speedup.

Did the answers change?

The exact-cache test returned the same output tokens in 99/99 cases, including repeat questions, edits and later turns. It covered prompts up to 30K; separate cold-repeat checks also matched at 120K.

How the cache stayed exact

This model combines ordinary attention with layers that carry a running state. The lab’s cache saves only states created while reading a prompt, and uses consistent 832-token reading chunks. It avoids mixing those states with ones created through a different computation path while generating an answer.

With drafting enabled, 99/99 cases returned the same token sequence as the matched uncached reference, including second and third conversation turns. The test covered eleven prompts up to 30K and up to 48 output tokens per case. Separate 120K cold repeats matched too. Matching output tokens is narrower than a scores-level or every-input guarantee.

What does it cost?

Smaller reading chunks made the initial read slower: a 30K prompt took about 11.4 seconds instead of 9.8; a 120K cold check took 69.6 instead of about 55. Saved checkpoints also consume GPU memory. Repeated questions recover that cost by skipping much of the reading.

An edit keeps only the matching prefix reusable. The 30K middle-edit test started in 6.7 seconds on its first edited request and 0.7 seconds on a repeat. Putting frequently updated state near the end reduces how much text must be recomputed. Eviction, a changed prompt or a restarted server can remove reuse.

When did reuse pay off in the long task?

A reconstruction of the large-window ledger estimates that caching repaid its cold-read cost by call 4. Estimated reading time was 110 seconds with reuse, versus 866 without it. The cache-off total is a calculation from saved calls, not another timed run.

Reuse rules predicted the cached count exactly on 226/296 calls and within one 832-token block on 292/296. The estimate converted rendered characters to tokens. Edits near the start and reformatting old messages explain much of the lost reuse.

Time and reuse analysis

Exact prefix reuse versus approximate suffix reuse

If an early passage changes, later internal states were calculated using the old passage. Reusing those states is an approximation even when the later words are unchanged. The CLM paper’s suffix-cache technique explores that tradeoff; it was not used in these exact-prefix results.

Detailed engine rules · vLLM’s explanation of prefix caching