Longer context: the measured results
Task outcomes, speed, recall and the source records behind the explanations.
Files completed the ledger fastest. Caching shortened the wait before an answer. Longer active context slowed generation.
Completing a long task
A synthetic ledger delivered 20 batches totaling about 121K input tokens. Updates set, add to or delete roughly 160 counters; accompanying memos were irrelevant. The final question asked for 24 current values. Each batch could be fetched only once. Original arms used seed 0; revised self-editing used seeds 0 and 1.

Every task result, context budget and token count
| Strategy | External files | Context budget | Correct / 24 | Elapsed time | Tokens generated |
|---|---|---|---|---|---|
| Large window, exact cache on | no | full window | 24 | 26 min | 89K tokens |
| Original self-editing agent | no | 32K | 19 | 64 min | 233K tokens |
| Summarise at 75% of the budget | no | 32K | 24 | 41 min | 211K tokens |
| Self-editing, files allowed | yes | 32K | 24 | 1.9 min | 7K tokens |
| No management, files allowed | yes | 32K | 24 | 1.9 min | 8K tokens |
| Large window, earlier reasoning removed | no | full window | no answer | stopped at 2.6 h | 397K tokens |
| Revised self-editing, seed 0 | no | 32K | 24 | 15 min | 67K tokens |
| Revised self-editing, seed 1 | no | 32K | 21 | 21 min | 91K tokens |
Both file-using agents answered 24/24 in about 1.9 minutes. Their peak active contexts were 8.2K and 8.9K tokens. The equal-score large-window and summary runs took about 26 and 41 minutes: approximately 14× and 22× the elapsed time in this particular trial.
These totals combine generation, prompt processing and tool use. They are not per-token decode speeds. The “large window” arm imposed no context-management budget, but the model still filtered batches through code; it did not retain every raw memo in its conversation.
Why the failures matter
The original self-editing run got 19/24. Its five missed answers exactly matched a ledger with batches 2, 3, 4, 5 and 7 omitted: the harness rolled back those already-delivered batches. The model’s recorded state edits were not the source of these five misses.
The earlier 12-trial key-value comparison also does not rank context policies: all arms wrote data to files, and incomplete file writes caused the losses. Some summary calls were empty after reasoning used the output allowance.
The separate old-thinking-removal trial was stopped after 2.6 hours without a final answer. It generated about 397K tokens while repeatedly reconstructing state and hitting the output cap.
Revised self-editing scored 24/24 in 15 minutes and 21/24 in 21 minutes. Every batch was retained. On the second seed, a bug in the model’s rewritten update script stopped applying deletes from batch 17, causing three wrong totals.
Verified-fold and beyond-window trials
A fix that retains and checks the update script is under test. It and the 480K stream have no completed scored result in this snapshot.
The completed 121K ledger exceeded the 32K working budget but remained below the model’s full 262K position limit. The tasks were locally constructed, not a reproduction of the paper’s benchmark scores.
Where the time went: reconstruction from saved calls
| Run | Total seconds | Estimated writing + call overhead | Estimated reading | Tool seconds |
|---|---|---|---|---|
| Large window | 1589 | 1461 | 110 | 2.7 |
| Original self-editing | 3841 | 3476 | 336 | 5.8 |
| Summaries | 2464 | 2068 | 253 | 4.4 |
| Files + self-editing | 112 | 76 | 16 | 2.7 |
| Files + plain agent | 111 | 80 | 15 | 2.5 |
| Revised self-editing, seed 0 | 896 | 758 | 109 | 4.5 |
| Revised self-editing, seed 1 | 1258 | 1095 | 130 | 5.6 |
Wall and tool times come from run records. Reading time is estimated from cold-read rates; writing is the remaining call time and includes call overhead. Summary-call time is estimated too. The columns omit other harness and setup time.
Most time went into generation, much of it thinking. All file-using tool commands together took 2.5–2.7 seconds. This does not isolate disk I/O, buffered versus durable writes or per-token decode rates.
Longer context slows generation
The server accepted a 262,144-token configured window; the probes supplied up to about 250K input tokens with room for an answer. Weights were FP8; the attention cache stayed at 16-bit precision.

Speed values and how they were measured
| Context | Output tokens/s | Replies measured |
|---|---|---|
| 8K | 127 | 1 |
| 30K | 99 | 1 |
| 60K | 70 | 20 |
| 120K | 49 | 20 |
| 160K | 40 | 20 |
| 200K | 35 | 20 |
| 230K | 31 | 20 |
| 250K | 27 | 20 |
8K and 30K: one chat recall prompt each, cache disabled. 60K–250K: medians over 20 six-code replies per length, two filler styles, exact prefix cache enabled. These are separate prompt tests, not a matched speed comparison across methods. Short replies and draft acceptance affect the measured output rates.
Cold reading speed and the first empty-answer result
| Input tokens | Cold time to first token | Input tokens/s |
|---|---|---|
| 7,848 | 2.4 s | 3288 |
| 29,865 | 9.8 s | 3059 |
| 59,836 | 22.1 s | 2704 |
| 119,720 | 54.7 s | 2187 |
| 199,588 | 114.4 s | 1745 |
| 249,599 | 160.8 s | 1552 |
These are the initial uncached bare-text probe’s input-processing measurements. Its short rows ran out of the 96-token output allowance, and at 250K both drafting settings immediately ended the answer. A follow-up using chat formatting answered at 213.5K and 250K, so the apparent bare-text cutoff was not an engine window limit.
The first probe’s drafting-on and drafting-off token sequences matched at all tested lengths, including the empty 250K response. That equality does not itself establish successful recall at every length. Later chat and recall results answer that separate question.
How caching changes the wait · Original protocol, corrections and results
Can it find the right record?
Two prompt styles, six lengths, ten questions per length and six codes per question: 708/720 codes correct. Each question sampled records spread through the context.

Recall data table
| Context | Codes buried in ordinary words | Wall of look-alike codes |
|---|---|---|
| 60K | 60 of 60 | 60 of 60 |
| 120K | 60 of 60 | 58 of 60 |
| 160K | 59 of 60 | 58 of 60 |
| 200K | 58 of 60 | 57 of 60 |
| 230K | 60 of 60 | 60 of 60 |
| 250K | 60 of 60 | 58 of 60 |
All 12 misses returned another record’s real code. Half were the preceding record; half had a similar record number. Errors appeared at different positions, and results did not worsen steadily: both styles were perfect at 230K.
The preset rule required at least 59/60 in both styles at a length and every shorter tested length. Only 60K met that continuous-range rule. It is a result on these similar-code prompts, not a universal reliable-context limit.
Raw data: codes among ordinary words · Raw data: wall of codes
Reusing an unchanged prefix
| Prompt or edit | First token without reuse | First token with reuse |
|---|---|---|
| Repeated 30K prompt | 11.4 s | 0.8 s |
| Middle edit in 30K, first request | 11.3 s | 6.7 s |
| Same edit, repeated | 11.3 s | 0.7 s |
| Question over 200K (separate probes) | 114 s | 2.0 s |
With drafting on, the matched cache experiment produced identical output token sequences in 99/99 cases: eleven prompts, nine variants each, up to 30K input and 48 output tokens. Both servers used 832-token reading chunks. Two additional cold-versus-cached checks at 120K also matched.
| Server | First pass | Repeat pass | Exact per pass |
|---|---|---|---|
| Cache off | 88.5 tokens/s | 88.7 tokens/s | 12/12 |
| Cache on | 89.6 tokens/s | 89.7 tokens/s | 12/12 |
The separate standard gate compared the shipped 4,096-token read configuration with the 832-token exact-cache configuration. Both returned 12/12 reference answers on the first and second passes. Their roughly 1% decode-rate difference does not establish a speed gain.
The cached 200K questions started in about 2 seconds; the original separate cold run took 114 seconds. Fixed small chunks raised cold-read latency: 30K went from about 9.8 to 11.4 seconds (16% longer), and 120K from about 55 to 69.6 seconds (about 27% longer).
Checkpoint cost and equality scope
The target model has recurrent layers as well as full attention. A reusable point needs the matching attention prefix and saved recurrent state. Checkpoints cost approximately three 832-token attention blocks’ worth of memory, with drafting and allocator details affecting actual allocation.
The add-on excludes generation-produced blocks and retains periodic reading checkpoints. An edit resumes before the change and recomputes what follows; it does not reuse stale post-edit states. The 99-case drafting test compared output tokens, not scores. Multi-user scheduling, one-card execution and disk restore were not established by this test.
How the cache works · Cache experiments · 99-case data · Standard decode gate data
CPU cleanup and reasoning history
| Rule | Context freed |
|---|---|
| Conservative cleanup on three transcripts (combined rules) | 1.83% of approximately 4.99M tokens |
| Large tool outputs offloaded (separate rule) | 7.50% of transcript tokens; later hidden-line quotation in 54/152 outputs |
| Earlier reasoning in 277 separate 27B model calls | 9.9% median input share, 55.8% maximum |
The cleanup census used roughly 4.99 million tokens from three existing agent transcripts. The 1.83% combined cleanup result is a text reduction, not an equal-answer model test. The 7.50% output-offload rule was measured separately and cannot simply be added to it.
54/152 large outputs had a line from the hidden middle quoted later, suggesting possible read-back work. It was not a measured retrieval experiment. Reasoning-history shares came from a different set of 277 model calls.
Disk cache: calculations, not a completed benchmark
The full-attention cache payload is 64 KiB/token for this target model at 16 bits. At 200,000 tokens that is 13.1072 GB, or 12.207 GiB, before other state. Saving text, parking this cache and streaming it while generating are different operations.
| Assumed sustained storage read rate | Time to move the 200K cache once |
|---|---|
| 1 GB/s | At least 13.11 seconds |
| 5 GB/s | At least 2.62 seconds |
| 10 GB/s | At least 1.31 seconds |
These are hypothetical single-transfer lower bounds, not measured tokens/s. Neither disk parking/restoration nor active disk streaming has a completed result in this snapshot. Explanation, assumptions and memory sizes.
Related result: short label decisions
One-step constrained labels and normal decoding without thinking agreed on 80/80 easy items; both got 78/80 correct. On the 24-item thinking subset, labels agreed on 24/24 and got 23/24 correct. Median response times were about 0.09 seconds without thinking and 0.59 with thinking.
Skipping reasoning supplied the saving; decoding a short label was already about 0.09 seconds. This is not a hard-reasoning benchmark or a context-offload speedup. Full one-step result.
How far to generalize these results
The agent comparison uses one synthetic task family and one machine. Original arms use one seed; revised self-editing uses two. The task was designed to have compact current state and irrelevant memos. Arbitrary documents or exact quotations can require much more retained information. Exact-match scoring checks the requested answers, not every possible question about the history.
One request ran at a time. Timings include different prompt shapes and reading configurations where identified. Do not use them as an SSD benchmark, a many-user capacity estimate, or a guarantee for another model. The source record includes proposals as well as completed trials; pending work has no measured score here.
Hardware and research snapshot
Qwen3.8-27B with FP8 weights and a 16-bit attention cache, two Intel Arc Pro B70 GPUs, R314 runtime, one request at a time. Research snapshot: 2026-10-05 14:37 UTC (10:37 EDT). K means approximately one thousand tokens; output speeds include drafting unless stated.
Original research
- Research summary (Markdown)Completed findings and pending work in the October 5 snapshot.
- Window, recall and one-step decisions (Markdown)Preregistered probes, corrections and results.
- Exact prefix cache tests (Markdown)99-case comparison and the separate decode-rate gate.
- How the cache works (Markdown)Detailed engine analysis, state sizes and reuse rules.
- Self-editing and ledger comparisons (Markdown)Task rules, failure analysis and revised-agent design.
- CPU cleanup census (Markdown)Token accounting and potential retrieval costs.
- Where the time went (Markdown)Saved-run timing reconstruction and prefix reuse.
- Paper and related research (Markdown)CLM, alternatives and proposed disk-cache experiments.
- Run instructions and scripts
Website measurement record and aggregation notes · Latest research scripts