Past the window: what the agent still remembers
A 240K-token history behind a 16K or 32K live window, with questions planted at known distances back: which ways of managing the window still answer them.
Short answer: keeping old text word for word in a searchable archive, beside a small pinned state, was the only method that stayed above the simple baselines on every kind of question at a 32K window. It was also among the quickest: 15 minutes per run, against 13 to 20 for the others. Methods that only drop or compress old text forgot what lay more than about 25 batches back, or reported old values as if they were still current.
The question
An agent that works on one task for a long time builds up far more history than its live window can hold. Something has to leave the window. What can the agent still answer about things that happened long ago?
The same Qwen3.8-27B model read a 240K-token stream of short stories in 147 batches, with about three counter changes per batch. Its live window was capped at 16K or 32K tokens, so the history was 7 to 15 times larger than what it could see at once. Every 25 batches it was asked three questions about a batch a known distance back:
- Current value: what a counter is now.
- Old value: what a counter was before it was overwritten.
- Who did it: a small detail that only appeared once, in the story text.
At the end it answered 14 final questions about the current state. Distance is counted in batches back from the question.
Four ways to manage the window
- Window-drop
- When the window is full, the oldest turns are dropped. Nothing is kept. This is the simplest baseline, and it is lossy.
- Compaction
- What coding agents commonly do: old tool output is replaced with a short stub and old thinking is removed, then the oldest turns go. Also lossy. Called D or P in the evidence files.
- Self-editing, read mode
- The model reads each batch and keeps
name valuelines in a pinned state at the top; processed batches leave the window. How self-editing with a protected state works. - Self-editing + archive and search
- The same pinned state, and every batch that leaves the window goes word for word to a read-only archive. A
recalltool searches it. How the archive works.
Results
Each score is right/asked. The distance columns split the same probe questions by how far back the answer was stated; the type columns split them by question type. At 16K the scores are pooled over two seeds, except where noted. The 32K rows are one seed each.
| Method | Window | Seeds | Probes right | 1–4 back | 5–24 back | 25–99 back | Current value | Old value | Who did it | Final questions | All | Minutes per run |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Window-drop (W32) | 32K | 1 | 11/15 | 8/8 | 3/5 | 0/2 | 5/5 | 2/5 | 4/5 | 14/14 | 25/29 | 20 |
| Compaction (P32) | 32K | 1 | 13/15 | 7/8 | 4/5 | 2/2 | 5/5 | 5/5 | 3/5 | 4/14 (6 stale) | 17/29 | 13 |
| Self-editing, read mode (B32ir) | 32K | 1 | 8/15 | 4/8 | 3/5 | 1/2 | 5/5 | 3/5 | 0/5 | 14/14 | 22/29 | 16 |
| Self-editing + archive and search (B32ira) | 32K | 1 | 13/15 | 7/8 | 4/5 | 2/2 | 5/5 | 4/5 | 4/5 | 14/14 | 27/29 | 15 |
| Window-drop (W16) | 16K | 2 | 16/30 | 11/11 | 5/14 | 0/5 | 9/10 | 3/10 | 4/10 | 34/35 | 50/65 | 22 |
| Compaction (D16) | 16K | 2 | 19/30 | 8/11 | 9/14 | 2/5 | 10/10 | 8/10 | 1/10 | 35/35 | 54/65 | 19-20 |
| Self-editing, read mode (B16ir) | 16K | 2 | 9/30 | 5/11 | 4/14 | 0/5 | 8/10 | 1/10 | 0/10 | 35/35 | 44/65 | 14-20 |
| Self-editing + archive and search (B16ira) | 16K | 1 | 12/15 | 7/8 | 4/5 | 1/2 | 3/5 | 4/5 | 5/5 | 8/14 (2 stale) | 20/29 | 25-36 |
Second round: two B70 GPUs, server window 65,536 tokens, exact prefix cache, one run at a time. Seed 1 is void: it hit the step cap and answered keys that were not asked. Only seed 0 is scored. Probe analysis for this table.
First round: 16K window, seed 0, one GPU
| Method | Window | Seeds | Probes right | 1–4 back | 5–24 back | 25–99 back | Current value | Old value | Who did it | Final questions | All | Minutes per run |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Window-drop (W16) | 16K | 1 | 10/15 | 8/8 | 2/5 | 0/2 | 5/5 | 2/5 | 3/5 | 14/14 | 24/29 | 22 |
| Compaction (D16) | 16K | 1 | 10/15 | 5/8 | 3/5 | 2/2 | 5/5 | 5/5 | 0/5 | 14/14 | 24/29 | 19 |
| Self-editing, read mode (B16ir) | 16K | 1 | 4/15 | 3/8 | 1/5 | 0/2 | 3/5 | 1/5 | 0/5 | 14/14 | 18/29 | 20 |
| Self-editing + archive and search (B16ira) | 16K | 1 | 12/15 | 7/8 | 4/5 | 1/2 | 3/5 | 4/5 | 5/5 | 8/14 (2 stale) | 20/29 | 36 |
One B70 GPU with a 26,624-token server window. These seed-0 runs are included in the pooled 16K rows above. Probe analysis for this table.
Cost of the 32K runs
| Method at 32K | Tokens written | Model calls | Peak live context | Archive searches |
|---|---|---|---|---|
| Window-drop (W32) | 73,960 | 159 | 27,949 | – |
| Compaction (P32) | 51,926 | 157 | 21,128 | – |
| Self-editing, read mode (B32ir) | 55,160 | 346 | 23,887 | – |
| Self-editing + archive and search (B32ira) | 37,577 | 360 | 27,179 | 20 |
The archive method wrote the fewest tokens of the 32K methods and searched its archive 20 times, with one rollback.
Earlier run: 48K window over a 480K history (2026-10-08)
| Method | Window | Seeds | Probes right | 1–4 back | 5–24 back | 25–99 back | 100–399 back | Current value | Old value | Who did it | Final questions | All | Minutes per run |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Window-drop (W48) | 48K | 1 | 23/33 | 9/9 | 13/14 | 1/6 | 0/4 | 11/11 | 8/11 | 4/11 | 10/10 | 33/43 | 71 |
| Compaction (D48) | 48K | 1 | 20/33 | 9/9 | 9/14 | 2/6 | 0/4 | 11/11 | 8/11 | 1/11 | 10/10 | 30/43 | 57 |
Reported in the program charter. Per-type splits and run times are from the same run's probe analysis; its raw files are not in the lab repository yet. The summarising arm at 48K was cancelled at the 4-hour limit after 17 summaries; the self-editing arms at 48K never ran. Program charter.
What it shows
- Archive and search at 32K is the first method above the lossy baselines on every axis at once. It had the best probe score (13/15), all 14 final answers, the best total (27/29), and at 15 minutes it beat window-drop and read mode. Only compaction finished sooner (13 minutes), with far worse final answers. The verbatim text is still there, so “who did it” questions can be answered.
- Compaction that keeps old state visible reports stale values as current. At 32K it answered the probes well (13/15), then got only 4 of 14 final questions right: six answers were values that used to be true. It also loses “who did it” details, because those live in the tool output it stubs out (1/10 at 16K).
- Window-drop is honest but forgets. It answered nothing correctly beyond about 25 batches back in any run: 0/5 at 16K, 0/2 at 32K, and 1/6 then 0/4 even with a 48K window.
- 16K starves the methods that work inside the window. With the batch, the state, the instructions and room to think all inside 16K, the self-editing read mode fell to 9/30. The archive method had no room either: the harness refused to fetch the next batch 121 times and refused to fold old batches 35 times, so old values stuck (8/14 final). That was lack of room, not a fault in how recalled text was written back. At 32K neither failure appeared.
- Speed did not degrade as history grew. Every method took about 4–10 minutes per 100K tokens of history, in both the first and second half of the stream.
- End-of-run questions alone do not separate methods. At 16K three methods answered every final question while scoring 4/15 to 10/15 on the probes. They mostly ask about state that is still in view.
Limits
The 32K results are one seed; the 16K results are two seeds, and the archive method has only one valid 16K seed. Each distance bin holds few questions (as few as 2 at 25–99 batches back), so a difference of one or two answers is not decisive. One synthetic task family, one model and one machine. The stream reached 240K tokens, not the millions this program is aimed at.
Lab note with both rounds and the post-mortem · Raw evidence: checks, logs, probe analyses and server settings · Website measurement record
What we are trying to solve now
The target is an agent that works on one task for weeks or months, with 1 to 10 million tokens of history, while its live window stays capped: 200K tokens on a dedicated GPU, or 32K–48K when the GPU is shared. The model must stay fast, accurate and on goal. The model’s own arithmetic stays exact; how the history is managed may lose information, but then the loss is measured.
Why the model limits the options
Qwen3.8-27B is a hybrid. Most of its layers carry a running summary state forward instead of keeping one entry per past token. That state cannot have a piece cut out of the middle. The exact operations are: add new text at the end, or roll back to a saved checkpoint and read everything after it again. Cutting a block out of the middle of an ordinary attention cache is only approximate too: the later entries were computed while looking at the removed text. So an exact edit in the middle means rolling back and re-reading the tail, which the second, idle GPU could do between turns.
How the next tests will measure it
- Probe families at 1–4, 5–24, 25–99, 100–399 and 400+ batches back: current values, overwritten values, small details, updated facts and questions the history cannot answer.
- A rule stated at the start that must still hold at the end, and a procedure learned early that must be reused late.
- A compaction-count axis: force a set number of compactions and plot accuracy against it.
- Costs for every method: time per 100K tokens of history, tokens spent on management, peak window and cache reuse.
Next steps
- A second seed at 32K for all four methods.
- Harden the archive method for small windows: limit how much one search returns, tell the model how to free room instead of refusing again, and answer with what it has when steps run out.
- Add the rule-adherence and procedure-reuse questions.
- Measure saving and restoring the attention cache from fast disk for recalled batches, so recalled text need not be read again from scratch.
The full program charter: the problem, every idea so far and the benchmark design · Recommendations so far