SkinDeepRESEARCHSteve Seguin

Past the window: what the agent still remembers

A 240K-token history behind a 16K or 32K live window, with questions planted at known distances back: which ways of managing the window still answer them.

Short answer: keeping old text word for word in a searchable archive, beside a small pinned state, was the only method that stayed above the simple baselines on every kind of question at a 32K window. It was also among the quickest: 15 minutes per run, against 13 to 20 for the others. Methods that only drop or compress old text forgot what lay more than about 25 batches back, or reported old values as if they were still current.

The question

An agent that works on one task for a long time builds up far more history than its live window can hold. Something has to leave the window. What can the agent still answer about things that happened long ago?

The same Qwen3.8-27B model read a 240K-token stream of short stories in 147 batches, with about three counter changes per batch. Its live window was capped at 16K or 32K tokens, so the history was 7 to 15 times larger than what it could see at once. Every 25 batches it was asked three questions about a batch a known distance back:

At the end it answered 14 final questions about the current state. Distance is counted in batches back from the question.

Four ways to manage the window

Window-drop
When the window is full, the oldest turns are dropped. Nothing is kept. This is the simplest baseline, and it is lossy.
Compaction
What coding agents commonly do: old tool output is replaced with a short stub and old thinking is removed, then the oldest turns go. Also lossy. Called D or P in the evidence files.
Self-editing, read mode
The model reads each batch and keeps name value lines in a pinned state at the top; processed batches leave the window. How self-editing with a protected state works.
Self-editing + archive and search
The same pinned state, and every batch that leaves the window goes word for word to a read-only archive. A recall tool searches it. How the archive works.

Results

Each score is right/asked. The distance columns split the same probe questions by how far back the answer was stated; the type columns split them by question type. At 16K the scores are pooled over two seeds, except where noted. The 32K rows are one seed each.

MethodWindowSeedsProbes right1–4 back5–24 back25–99 backCurrent valueOld valueWho did itFinal questionsAllMinutes per run
Window-drop (W32)32K111/158/83/50/25/52/54/514/1425/2920
Compaction (P32)32K113/157/84/52/25/55/53/54/14 (6 stale)17/2913
Self-editing, read mode (B32ir)32K18/154/83/51/25/53/50/514/1422/2916
Self-editing + archive and search (B32ira)32K113/157/84/52/25/54/54/514/1427/2915
Window-drop (W16)16K216/3011/115/140/59/103/104/1034/3550/6522
Compaction (D16)16K219/308/119/142/510/108/101/1035/3554/6519-20
Self-editing, read mode (B16ir)16K29/305/114/140/58/101/100/1035/3544/6514-20
Self-editing + archive and search (B16ira)16K112/157/84/51/23/54/55/58/14 (2 stale)20/2925-36

Second round: two B70 GPUs, server window 65,536 tokens, exact prefix cache, one run at a time. Seed 1 is void: it hit the step cap and answered keys that were not asked. Only seed 0 is scored. Probe analysis for this table.

First round: 16K window, seed 0, one GPU
MethodWindowSeedsProbes right1–4 back5–24 back25–99 backCurrent valueOld valueWho did itFinal questionsAllMinutes per run
Window-drop (W16)16K110/158/82/50/25/52/53/514/1424/2922
Compaction (D16)16K110/155/83/52/25/55/50/514/1424/2919
Self-editing, read mode (B16ir)16K14/153/81/50/23/51/50/514/1418/2920
Self-editing + archive and search (B16ira)16K112/157/84/51/23/54/55/58/14 (2 stale)20/2936

One B70 GPU with a 26,624-token server window. These seed-0 runs are included in the pooled 16K rows above. Probe analysis for this table.

Cost of the 32K runs
Method at 32KTokens writtenModel callsPeak live contextArchive searches
Window-drop (W32)73,96015927,949–
Compaction (P32)51,92615721,128–
Self-editing, read mode (B32ir)55,16034623,887–
Self-editing + archive and search (B32ira)37,57736027,17920

The archive method wrote the fewest tokens of the 32K methods and searched its archive 20 times, with one rollback.

Earlier run: 48K window over a 480K history (2026-10-08)
MethodWindowSeedsProbes right1–4 back5–24 back25–99 back100–399 backCurrent valueOld valueWho did itFinal questionsAllMinutes per run
Window-drop (W48)48K123/339/913/141/60/411/118/114/1110/1033/4371
Compaction (D48)48K120/339/99/142/60/411/118/111/1110/1030/4357

Reported in the program charter. Per-type splits and run times are from the same run's probe analysis; its raw files are not in the lab repository yet. The summarising arm at 48K was cancelled at the 4-hour limit after 17 summaries; the self-editing arms at 48K never ran. Program charter.

What it shows

Limits

The 32K results are one seed; the 16K results are two seeds, and the archive method has only one valid 16K seed. Each distance bin holds few questions (as few as 2 at 25–99 batches back), so a difference of one or two answers is not decisive. One synthetic task family, one model and one machine. The stream reached 240K tokens, not the millions this program is aimed at.

Lab note with both rounds and the post-mortem · Raw evidence: checks, logs, probe analyses and server settings · Website measurement record

What we are trying to solve now

The target is an agent that works on one task for weeks or months, with 1 to 10 million tokens of history, while its live window stays capped: 200K tokens on a dedicated GPU, or 32K–48K when the GPU is shared. The model must stay fast, accurate and on goal. The model’s own arithmetic stays exact; how the history is managed may lose information, but then the loss is measured.

Why the model limits the options

Qwen3.8-27B is a hybrid. Most of its layers carry a running summary state forward instead of keeping one entry per past token. That state cannot have a piece cut out of the middle. The exact operations are: add new text at the end, or roll back to a saved checkpoint and read everything after it again. Cutting a block out of the middle of an ordinary attention cache is only approximate too: the later entries were computed while looking at the removed text. So an exact edit in the middle means rolling back and re-reading the tail, which the second, idle GPU could do between turns.

How the next tests will measure it

Next steps

  1. A second seed at 32K for all four methods.
  2. Harden the archive method for small windows: limit how much one search returns, tell the model how to free room instead of refusing again, and answer with what it has when steps run out.
  3. Add the rule-adherence and procedure-reuse questions.
  4. Measure saving and restoring the attention cache from fast disk for recalled batches, so recalled text need not be read again from scratch.

The full program charter: the problem, every idea so far and the benchmark design · Recommendations so far