Remembering what was dropped
An archive of dropped text with a search tool answered questions about the past as well as summarising, in a quarter of the time.
The goal: when an agent has to make room, lose less than summarising does, and spend less time doing it. Moving dropped text to a searchable archive answered every question, including those about text it had dropped, in a quarter of the summariser’s time.
The test
The 119K-token narrative reading task, with twelve extra questions asked only at the end: values that were later overwritten, and small details in the text. 36 questions in all. Every approach had a 32K working budget except keep-everything, which used the full window.
| Approach | Files | Right | Time | Peak context | Tokens written |
|---|---|---|---|---|---|
| Table plus archive and recall | archive only | 36 of 36 | 6.1 min | 23K | 14K tokens |
| Table only | no | 25 of 36 | 5.7 min | 20K | 15K tokens |
| Summarise at 75% | no | 36 of 36 | 26 min | 23K | 121K tokens |
| Keep everything in the window | no | 34 of 36 | 14 min | 144K | 33K tokens |
| No management, files allowed | yes | 3 of 36 | 6.6 min | 27K | 26K tokens |
What it showed
- The archive beat summarising on time and writing, at equal correctness. Both answered 36 of 36. The archive agent took 6.1 min and wrote 14K tokens; the summariser took 26 min and wrote 121K tokens.
- A running table alone cannot answer about the past. It keeps only current values, so questions about overwritten values and details went blank or wrong: 25 of 36.
- The plain files agent collapsed because it has no guards. It spent its 32K budget thinking before it had finished reading and was stopped: 3 of 36. The same agent got every answer on the version without the extra questions.
How the archive works
- When the agent drops a batch from its working context, the batch goes word for word to an archive the model can read but not change.
- A
recallsearch tool finds passages in that archive. - Recalled text counts against the 32K budget like any other tool output.
- The grader voids the run if any other file holds stream data, so the archive is the only memory outside the context.
Nothing is deleted, only moved out of view. In this run the agent searched the archive 12 times. How summaries work · Why files helped on the ledger.
Limits
One seed and one task family, so treat the ranking as a first result. Summarising also got every answer here: with thinking on, its summaries kept the old values. Its cost was time and writing, not accuracy.
Lab note: Retention · Per-run summaries · How to reproduce · All results