SkinDeepRESEARCHSteve Seguin

Remembering what was dropped

An archive of dropped text with a search tool answered questions about the past as well as summarising, in a quarter of the time.

The goal: when an agent has to make room, lose less than summarising does, and spend less time doing it. Moving dropped text to a searchable archive answered every question, including those about text it had dropped, in a quarter of the summariser’s time.

The test

The 119K-token narrative reading task, with twelve extra questions asked only at the end: values that were later overwritten, and small details in the text. 36 questions in all. Every approach had a 32K working budget except keep-everything, which used the full window.

ApproachFilesRightTimePeak contextTokens written
Table plus archive and recallarchive only36 of 366.1 min23K14K tokens
Table onlyno25 of 365.7 min20K15K tokens
Summarise at 75%no36 of 3626 min23K121K tokens
Keep everything in the windowno34 of 3614 min144K33K tokens
No management, files allowedyes3 of 366.6 min27K26K tokens

What it showed

How the archive works

  1. When the agent drops a batch from its working context, the batch goes word for word to an archive the model can read but not change.
  2. A recall search tool finds passages in that archive.
  3. Recalled text counts against the 32K budget like any other tool output.
  4. The grader voids the run if any other file holds stream data, so the archive is the only memory outside the context.

Nothing is deleted, only moved out of view. In this run the agent searched the archive 12 times. How summaries work · Why files helped on the ledger.

Limits

One seed and one task family, so treat the ranking as a first result. Summarising also got every answer here: with thinking on, its summaries kept the old values. Its cost was time and writing, not accuracy.

Lab note: Retention · Per-run summaries · How to reproduce · All results