SkinDeepRESEARCHSteve Seguin

Keep records in files

A long task can use a small working context.

Keep the records outside the conversation and load what matters. Files without context editing and files with editing available use code to keep the working conversation small.

A shelf holds many records, while an arrow brings a few selected excerpts to a small working area.
A large archive can support a small working context.

What changed in the ledger test?

The model processed a long stream of updates, then answered 24 questions about the final totals. With files, both agents answered all 24 in about 1.9 minutes and kept active context below 9K tokens.

Large window: 26 minutes, 24 of 24 correct. Summaries: 41 minutes, 24 correct. Original self-editing: 64 minutes, 19 correct. Revised self-editing: 15 minutes and 24 correct on seed 0, 21 minutes and 21 correct on seed 1. Both file-using agents: 1.9 minutes, 24 correct.
121K ledger: seed 0 for each approach, plus seed 1 for revised self-editing. The large-window seed-1 run returned no answer; see the full comparison.

Files cut the amount of text generated to about 7–8K tokens, versus 89K with the large window and 211K with summaries. Code kept the totals; the model did much less reading and writing.

Does writing a file slow decoding?

A file operation takes time between model calls. It does not make the model read the entire archive for every new output token. The smaller active context can also make generation faster.

All tool commands combined took 2.5–2.7 seconds in the file-using runs, including other command work. This does not isolate disk-write time or show that disk access is free. Streaming the model’s numerical cache from disk is a different operation.

Where the rest of the time went

The saved-call analysis estimates that generation and call overhead used about 76–80 seconds of the 111–112-second file-using runs. Reading used roughly 15–16 seconds. Much of the generated text was thinking. These are reconstructed phases, not per-token timing measurements.

Timing reconstruction and file operations

What should be saved?

A table of current totals was enough for this task. An exact quotation or a question about an old change may need the original records. Save those too when later questions could depend on them.

Saving a record preserves its bytes; finding the right passage and using it correctly still matter.

A failure to avoid: saving only part of a record

In the earlier key-value trials, some shell pipelines saved and displayed a batch at the same time. Stopping the display early also stopped the save, leaving truncated files. Save the complete record first and preview it afterward.

Original failure analysis