Keep records in files
A long task can use a small working context.
Keep the records outside the conversation and load what matters. Files without context editing and files with editing available use code to keep the working conversation small.

What changed in the ledger test?
The model processed a long stream of updates, then answered 24 questions about the final totals. With files, both agents answered all 24 in about 1.9 minutes and kept active context below 9K tokens.

Files cut the amount of text generated to about 7–8K tokens, versus 89K with the large window and 211K with summaries. Code kept the totals; the model did much less reading and writing.
Does writing a file slow decoding?
A file operation takes time between model calls. It does not make the model read the entire archive for every new output token. The smaller active context can also make generation faster.
All tool commands combined took 2.5–2.7 seconds in the file-using runs, including other command work. This does not isolate disk-write time or show that disk access is free. Streaming the model’s numerical cache from disk is a different operation.
Where the rest of the time went
The saved-call analysis estimates that generation and call overhead used about 76–80 seconds of the 111–112-second file-using runs. Reading used roughly 15–16 seconds. Much of the generated text was thinking. These are reconstructed phases, not per-token timing measurements.
What should be saved?
A table of current totals was enough for this task. An exact quotation or a question about an old change may need the original records. Save those too when later questions could depend on them.
Saving a record preserves its bytes; finding the right passage and using it correctly still matter.
A failure to avoid: saving only part of a record
In the earlier key-value trials, some shell pipelines saved and displayed a batch at the same time. Stopping the display early also stopped the save, leaving truncated files. Save the complete record first and preview it afterward.