SkinDeepRESEARCHSteve Seguin

Ways to work with more context

Compare what each approach keeps, what it costs and what the experiments showed.

Files, summaries and self-editing can keep a task going beyond one context window. They do it by controlling what stays in view. A larger window holds more at once; a prefix cache avoids repeated reading. Neither creates unlimited memory.

Choose an approach for a diagram, a worked example and its tradeoffs. “Decode” means writing the answer; total task time also includes reading and tool use.

ApproachWhat changes, and the result
Bigger active windowHolds more text at once. The window opened to 262K positions; longer active context slowed generation.
Exact prefix cacheReuses the unchanged start. Repeat questions began sooner; writing stayed near 89 tokens/s in the standard check.
Files, no context editingCode reads saved records and keeps the totals. 24/24 in 1.9 minutes on the 121K ledger.
Files, editing availableAdds permission to rewrite the conversation. 24/24 in 1.9 minutes on the 121K ledger; no edit was needed.
CLM self-editingRewrites selected parts of the working history. Original agent: 19/24 in 64 minutes; its harness lost five delivered batches.
Periodic summariesReplaces old history with a short account. It got 24/24 in 41 minutes; making summaries added work.
Revised self-editingProtects incoming batches and pins the current state. 24/24 and 21/24 on two 121K runs; 24/24 on a 478K stream.
Remove old thinkingDrops earlier reasoning from later calls. One ledger run repeatedly rebuilt its state and returned no answer after 2.6 hours.
Park a cache on diskSaves an inactive session for later restoration. Save/restore takes time; not benchmarked here.
Stream active cache from disk/RAMTrades repeated transfers or CPU work for capacity. Requires engine support; not benchmarked here.
CPU text cleanupRemoves repetition and formatting noise. Conservative cleanup reduced tokens by 1.83%.

What worked best in this experiment?

The task was a running ledger: process 20 batches of updates, then report 24 current values. The 121K tokens of input contained much more text than the small table needed to answer. That made it a good task for code and external files.

Why files helped so much · Full comparison, including failures

Can any of these remember everything forever?

No finite context can hold every detail of an ever-growing history. Files can retain an archive up to available storage; retrieval decides what to read. A summary or state table can support a long task when it preserves everything that task needs. If future questions may require arbitrary original details, retain the originals too.

Other approaches: retrieval, rolling windows and compression

Retrieval-augmented generation (RAG) selects passages from an external collection. File search is a simple version of the same idea; this ledger used code and files, not a separately benchmarked vector database. A rolling window discards old text, and a model-written summary can omit details. Both can keep a session moving while losing information. Reducing the precision of the numerical cache saves memory but changes its numbers; that was not part of these context experiments.

Recursive reading of documents and selective cache eviction are related approaches, not measured competitors in the ledger table. Follow the research review for the wider literature.