Ways to work with more context
Compare what each approach keeps, what it costs and what the experiments showed.
Files, summaries and self-editing can keep a task going beyond one context window. They do it by controlling what stays in view. A larger window holds more at once; a prefix cache avoids repeated reading. Neither creates unlimited memory.
Choose an approach for a diagram, a worked example and its tradeoffs. “Decode” means writing the answer; total task time also includes reading and tool use.
| Approach | What changes, and the result |
|---|---|
| Bigger active window | Holds more text at once. The window opened to 262K positions; longer active context slowed generation. |
| Exact prefix cache | Reuses the unchanged start. Repeat questions began sooner; writing stayed near 89 tokens/s in the standard check. |
| Files, no context editing | Code reads saved records and keeps the totals. 24/24 in 1.9 minutes on the 121K ledger. |
| Files, editing available | Adds permission to rewrite the conversation. 24/24 in 1.9 minutes on the 121K ledger; no edit was needed. |
| CLM self-editing | Rewrites selected parts of the working history. Original agent: 19/24 in 64 minutes; its harness lost five delivered batches. |
| Periodic summaries | Replaces old history with a short account. It got 24/24 in 41 minutes; making summaries added work. |
| Revised self-editing | Protects incoming batches and pins the current state. 24/24 and 21/24 on two 121K runs; 24/24 on a 478K stream. |
| Remove old thinking | Drops earlier reasoning from later calls. One ledger run repeatedly rebuilt its state and returned no answer after 2.6 hours. |
| Park a cache on disk | Saves an inactive session for later restoration. Save/restore takes time; not benchmarked here. |
| Stream active cache from disk/RAM | Trades repeated transfers or CPU work for capacity. Requires engine support; not benchmarked here. |
| CPU text cleanup | Removes repetition and formatting noise. Conservative cleanup reduced tokens by 1.83%. |
What worked best in this experiment?
The task was a running ledger: process 20 batches of updates, then report 24 current values. The 121K tokens of input contained much more text than the small table needed to answer. That made it a good task for code and external files.
Why files helped so much · Full comparison, including failures
Can any of these remember everything forever?
No finite context can hold every detail of an ever-growing history. Files can retain an archive up to available storage; retrieval decides what to read. A summary or state table can support a long task when it preserves everything that task needs. If future questions may require arbitrary original details, retain the originals too.
Other approaches: retrieval, rolling windows and compression
Retrieval-augmented generation (RAG) selects passages from an external collection. File search is a simple version of the same idea; this ledger used code and files, not a separately benchmarked vector database. A rolling window discards old text, and a model-written summary can omit details. Both can keep a session moving while losing information. Reducing the precision of the numerical cache saves memory but changes its numbers; that was not part of these context experiments.
Recursive reading of documents and selective cache eviction are related approaches, not measured competitors in the ledger table. Follow the research review for the wider literature.