Working past the context limit
Which ways of managing an AI’s memory actually helped?
How can an AI finish a long job when its working memory is limited? We tested a stream of counter updates, like transactions in a ledger, then asked for 24 final totals.
| Approach | Correct | Time |
|---|---|---|
| Large window | 24/24 | 26 min |
| Summaries | 24/24 | 41 min |
| Self-editing (CLM) | 19/24 | 64 min |
| Revised self-editing, run 1 | 24/24 | 15 min |
| Revised self-editing, run 2 | 21/24 | 21 min |
| Files + plain agent | 24/24 | 1.9 min |
| Files + self-editing | 24/24 | 1.9 min |
| Remove old thinking | No answer | stopped at 2.6 h |
Files let the model keep a small working context while code maintained the totals. Both file-using approaches finished in 1.9 minutes with all 24 answers right.
The stream contained 121K tokens; the file-using runs held under 9K at once. The original self-editing run lost five delivered batches. Full comparison and test details.
Other ways to handle more context
| Approach | Result |
|---|---|
| Prefix caching | Repeated 30K prompt: first token in 0.8 s instead of 11.4 s. Standard-check writing speed stayed near 89 tokens/s. |
| Bigger active context | Observed writing speed: about 127 tokens/s at 8K, versus 27 at 250K, in separate prompt tests. |
| CPU text cleanup | 1.83% fewer tokens with conservative cleanup. |
| Active cache on disk | Not benchmarked. The 200K-token attention cache alone needs about 12.2 GiB. |
The practical difference
Files store the records. Active context is what the model can use right now. Keeping records outside the conversation can extend a job. It does not give the GPU unlimited working memory.