SkinDeepRESEARCHSteve Seguin

Working past the context limit

Which ways of managing an AI’s memory actually helped?

How can an AI finish a long job when its working memory is limited? We tested a stream of counter updates, like transactions in a ledger, then asked for 24 final totals.

ApproachCorrectTime
Large window24/2426 min
Summaries24/2441 min
Self-editing (CLM)19/2464 min
Revised self-editing, run 124/2415 min
Revised self-editing, run 221/2421 min
Files + plain agent24/241.9 min
Files + self-editing24/241.9 min
Remove old thinkingNo answerstopped at 2.6 h

Files let the model keep a small working context while code maintained the totals. Both file-using approaches finished in 1.9 minutes with all 24 answers right.

The stream contained 121K tokens; the file-using runs held under 9K at once. The original self-editing run lost five delivered batches. Full comparison and test details.

Other ways to handle more context

ApproachResult
Prefix cachingRepeated 30K prompt: first token in 0.8 s instead of 11.4 s. Standard-check writing speed stayed near 89 tokens/s.
Bigger active contextObserved writing speed: about 127 tokens/s at 8K, versus 27 at 250K, in separate prompt tests.
CPU text cleanup1.83% fewer tokens with conservative cleanup.
Active cache on diskNot benchmarked. The 200K-token attention cache alone needs about 12.2 GiB.

The practical difference

Files store the records. Active context is what the model can use right now. Keeping records outside the conversation can extend a job. It does not give the GPU unlimited working memory.