Let the model edit its own context
CLM, summaries and state files: how a long task can use a short working history.
CLM gives the model an editable conversation. Here CLM means Context Language Model: the model can replace obsolete parts of its live transcript with useful state, rather than only appending new messages.
Explore the mechanisms: original self-editing, protected state or periodic summaries.
How is that different from a summary or a plan file?
A periodic summary replaces a large portion of the conversation at a chosen threshold. CLM can make smaller, targeted edits: update a counter, delete a superseded log, or keep an exact quotation. In both cases, information removed without a saved copy is gone.
An ordinary plan file lives outside the live conversation. The model sees it when the application loads it. The CLM mirror is different: an accepted edit changes the message history sent on the next call. Merely saving a transcript file does not make this happen.
What did the experiment find?
With external data files forbidden, the revised agent was faster, but only one of its two runs was fully correct.
| Self-editing agent | Correct | Time |
|---|---|---|
| Self-editing (CLM) | 19/24 | 64 min |
| Revised self-editing, run 1 | 24/24 | 15 min |
| Revised self-editing, run 2 | 21/24 | 21 min |
In the original run, the surrounding software discarded five batches after delivering them, and they could not be fetched again. That explains its five wrong answers. The model’s own state edits did not cause those misses.
The revised agent retained every batch. Its second run made a different mistake: the model rewrote its update script and stopped applying deletes from batch 17. That caused three wrong totals. Keeping data available and computing the answer correctly are separate requirements.
Why not just delete old thinking?
The model may have stored its working state there. Removing earlier thinking left one run repeatedly rebuilding its state, without a final answer after 2.6 hours. Keep explicit task state before attempting this optimization.
How much room did earlier reasoning use?
A median 9.9% of input across 277 calls in the initial key-value comparison, reaching 55.8% on one call. The stuck trial was a separate large-window ledger run. Measurements and outcome.
Can a small CPU cleaner solve the problem?
Conservative cleanup removed 1.83% of roughly five million tokens. That is too little to turn a 32K window into 100K. Moving large outputs to files saved more visible space, but later questions can require reading them back. Cleanup measurements.
Revised-agent details and the next tests
Runs 1 and 2 used ledger seeds 0 and 1. They took 15 and 21 minutes, generated 67K and 91K tokens, and peaked at 20K and 15K active context. The harness protects delivered batches, checks room before reading and keeps explicit state visible.
A verified-fold fix is being tested: retain one update script and check its self-test and handled-line count. The separate 478K stream scored 24/24 in 8.2 minutes; code processed most of its batches. The verified-fold fix has no completed score in this snapshot.
Two-seed results and script-bug analysis · Original comparison · Run instructions
Paper and wider research
The Context Language Models paper introduced this editable-context approach and evaluated it on other models and workloads. The lab’s key-value generator and running ledger are local experiments, not the paper’s released benchmark or a reproduction of its headline scores.
Context Language Models paper · Authors’ implementation · CPU-cleaning census