CLM or files: what changes?
CLM edits what the model sees next. Files keep records it can look up later. They can be used together; “files with editing available” is that combination.
Same model, different memory rules
Every row used Qwen3.8-27B and could run code. CLM was an editing method, not a different trained model. The comparison changed the prompts, the surrounding application and permission to keep data outside context.
| Approach | Where useful information lives |
|---|---|
| Original CLM | Inside the editable message history. No separate data archive was allowed in this arm. |
| Revised CLM | Inside a protected working context, with a state file loaded into every call. Its contents still use the 32K budget. |
| Files, no editing | In an external archive and state files. Tools return selected text or calculated results; message history is not rewritten. |
| Files, editing available | The same external storage, plus the original CLM mechanism for rewriting message history. |
But revised CLM has a file too?
What matters is how the file is used. The application automatically places the revised agent’s STATE.txt into the next prompt. A 2K-token state therefore uses 2K tokens of context. It is a convenient editing interface, not extra memory outside the budget.
An archive can contain a million tokens while a tool returns only a 100-token excerpt. Only the returned text joins the prompt. The numerical state for that active prompt still lives in GPU memory.
What can each remember?
Suppose “10 apples, add 3, remove 2” becomes “apples: 11.” That compact state answers a question about the current total. If the original updates were discarded, it cannot explain which transaction removed two apples. A complete saved archive lets a tool find that transaction later.
CLM can also use such an archive when permitted. The lab’s context-only rule isolated its ability to maintain useful state inside a small window; it is not a restriction inherent to CLM.
Why did the file agents finish sooner?
Saved scripts updated saved totals with little model output. On the 121K ledger, both file agents got 24/24 in 1.9 minutes; revised CLM took 15 and 21 minutes, scoring 24/24 and 21/24. The file-and-editing run never needed a context edit, so its result does not show a benefit from adding editing.
All results · What the revised agent changed and what failed
Who was told to do what?
The original CLM prompt teaches transcript editing. Revised CLM adds an explicit routine: update the state immediately after a batch, then remove the processed batch. File-enabled agents receive permission to save data, but choose their own file layout and processing code.
Read the prompt sources and model setup · Detailed run analysis