Longer context, plain answers
Common questions about memory, disk, speed and what “unlimited” can mean.
The basics
What is a context window?
It is the amount of input and generated text the model can work with in one request. A token is a piece of text, not necessarily a word. Instructions, conversation history, tool results and the answer all use room. Reserve space for the answer instead of filling the entire limit with input.
Which methods allow unlimited context?
None gives a finite model unlimited simultaneous attention. Files and retrieval can support an archive larger than the window. Summaries and CLM self-editing can support an ongoing task by replacing old text with useful state. These approaches can keep going while storage, time and task requirements permit.
The 478K-token ledger exceeded the full window. Both file agents and the revised self-editing agent answered all 24 questions using code to process the stream with a small active context. Larger-stream results · How each method works.
What does CLM mean here? Is it a new model?
CLM means Context Language Model: a model that can edit its own live conversation through a file. These experiments used an existing Qwen3.8-27B model with tools and a surrounding agent harness. They did not train a new foundation model or enlarge its built-in position limit.
If the transcript is on disk, does the model remember it?
The application has a record. The model sees only what is supplied in the current call. A tool must load the relevant file, retrieve an excerpt, or compute an answer from the records. A filename by itself does not put the file’s contents into the model.
How is this different from keeping a plan or notes file?
A notes file is external state that can survive a session reset; the application chooses when to show it again. A CLM context-file edit changes the live message history used by the next model call. Both can preserve current goals and facts, but neither recovers information that was never recorded.
Speed
Does saving context to disk hurt decode speed?
It depends on what is saved. A text-file write is a tool operation; it does not add a read of the entire archive to every generated token. Parking a numerical cache adds a save/restore wait. Streaming an active cache during generation adds repeated transfers and can slow decoding substantially.
The file-using ledger finished much faster overall, but there was no isolated disk-write-overhead test and no disk-streamed cache benchmark. The three uses of disk.
Did the exact prefix cache change token generation speed?
Not meaningfully in the standard prompt check: approximately 88.5–88.7 tokens/s without it and 89.6–89.7 with it. The benefit was a shorter wait before the first token. One run per server configuration cannot turn that small difference into a reliable speedup claim.
Why is a cached long context still slower to write against?
Caching avoids recomputing old prompt states. Full-attention layers still consult those states during generation. More active context means more attention work and data movement. The observed output rate fell from about 127 tokens/s near 8K to 27 near 250K.
Why were files so much faster than keeping everything in context?
Code handled the ledger, active context stayed below 9K, and the agents generated only about 7–8K tokens. The large-window run generated 89K; the summary run generated 211K. The saved work was much larger than the file-operation cost in this task.
That does not imply a 14× speedup for reading any large document. What made this task suitable.
Does summarizing always make a task faster?
No. It shortens future context, but creating the summary costs model calls and generated tokens. An early rewrite also forces later prompt text to be processed again. Here, summaries got 24/24 but took 41 minutes, versus 26 minutes with the larger context and its prefix cache.
Is the first read faster too?
Not in this exact-cache implementation. Its smaller, consistent reading chunks made cold reads longer: about 11.4 versus 9.8 seconds at 30K, and 69.6 versus about 55 seconds at 120K. Reuse pays off on later requests with a matching prefix.
Disk and GPU memory
Can a GPU with a few KB free use a full 200K-token window from disk?
No, not by saving a file. For this model, 200,000 tokens require about 12.2 GiB of attention cache across two GPUs, plus other state and working memory. An offloading engine can trade transfers or CPU work for capacity, but still needs space to compute. This research did not demonstrate that setup.
Do 200 KB of text and a 200K context mean the same thing?
No. KB measures bytes; K in a context-window label normally means thousands of tokens. The tokenizer determines how many tokens a file contains. The numerical attention cache is much larger than the original text, and its size also depends on the model.
Could system RAM replace the SSD?
RAM removes the storage-read step, but offloaded state still has to reach the processor doing the attention work. Capacity, upload bandwidth, buffer space and engine support matter. Neither active RAM offload nor disk offload was benchmarked in these runs.
Can I save a conversation and resume it without rereading?
A compatible snapshot of the full numerical state can be designed for that. It must include the required attention and recurrent state, positions and metadata. Restoring it can avoid recomputing the prompt, but loading is not instantaneous and the active state must still fit the chosen execution setup. This was proposed, not demonstrated here.
Why not just open the biggest window?
The two-card setup accepted the model’s 262,144-position limit. That gives useful headroom, but longer prompts take longer to read, slow generation and can make precise lookup harder. A shorter working context can still be worthwhile when the hardware can hold more.
Accuracy and practical choices
What does “lossless” or “exact” actually mean?
Three different claims need different checks. A saved file can preserve every byte. A state table can preserve the answers needed for one task while omitting history. A numerical cache can preserve values, while equality to a fresh computation also depends on how the runtime uses them.
The exact-prefix trial matched output token sequences on 99 cases. That does not make summarization lossless, guarantee correct answers, or establish every future cache continuation. What was compared.
Can it reliably find a fact at 200K or 250K?
It answered questions at those lengths, but sometimes picked another record’s code. Across the recall test, 708 of 720 codes were correct; all 120 lookups at 60K were correct. At 250K it got 60/60 with ordinary-word filler and 58/60 in a wall of similar records. These are test results, not a guarantee for every document.
For exact record lookup, use deterministic search or code and show the relevant record. The full recall table.
Did CLM itself forget the five missing answers?
The run failed on five queried values, but its answers exactly matched a ledger missing five delivered batches. The harness had rolled those batches out after overflow. The recorded state edits did not account for those losses. Correct delivery and recovery rules matter as much as the editing strategy.
Should we remove old reasoning and redundant text?
Only with the task state preserved. Cleanup removed about 1.8% of the sampled transcripts, so it was a small saving. Old reasoning took more space, but removing it blindly left one run stuck for 2.6 hours without an answer. A short label task is a separate case: disabling new reasoning matched the thinking answers on 24 easy items.
Would a second GPU be better used to prune in the background?
This comparison used both GPUs together for the main model. A separate worker/pruner arrangement was not measured. It would change the main model’s available memory and speed and introduce another model’s editing decisions. The present results do not establish that it is a better use of the second card.
Will the same numbers apply to my model, GPU or many users?
No. Cache size depends on architecture and precision; decode rate depends on hardware, active length, workload and serving choices. These measurements used one Qwen3.8-27B request at a time on two B70 GPUs. Shared-cache concurrency and other hardware need separate measurements.
What should I take away for a real application?
Keep exact source records when later questions may need them. Use tools to retrieve or compute what the model needs now. Preserve explicit current state across edits, and judge the system by correct completed tasks and total elapsed time. Use the larger window as headroom and prefix caching to avoid repeated reading when the engine supports it correctly.