Can disk replace GPU memory?
Text files, saved numerical state and active offloading solve different problems.
A nearly full GPU cannot gain a full 200K-token active context just by saving it to disk. The model still needs working memory. An offloading engine can trade transfers or CPU work for capacity, but needs buffers and a place to compute.
Three different uses of disk
- Save text and records
- Keep an archive and read selected parts. This is the approach that completed the ledger in 1.9 minutes with under 9K active context.
- Park an inactive cache
- Save the model’s numerical state between sessions, then restore it before use.
- Stream an active cache
- Move numerical state during generation because it cannot all stay in GPU memory.
How much memory does 200K need?
For this model, about 12.2 GiB for the attention cache alone, spread across two GPUs. Model weights and other working state need additional space.
That is much larger than the original text. Also, 200 KB is not 200K tokens: KB measures file bytes; tokens are pieces of text counted by the model.
Cache sizes and calculation
| Active tokens | Attention-cache payload across both GPUs |
|---|---|
| 8,192 | 0.50 GiB |
| 32,768 | 2.00 GiB |
| 60,000 | 3.66 GiB |
| 120,000 | 7.32 GiB |
| 200,000 | 12.21 GiB |
| 262,144 | 16.00 GiB |
The target model has 16 full-attention layers. At 16-bit precision: 16 layers × 2 arrays × 4 KV heads × 256 values × 2 bytes = 65,536 bytes per token across both GPUs. At 200,000 tokens this is 13.1072 GB, or 12.207 GiB.
This excludes weights, recurrent state, drafting, checkpoints, allocation padding and temporary workspace. Model shapes and accounting.
Does parking a cache slow writing?
Saving and restoring adds waiting time. Once the complete state is restored to active memory, ordinary generation need not keep reading disk. The state must still fit the chosen execution arrangement.
The lab proposed this path but did not complete a save/restore benchmark. A correct restore needs compatible weights, positions and all required runtime state.
What if it keeps fetching from disk?
Full attention consults earlier keys and values while generating. Repeated transfers can become a bottleneck. This is different from saving text files between model calls.
The lab did not benchmark active disk or RAM offload. Its roughly 35 tokens/s at 200K used an in-memory cache.
Example: moving the 200K cache once
| Assumed sustained storage read rate | Time to move the 200K cache once |
|---|---|
| 1 GB/s | At least 13.11 seconds |
| 5 GB/s | At least 2.62 seconds |
| 10 GB/s | At least 1.31 seconds |
These are hypothetical transfer lower bounds: 13.1072 GB divided by storage bandwidth. They are not measured model speeds. Computation and any additional unoverlapped upload cost more; partial residence, pipelining and accepted draft tokens change the calculation.
What should a smaller machine do?
Use a smaller active context with files or retrieval when the task allows it. Keeping everything active requires enough memory, a smaller model or cache representation, or an engine with explicit offloading support. Saving a transcript does not enlarge the model’s native window.
Original research and related storage documentation
Disk and offload research notes · LMCache’s storage explanation