How to reproduce the context tests
The tasks, graders, agent harness, comparison driver and probes, with what others can and cannot rerun.
Everything except the server is public. The tasks, graders, agent harness, comparison driver and probes are in the lab’s repository. They run against any OpenAI-compatible server with tool calling, and the harness can be checked on an ordinary computer with a fake server, no GPU needed.
What you can and cannot rerun
All measurements used Qwen3.8-27B with FP8 weights and a 16-bit attention cache, on two Intel Arc Pro B70 GPUs, one request at a time.
| Part | Reproducible by others? |
|---|---|
| Task generators and graders | Yes. Python, CPU only; the same seed gives the same task. |
| Agent harness patches | Yes. The summary, plain and improved self-editing agents subclass the paper’s released harness without changing it (clm_baselines.py, clm_improved.py, ctxfold.py). |
| Stub checks | Yes. A fake OpenAI server plays the model; needs Docker and the harness, no GPU. |
| Comparison driver and probes | Yes, against any OpenAI-compatible server with tool calling. Scores and times will differ with another model, server or machine. |
| The lab’s server | No. It runs the lab’s own vLLM build with its add-ons, started by the lab’s launcher. The exact-cache add-on’s source is public, but the image it runs in is not. |
The agents run on the upstream CLM harness at commit 18dc111 (CC BY-NC 4.0, non-commercial), installed with pip install -e in a Python 3.12 environment. The scripts’ default paths point at the lab’s machine: set VENV to your environment and pass --tokenizer to the task generators.
git clone https://github.com/steveseguin/b70-optimization-lab
D=b70-optimization-lab/experiments/qwen38-27b-b70/scripts/context
export VENV=/path/to/clm-venv
Full harness notes: how requests are built, budgets, sandboxing and the exact prompts.
1. The server
The lab started one two-card server with its research launcher and kept it up while a client script ran. Its settings:
- Window: 262,144 tokens (
--max-model-len 262144), one sequence at a time. - Exact prefix cache:
--prefix-cache align; the prompt is read in 832-token pieces, one cache block each, so a resumed read does the same arithmetic as a fresh one; a kept state every 13,312 tokens (--prefix-cache-retention-interval=13312). Add-on: b70-prefix-cache-exact. - Drafting on: the model’s own draft head, five tokens ahead.
- Tool calling:
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3. The harness needs these; without them requests with tools are refused.
On another server, the same tool-calling flags and a window that holds the budget plus the answer allowance are what matter. How the exact cache works.
2. Check the harness without a GPU
stub-checks.sh starts fake_openai_server.py on a free port and runs tiny trials against it: summaries fire, old thinking is dropped from requests, delivered batches are never lost, guards refuse what they should, and the archive is accepted by the grader. It prints one PASS or FAIL line per check.
stub-checks.sh <out_dir>
STUB=1 $D/smoke.sh <out_dir>/context-stub-smoke
STUB=1 SIZES=4000 OUT_DIR=<out_dir>/context-stub-second $D/second-comparison.sh
The second-comparison stub run also includes two rule probes: an agent that stores stream data in a file must be voided; one that keeps it in context must not.
3. Make the tasks
Each generator writes Harbor tasks with their grader. In memory mode stream data may live only in the model’s context, and a file holding it voids the run; in notes mode files are allowed.
- make_ledger_tasks.pyLine-format ledger: about 160 counters, SET/ADD/DEL lines with irrelevant memos, 24 questions at the end.
- make_kvstream_tasks.pyKey-value stream: batches of 100 SETs of 24-word values, then 24 lookups.
- make_sparse_prose_tasks.pyNarrative text with a few real counter changes.
--densitysets changes per 2,000 tokens,--wordswrites numbers as words,--surprise Kadds K hidden questions about earlier text (the retention test). - make_prose_ledger_tasks.pyDense narrative ledger with pronouns, corrections and distractors; harder than one call can read reliably.
make_ledger_tasks.py OUT_DIR --tokens 60000 120000 180000 --mode memory|notes [--seeds 0 1]
make_kvstream_tasks.py OUT_DIR --tokens N --mode memory|notes
make_sparse_prose_tasks.py OUT_DIR --tokens 60000 120000 480000 [--density 3] [--words] ...
[--mode memory|notes] [--seeds 0 1] [--batch-tokens 2000] [--n-batches N]
make_prose_ledger_tasks.py OUT_DIR --tokens 60000 120000 480000 --mode memory|notes [--seeds 0 1]
4. Run a comparison
second-comparison.sh runs each arm on each task as its own trial and skips finished trials on a re-run. Choose with ARMS (run in the order given), KINDS (for example sparse), SIZES, SEEDS and SPARSE_ARGS (options for the sparse prose generator). DRY_RUN=1 prints the plan and a time estimate; touch $OUT_DIR/STOP stops between trials.
OUT_DIR=/tmp/x API_BASE=http://unused DRY_RUN=1 SUBSET=core $D/second-comparison.sh # plan + estimate
API_BASE=... OUT_DIR=<out_dir>/context-clm-second-$(date +%Y%m%d) SUBSET=core $D/second-comparison.sh
The retention test is what retention-run.sh runs, in two blocks (B32ira B32ir E32r, then C32 Ar):
SPARSE_ARGS="--density 3 --words --surprise 12" ARMS="B32ira B32ir E32r" KINDS=sparse SEEDS="0" SIZES=120000 \
OUT_DIR=<out_dir>/ret120 SUBSET=core ./second-comparison.sh
| Arm | What it is |
|---|---|
| A | No management and no budget: the whole task in the 262K window; stream data only in context. |
| Aw | A, plus every tool result states the window used and left; a fetch that cannot fit is refused with a request to write down what is needed first. |
| Ar | Keep everything with the window shown: the baseline for the reading tasks. |
| B32 | The paper’s self-editing agent, unmodified, 32K budget, data only in context. |
| C32 | Summarise when the context reaches 75% of a 32K budget, data only in context. |
| D32 | Self-editing, 32K budget, files allowed. |
| E32 | No management, 32K budget, files allowed. |
| E32r | The same plain files-allowed agent, on the reading tasks. |
| B32i | Revised self-editing: delivered batches are never rolled back, a room check before each fetch, a pinned state file, old thinking dropped. |
| B32in | B32i with thinking off on routine calls and on where judgement is needed. |
| B32io | B32i with thinking off on every call. |
| B32ir | Read mode for narrative text: the model reads each report itself and keeps one line per counter; a drop is refused unless every counter mentioned has a line. |
| B32ira | B32ir plus a read-only archive of dropped batches and a recall search tool. |
5. Check each batch and report
- replay_state_check.pyCompares the agent’s working table after every batch with the true state, and names each error: missed update, missed delete, distractor applied, lost at an edit and so on. CPU only.
- reading_run_report.pyPer trial: score, answer counts, per-batch accuracy, calls with thinking on and off, tokens written, a reading/writing/tool time split, refusals, and whether the model wrote a parser.
replay_state_check.py TRIAL_DIR|JOB_DIR|runs/ [...] [--json out.json]
reading_run_report.py PATH [PATH ...] [--json out.json]
6. Probes
- qwen38-fp8-long-context-probe.pyRecall mode (
--question-sets): several questions per length over one ledger. Edge mode (--bisect): narrows the longest length that still answers every code. - qwen38-fp8-prefix-cache-exactness-probe.pyIs an answer served from the cache identical to a cold one? Save against a cache-off server, then compare on the cache server.
- qwen38-fp8-edit-cost-probe.pyWait for the first token after one edit, by context length and edit position.
- prose_fold_probe.pyCan one call read a narrative batch? Each call gets the true state before the batch, so errors do not compound.
- calibrate-reading.shRuns that probe over a grid of densities and difficulty switches and prints the hardest setting that passes as
SPARSE_ARGS.
python3 qwen38-fp8-long-context-probe.py --base-url http://127.0.0.1:18196 --model qwen38-27b-fp8 \
--tokenizer <model_dir> --lengths 8000,30000,60000,120000,200000,250000 --out probe.json
python3 qwen38-fp8-prefix-cache-exactness-probe.py --base-url http://127.0.0.1:18132 --save control.json
python3 qwen38-fp8-prefix-cache-exactness-probe.py --base-url http://127.0.0.1:18132 --compare-with control.json \
--out cache.json
python3 qwen38-fp8-edit-cost-probe.py --base-url http://127.0.0.1:18196 --lengths 30000,120000,200000 --out edit-cost.json
API_BASE=http://127.0.0.1:18124/v1 MODEL_NAME=qwen38-27b-fp8 \
prose_fold_probe.py TASK_DIR [TASK_DIR ...] [--thinking on|off|both] [--variant plain|reason|both]
Measured results · Remembering what was dropped · Lab research summary