SkinDeepRESEARCHSteve Seguin

Longer context: the measured results

Task outcomes, speed, recall and the source records behind the explanations.

Files completed the ledger fastest. Caching shortened the wait before an answer. Longer active context slowed generation.

Completing a long task

A synthetic ledger delivered 20 batches totaling about 121K input tokens. Updates set, add to or delete roughly 160 counters; accompanying memos were irrelevant. The final question asked for 24 current values. Each batch could be fetched only once. Original arms used seed 0; revised self-editing used seeds 0 and 1.

Large window: 26 minutes, 24 of 24 correct. Summaries: 41 minutes, 24 correct. Original self-editing: 64 minutes, 19 correct. Revised self-editing: 15 minutes and 24 correct on seed 0, 21 minutes and 21 correct on seed 1. Both file-using agents: 1.9 minutes, 24 correct.
Completed 121K-token ledger runs. Revised self-editing used two seeds; the other approaches used seed 0. Full results
Every task result, context budget and token count
StrategyExternal filesContext budgetCorrect / 24Elapsed timeTokens generated
Large window, exact cache onnofull window2426 min89K tokens
Original self-editing agentno32K1964 min233K tokens
Summarise at 75% of the budgetno32K2441 min211K tokens
Self-editing, files allowedyes32K241.9 min7K tokens
No management, files allowedyes32K241.9 min8K tokens
Large window, earlier reasoning removednofull windowno answerstopped at 2.6 h397K tokens
Revised self-editing, seed 0no32K2415 min67K tokens
Revised self-editing, seed 1no32K2121 min91K tokens

Both file-using agents answered 24/24 in about 1.9 minutes. Their peak active contexts were 8.2K and 8.9K tokens. The equal-score large-window and summary runs took about 26 and 41 minutes: approximately 14× and 22× the elapsed time in this particular trial.

These totals combine generation, prompt processing and tool use. They are not per-token decode speeds. The “large window” arm imposed no context-management budget, but the model still filtered batches through code; it did not retain every raw memo in its conversation.

Why the failures matter

The original self-editing run got 19/24. Its five missed answers exactly matched a ledger with batches 2, 3, 4, 5 and 7 omitted: the harness rolled back those already-delivered batches. The model’s recorded state edits were not the source of these five misses.

The earlier 12-trial key-value comparison also does not rank context policies: all arms wrote data to files, and incomplete file writes caused the losses. Some summary calls were empty after reasoning used the output allowance.

The separate old-thinking-removal trial was stopped after 2.6 hours without a final answer. It generated about 397K tokens while repeatedly reconstructing state and hitting the output cap.

Call-by-call analysis · Readable explanation

Revised self-editing scored 24/24 in 15 minutes and 21/24 in 21 minutes. Every batch was retained. On the second seed, a bug in the model’s rewritten update script stopped applying deletes from batch 17, causing three wrong totals.

Verified-fold and beyond-window trials

A fix that retains and checks the update script is under test. It and the 480K stream have no completed scored result in this snapshot.

The completed 121K ledger exceeded the 32K working budget but remained below the model’s full 262K position limit. The tasks were locally constructed, not a reproduction of the paper’s benchmark scores.

Where the time went: reconstruction from saved calls
RunTotal secondsEstimated writing + call overheadEstimated readingTool seconds
Large window158914611102.7
Original self-editing384134763365.8
Summaries246420682534.4
Files + self-editing11276162.7
Files + plain agent11180152.5
Revised self-editing, seed 08967581094.5
Revised self-editing, seed 1125810951305.6

Wall and tool times come from run records. Reading time is estimated from cold-read rates; writing is the remaining call time and includes call overhead. Summary-call time is estimated too. The columns omit other harness and setup time.

Most time went into generation, much of it thinking. All file-using tool commands together took 2.5–2.7 seconds. This does not isolate disk I/O, buffered versus durable writes or per-token decode rates.

Method, engine-log cross-checks and reuse calculations

Why files helped · Original reported task results

Longer context slows generation

The server accepted a 262,144-token configured window; the probes supplied up to about 250K input tokens with room for an answer. Weights were FP8; the attention cache stayed at 16-bit precision.

Observed output rate: 127 tokens per second at 8K context, 99 at 30K, 70 at 60K, 49 at 120K, 40 at 160K, 35 at 200K, 31 at 230K and 27 at 250K.
Separate prompt tests: one reply each at 8K and 30K; medians of 20 replies at the longer lengths.
Speed values and how they were measured
ContextOutput tokens/sReplies measured
8K1271
30K991
60K7020
120K4920
160K4020
200K3520
230K3120
250K2720

8K and 30K: one chat recall prompt each, cache disabled. 60K–250K: medians over 20 six-code replies per length, two filler styles, exact prefix cache enabled. These are separate prompt tests, not a matched speed comparison across methods. Short replies and draft acceptance affect the measured output rates.

Cold reading speed and the first empty-answer result
Input tokensCold time to first tokenInput tokens/s
7,8482.4 s3288
29,8659.8 s3059
59,83622.1 s2704
119,72054.7 s2187
199,588114.4 s1745
249,599160.8 s1552

These are the initial uncached bare-text probe’s input-processing measurements. Its short rows ran out of the 96-token output allowance, and at 250K both drafting settings immediately ended the answer. A follow-up using chat formatting answered at 213.5K and 250K, so the apparent bare-text cutoff was not an engine window limit.

The first probe’s drafting-on and drafting-off token sequences matched at all tested lengths, including the empty 250K response. That equality does not itself establish successful recall at every length. Later chat and recall results answer that separate question.

How caching changes the wait · Original protocol, corrections and results

Can it find the right record?

Two prompt styles, six lengths, ten questions per length and six codes per question: 708/720 codes correct. Each question sampled records spread through the context.

Codes correct out of 60 per style. Ordinary words and look-alike codes: 60K, 60 and 60; 120K, 60 and 58; 160K, 59 and 58; 200K, 58 and 57; 230K, 60 and 60; 250K, 60 and 58.
Correct codes out of 60 for each context length and filler style.
Recall data table
ContextCodes buried in ordinary wordsWall of look-alike codes
60K60 of 6060 of 60
120K60 of 6058 of 60
160K59 of 6058 of 60
200K58 of 6057 of 60
230K60 of 6060 of 60
250K60 of 6058 of 60

All 12 misses returned another record’s real code. Half were the preceding record; half had a similar record number. Errors appeared at different positions, and results did not worsen steadily: both styles were perfect at 230K.

The preset rule required at least 59/60 in both styles at a length and every shorter tested length. Only 60K met that continuous-range rule. It is a result on these similar-code prompts, not a universal reliable-context limit.

Raw data: codes among ordinary words · Raw data: wall of codes

Reusing an unchanged prefix

Prompt or editFirst token without reuseFirst token with reuse
Repeated 30K prompt11.4 s0.8 s
Middle edit in 30K, first request11.3 s6.7 s
Same edit, repeated11.3 s0.7 s
Question over 200K (separate probes)114 s2.0 s

With drafting on, the matched cache experiment produced identical output token sequences in 99/99 cases: eleven prompts, nine variants each, up to 30K input and 48 output tokens. Both servers used 832-token reading chunks. Two additional cold-versus-cached checks at 120K also matched.

ServerFirst passRepeat passExact per pass
Cache off88.5 tokens/s88.7 tokens/s12/12
Cache on89.6 tokens/s89.7 tokens/s12/12

The separate standard gate compared the shipped 4,096-token read configuration with the 832-token exact-cache configuration. Both returned 12/12 reference answers on the first and second passes. Their roughly 1% decode-rate difference does not establish a speed gain.

The cached 200K questions started in about 2 seconds; the original separate cold run took 114 seconds. Fixed small chunks raised cold-read latency: 30K went from about 9.8 to 11.4 seconds (16% longer), and 120K from about 55 to 69.6 seconds (about 27% longer).

Checkpoint cost and equality scope

The target model has recurrent layers as well as full attention. A reusable point needs the matching attention prefix and saved recurrent state. Checkpoints cost approximately three 832-token attention blocks’ worth of memory, with drafting and allocator details affecting actual allocation.

The add-on excludes generation-produced blocks and retains periodic reading checkpoints. An edit resumes before the change and recomputes what follows; it does not reuse stale post-edit states. The 99-case drafting test compared output tokens, not scores. Multi-user scheduling, one-card execution and disk restore were not established by this test.

How the cache works · Cache experiments · 99-case data · Standard decode gate data

CPU cleanup and reasoning history
RuleContext freed
Conservative cleanup on three transcripts (combined rules)1.83% of approximately 4.99M tokens
Large tool outputs offloaded (separate rule)7.50% of transcript tokens; later hidden-line quotation in 54/152 outputs
Earlier reasoning in 277 separate 27B model calls9.9% median input share, 55.8% maximum

The cleanup census used roughly 4.99 million tokens from three existing agent transcripts. The 1.83% combined cleanup result is a text reduction, not an equal-answer model test. The 7.50% output-offload rule was measured separately and cannot simply be added to it.

54/152 large outputs had a line from the hidden middle quoted later, suggesting possible read-back work. It was not a measured retrieval experiment. Reasoning-history shares came from a different set of 277 model calls.

Plain-language implications · Census, methods and caveats

Disk cache: calculations, not a completed benchmark

The full-attention cache payload is 64 KiB/token for this target model at 16 bits. At 200,000 tokens that is 13.1072 GB, or 12.207 GiB, before other state. Saving text, parking this cache and streaming it while generating are different operations.

Assumed sustained storage read rateTime to move the 200K cache once
1 GB/sAt least 13.11 seconds
5 GB/sAt least 2.62 seconds
10 GB/sAt least 1.31 seconds

These are hypothetical single-transfer lower bounds, not measured tokens/s. Neither disk parking/restoration nor active disk streaming has a completed result in this snapshot. Explanation, assumptions and memory sizes.

Related result: short label decisions

One-step constrained labels and normal decoding without thinking agreed on 80/80 easy items; both got 78/80 correct. On the 24-item thinking subset, labels agreed on 24/24 and got 23/24 correct. Median response times were about 0.09 seconds without thinking and 0.59 with thinking.

Skipping reasoning supplied the saving; decoding a short label was already about 0.09 seconds. This is not a hard-reasoning benchmark or a context-offload speedup. Full one-step result.

How far to generalize these results

The agent comparison uses one synthetic task family and one machine. Original arms use one seed; revised self-editing uses two. The task was designed to have compact current state and irrelevant memos. Arbitrary documents or exact quotations can require much more retained information. Exact-match scoring checks the requested answers, not every possible question about the history.

One request ran at a time. Timings include different prompt shapes and reading configurations where identified. Do not use them as an SSD benchmark, a many-user capacity estimate, or a guarantee for another model. The source record includes proposals as well as completed trials; pending work has no measured score here.

Hardware and research snapshot

Qwen3.8-27B with FP8 weights and a 16-bit attention cache, two Intel Arc Pro B70 GPUs, R314 runtime, one request at a time. Research snapshot: 2026-10-05 14:37 UTC (10:37 EDT). K means approximately one thousand tokens; output speeds include drafting unless stated.

Original research

Website measurement record and aggregation notes · Latest research scripts