Early tests
Stopping early
An easy task may not need every step of a model’s processing.
Compare accuracy and time on real messages. This demo uses a four-layer BERT model that can stop after layer 2; the Qwen tests below are separate.
A small classifier reads the numbers inside Qwen and predicts a label. At layers 6, 12 and 18, its confidence score decides whether to stop or continue.
Our 600-message chat test stopped at layer 12 every time. It missed 26 toxic messages, compared with 19 at full depth, and failed the reliability check. This was a fixed shortcut, not a detector of when an answer was ready.
Where this could be useful
Choose a smaller model, produce a result with fewer steps, or cancel a generation that is going wrong. These are different ways to save work; each needs its own quality and timing test.
- Route between small and large modelsUse the larger model when it adds value.
- Get an LLM result soonerUse less model work before returning an answer.
- Stop an unsafe image before it finishesCheck during generation and cancel early.
- Stop an image that misses your tasteSpend rendering time on more promising ideas.
More use cases: messages, devices and medical research
Messages and documents
- Live chat moderationReturn an allow, block or review decision.
- Route customer requestsSend a request to the right queue.
- Rank search resultsPut useful matches near the top.
- Sort an inboxSeparate receipts, requests and newsletters.
Actions
- Choose a game moveReturn a direction on each turn.
- Commands on a small deviceUse a short path for familiar commands.
Medical signals and scans
- ECG labels without a written reportReturn structured predictions from a heart trace.
- Check ECG recording qualityFlag an unusable trace for another recording.
- Mark regions in an MRIReturn a mask that a reviewer can inspect.
- MRI reconstruction with fewer model stepsStudy when further reconstruction stops helping.
- EEG sleep-stage labelsClassify recorded sleep segments directly.
Technical details and sources
Latest small experiments: learned stopping, model training and measured time · Adversarial review. These use 100 previously inspected messages and remain exploratory.
600 real chat decisions: results and measured runtime · Independent review · Watch recorded maze decisions.
Second-dataset validation · Testing an unknown-request output · Evidence requirements.
Separate implementation replay and portable classifiers · Changing-rule test: the current classifier failed · Adversarial reassessment.
Tests on 3,080 public banking queries | Stricter rule on a fresh reserve
Exactly how the classifier and stopping rule work · Real-query test protocol · Original per-query evidence
Checkpoints inside the model
Task heads at selected layers make candidate predictions. A gate decides whether to return, continue, or escalate. We begin with full-depth outputs, then measure intermediate readouts and actual early stopping.
Collecting intermediate states after running every layer is useful for training probes, but saves no inference computation. The runtime must skip later blocks. Timings must include the exit checks.
FastBERT, CALM, and LayerSkip provide relevant prior work in different settings.
Continuation and handoffs
Continuing within one model reuses compatible activations. Two independently trained models do not have interchangeable hidden states; a direct handoff needs an explicitly trained interface and comparison against simply reprocessing the original input.
CPU, GPU, and NPU placement are measured engineering choices. Copies, model residency, batching, cache repair, and fallback can eliminate theoretical savings. Parameter counts alone do not establish latency or cost.
Conditions for a useful shortcut
- Quality holds at a registered error tolerance.
- Accepted-request coverage is useful; rejecting everything is not success.
- Actual input-to-result time improves, including fallback.
- Changing rules, long context, unfamiliar inputs, and wrong high-confidence predictions remain visible in evaluation.