Research results
Ways to make a decision sooner
Compare the measured speed and accuracy, then open a method for its test results.
Browser BERT comparisons; the recorded Qwen results remain separate.
Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.
All six methods below return numerical classifier decisions. “Full Qwen” means all 24 layers followed by a trained classifier, not a written reply. Early-exit methods also skip later transformer layers; they do not stop halfway through spelling a word.
Classifiers and stopping early
| Method | Correct / 100 | Time / message |
|---|---|---|
| Full Qwen + classifier | 77 | 563 ms |
| Stop at layer 12 | 71 | 281 ms |
| Learned stopping | 79 | 421 ms |
| Tiny model only | 72 | 3.1 ms |
| Tiny model → Qwen | 79 | 325 ms |
| With distillation | 80 | 674 ms |
The learned stop reduced time by about 34% against its full-model reference. The tiny-model cascade reduced it by about 50% against its own Qwen reference, but introduced one new toxic-message miss.
Same 100 messages for accuracy; training differs between methods. Full Qwen and layer 12 share a new three-pass timing run. The learned gate was timed separately against its own 636 ms full-depth reference; its 34% saving is from that paired run. Warm CPU time includes input preparation: 50 timed messages for distillation, 100 for the other timed rows. The cascade’s own Qwen reference scores 78/100 at 648 ms.
Which tiny model? Google’s BERT miniature with two layers and 4.37 million trained parameters. We fine-tuned it on 384 ToxicChat messages to return SAFE or BLOCK directly, without generating text. Uncertain messages go to a separate Qwen2.5-0.5B-Instruct classifier.
Architecture, training and model revision
Saving work at the output
Dataset: ToxicChat0124 · 50 previously inspected messages: 25 toxic and 25 benign. Qwen2.5-0.5B, warm CPU, two timing passes. All 24 layers run.
| Output method | Correct / 50 | Toxic missed / 25 | Time / message |
|---|---|---|---|
| Generate one SAFE/BLOCK token | 31 | 3 | 712 ms |
| Read just the two label scores | 31 | 3 | 665 ms |
6.6% less time, with identical decisions on these 50 messages. Read only Qwen’s existing SAFE and BLOCK scores, then return the chosen label. No new training, full-vocabulary calculation or text generation is needed; no transformer layers are skipped.
Both paths still missed 3 toxic messages and blocked 16 benign ones. This improvement is relative to one-token generation. The six classifier methods above already avoid text generation, so their recorded numbers are unchanged.
How each approach works
Other datasets and stress tests · Original 600-message comparison · Try output formats on real messages
New-message follow-up: more training and int8
More training raised tiny BERT from 81 to 82 correct out of 100, but it missed more toxic messages. A stricter fallback preserved Qwen’s decisions with little speed benefit. Int8 Qwen ran faster but fell from 86 to 53 correct. These recipes do not establish a useful quality-preserving improvement.
Queue throughput: same answers, fewer separate calls
Reusing the fixed instructions: measured results