Research results
The original 600-message test
A real-message workload showed both the savings and the extra mistakes.
Run Qwen 0.5B on real messages and compare accuracy and time. All 24 layers run; this demo does not use the trained banking classifier.
Test results · Decision methods
Dataset: ToxicChat0124 · 600 archived, human-annotated messages: 81 toxic and 519 benign. No live moderation.
| Depth | Correct / 600 | Toxic missed / 81 |
|---|---|---|
| 6 of 24 layers | 506 | 23 |
| 12 of 24 layers | 519 | 26 |
| 18 of 24 layers | 516 | 23 |
| 24 of 24 layers | 522 | 19 |
The selected layer-12 path took 117.0 seconds versus 231.8 seconds at full depth. It missed seven more toxic messages and failed the reliability check.
Median of three complete passes per method, including input preparation. Full depth returns one trained label token. Only the selected shortcut and full-depth paths have this timing comparison; the table does not imply measured times at every depth.
Full results · Sampling and training · Independent audit · Later adaptive follow-up