Research results
Where should processing stop?
Qwen has 24 transformer layers. Here is what happens when we stop earlier.
Test results · Decision methods
Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.
A small classifier reads Qwen’s internal numbers at the chosen layer and returns SAFE or BLOCK. It does not write a reply.
Earlier is cheaper. Accuracy is uneven.
| Stop after | Correct / 100 | Toxic missed / 50 |
|---|---|---|
| Early: layer 6 | 75 | 17 |
| Halfway: layer 12 | 71 | 20 |
| Late: layer 18 | 78 | 16 |
| Full model: layer 24 | 77 | 18 |
These use the same small neural classifier design at each checkpoint. Layer 18 did slightly better than layer 24 here; a one-message difference is not a reliable ranking.
Skipped means transformer blocks, not the same percentage of model parameters or elapsed time. These fixed-depth rows come from stored internal states; each depth was not separately timed.
Linear classifiers and the larger test
| Layer | Linear correct / 100 | Toxic missed / 50 |
|---|---|---|
| 6 | 69 | 17 |
| 12 | 71 | 17 |
| 18 | 75 | 15 |
| 24 | 77 | 16 |
The nonlinear classifier has 896 inputs, 64 hidden units and two outputs. The linear classifier maps the same 896 numbers directly to two outputs. Both learn from 384 messages; Qwen stays frozen.
All four depths on 600 messages · Training and every comparison · Recorded results