SkinDeepRESEARCHSteve Seguin

Research results

Where should processing stop?

Qwen has 24 transformer layers. Here is what happens when we stop earlier.

Test results · Decision methods

Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.

Layer 6 of 2475% skipped
Layer 12 of 2450% skipped
Layer 18 of 2425% skipped
Layer 24 of 240% skipped
Each square is one transformer block. Filled squares run; outlined squares are skipped.

A small classifier reads Qwen’s internal numbers at the chosen layer and returns SAFE or BLOCK. It does not write a reply.

Earlier is cheaper. Accuracy is uneven.

Stop afterCorrect / 100Toxic missed / 50
Early: layer 67517
Halfway: layer 127120
Late: layer 187816
Full model: layer 247718

These use the same small neural classifier design at each checkpoint. Layer 18 did slightly better than layer 24 here; a one-message difference is not a reliable ranking.

Skipped means transformer blocks, not the same percentage of model parameters or elapsed time. These fixed-depth rows come from stored internal states; each depth was not separately timed.

Linear classifiers and the larger test
LayerLinear correct / 100Toxic missed / 50
66917
127117
187515
247716

The nonlinear classifier has 896 inputs, 64 hidden units and two outputs. The linear classifier maps the same 896 numbers directly to two outputs. Both learn from 384 messages; Qwen stays frozen.

All four depths on 600 messages · Training and every comparison · Recorded results

Let each message stop at a different point