SkinDeepRESEARCHSteve Seguin

Research results

Where should processing stop?

Qwen has 24 transformer layers. Here is what happens when we stop earlier.

Test results · Decision methods

Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.

Layer 6 of 2475% skipped
Layer 12 of 2450% skipped
Layer 18 of 2425% skipped
Layer 24 of 240% skipped
Each square is one transformer block. Filled squares run; outlined squares are skipped.

A small classifier reads Qwen’s internal numbers at the chosen layer and returns SAFE or BLOCK. It does not write a reply.

Earlier is cheaper. Accuracy is uneven.

Stop afterCorrect / 100Toxic missed / 50Time for 100
Early: layer 6751713.85 s
Halfway: layer 12712028.10 s
Late: layer 18781641.71 s
Full model: layer 24771856.26 s

These use the same small neural classifier design at each checkpoint. Layer 18 did slightly better than layer 24 here; a one-message difference is not a reliable ranking.

Skipped means transformer blocks, not the same percentage of model parameters or elapsed time. All four depths now have actual execution timings: three rotating paired passes on the same CPU. Every prediction and executed layer trace matched the stored results. Times above are the mean per 100 messages; loading is excluded.

All timing passes, mistakes and execution traces

Layer 12 skipped half the blocks and took about half the time, with six fewer correct answers. Layer 18 corrected eight errors but introduced seven others; its slightly higher total is not lossless performance.

Linear classifiers and the larger test
LayerLinear correct / 100Toxic missed / 50
66917
127117
187515
247716

The nonlinear classifier has 896 inputs, 64 hidden units and two outputs. The linear classifier maps the same 896 numbers directly to two outputs. Both learn from 384 messages; Qwen stays frozen.

All four depths on 600 messages · Training and every comparison · Recorded results

Let each message stop at a different point

Skip only the last one or two layers?

Separate follow-up: 50 reused messages, 25 toxic and 25 benign. New linear classifiers trained equally on 384 messages; three paired CPU timing passes.

Stop afterCorrect / 50Blocks skippedTime / messageNew toxic misses
Layer 22372/24 (8.3%)578 ms2
Layer 23361/24 (4.2%)618 ms1
Layer 24360/24 (0.0%)641 ms0

Layer 22 took 9.8% less time and layer 23 took 3.5% less. Both introduced toxic-message misses that the full-depth classifier avoided. “New toxic misses” compares the same messages with the layer-24 row, not the older 100-message test.

The classifier reads the internal state after the chosen block; later blocks never run. It does not wait for a word or inspect its first letter. This tests fixed shortcuts, not a detector of when an answer is ready.

Training, per-class errors and independent audit · All late-layer measurements · Different shortcut: avoid text generation