SkinDeepRESEARCHSteve Seguin

Research results

The original 600-message test

A real-message workload showed both the savings and the extra mistakes.

Test results · Decision methods

Dataset: ToxicChat0124 · 600 archived, human-annotated messages: 81 toxic and 519 benign. No live moderation.

Chosen shortcut: layer 1250% of blocks skipped
All 600 messages stopped at the same point. This was not an adaptive readiness detector.
DepthCorrect / 600Toxic missed / 81
6 of 24 layers50623
12 of 24 layers51926
18 of 24 layers51623
24 of 24 layers52219

The selected layer-12 path took 117.0 seconds versus 231.8 seconds at full depth. It missed seven more toxic messages and failed the reliability check.

Median of three complete passes per method, including input preparation. Full depth returns one trained label token. Only the selected shortcut and full-depth paths have this timing comparison; the table does not imply measured times at every depth.

Full results · Sampling and training · Independent audit · Later adaptive follow-up