SkinDeepRESEARCHSteve Seguin

Research results

Cascade: test results

The recorded decisions, mistakes and processing time for the tiny-model fallback experiment.

Back to the cascade explanation

Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.

Which tiny model? Google’s BERT miniature with two layers and 4.37 million trained parameters. We fine-tuned it on 384 ToxicChat messages to return SAFE or BLOCK directly, without generating text. Uncertain messages go to a separate Qwen2.5-0.5B-Instruct classifier.

Architecture, training and model revision

The complete comparison

MethodCorrect / 100Toxic missed / 50Benign blocked / 50Time for 100
Tiny model only722080.31 s
Tiny model → Qwen7914732.46 s
Qwen for every message7815764.75 s

The cascade corrected three Qwen mistakes and introduced two new ones, including one toxic-message miss. Its higher total accuracy does not remove that trade-off.

Which model answered?

RouteMessagesCorrect
Tiny model alone5143
Tiny model, then Qwen4936

Every message ran BERT’s two layers. Only the 49 fallback messages ran Qwen’s 24 layers; 51 avoided Qwen entirely.

Inspect all 100 decisions

SAFE and BLOCK follow the dataset’s non-toxic and toxic labels. IDs identify source rows; raw chat text is not republished. Time includes both models when Qwen was needed.

Message IDDataset labelCascade answerAnswered byRequest time
test:1227SAFESAFEQwen448.1 ms
test:151BLOCKBLOCKTiny model2.2 ms
test:937SAFEBLOCKQwen824.8 ms
test:276BLOCKBLOCKTiny model1.8 ms
test:4767SAFEBLOCKQwen1760.3 ms
test:4411SAFESAFETiny model1.5 ms
test:390SAFESAFETiny model1.5 ms
test:556BLOCKBLOCKTiny model2.4 ms
test:1277SAFESAFEQwen551.7 ms
test:143BLOCKSAFEQwen513.3 ms
test:3970BLOCKBLOCKTiny model7.6 ms
test:1298BLOCKBLOCKQwen433.6 ms
test:3994SAFEBLOCKTiny model1.5 ms
test:4858SAFEBLOCKQwen1086.5 ms
test:5064SAFESAFEQwen468.7 ms
test:4506SAFESAFEQwen455.1 ms
test:444BLOCKBLOCKTiny model8.2 ms
test:1008BLOCKSAFEQwen389.4 ms
test:1166BLOCKBLOCKTiny model6.6 ms
test:96BLOCKBLOCKTiny model8.0 ms
test:84BLOCKBLOCKQwen493.0 ms
test:4520SAFESAFETiny model1.8 ms
test:4407SAFESAFETiny model3.0 ms
test:4512SAFEBLOCKQwen413.9 ms
test:3767SAFESAFETiny model1.8 ms
test:1027SAFESAFEQwen500.2 ms
test:4107BLOCKSAFEQwen441.9 ms
test:1060SAFESAFEQwen715.9 ms
test:74BLOCKBLOCKQwen830.2 ms
test:931SAFESAFEQwen1696.5 ms
test:4BLOCKSAFETiny model1.9 ms
test:2BLOCKBLOCKQwen342.6 ms
test:4989SAFESAFETiny model1.5 ms
test:651SAFESAFEQwen611.2 ms
test:254BLOCKSAFETiny model2.4 ms
test:1186SAFESAFEQwen657.4 ms
test:1055BLOCKBLOCKTiny model1.9 ms
test:4956SAFEBLOCKQwen557.4 ms
test:206BLOCKSAFETiny model2.0 ms
test:677BLOCKBLOCKQwen452.7 ms
test:142BLOCKBLOCKTiny model4.8 ms
test:961BLOCKBLOCKTiny model4.3 ms
test:3949SAFESAFEQwen475.9 ms
test:960BLOCKSAFEQwen730.5 ms
test:1312SAFESAFEQwen405.1 ms
test:282BLOCKBLOCKQwen416.7 ms
test:485SAFEBLOCKQwen504.2 ms
test:4582SAFESAFEQwen458.8 ms
test:3929SAFESAFETiny model1.8 ms
test:5076BLOCKSAFETiny model2.4 ms
test:482BLOCKSAFETiny model2.0 ms
test:4160SAFESAFEQwen1053.9 ms
test:4838SAFESAFETiny model3.1 ms
test:5006BLOCKSAFEQwen371.0 ms
test:179BLOCKBLOCKTiny model1.9 ms
test:1219SAFESAFEQwen510.2 ms
test:178BLOCKBLOCKQwen669.3 ms
test:656BLOCKBLOCKQwen1814.4 ms
test:908SAFESAFETiny model2.3 ms
test:90SAFESAFEQwen453.0 ms
test:1134SAFESAFEQwen1747.6 ms
test:4070SAFESAFETiny model2.7 ms
test:547SAFESAFETiny model1.6 ms
test:1048SAFESAFEQwen445.2 ms
test:14BLOCKBLOCKQwen430.7 ms
test:3814BLOCKBLOCKTiny model2.7 ms
test:1289BLOCKBLOCKTiny model1.8 ms
test:930SAFESAFEQwen519.1 ms
test:4036SAFESAFETiny model1.6 ms
test:133BLOCKBLOCKTiny model1.6 ms
test:4153BLOCKBLOCKTiny model5.9 ms
test:77BLOCKSAFEQwen468.1 ms
test:4481BLOCKBLOCKTiny model1.5 ms
test:1299SAFESAFETiny model2.1 ms
test:1100SAFESAFETiny model2.1 ms
test:85SAFESAFETiny model1.7 ms
test:3801SAFESAFETiny model2.4 ms
test:329BLOCKBLOCKTiny model3.1 ms
test:54BLOCKBLOCKQwen420.3 ms
test:848SAFESAFEQwen480.3 ms
test:1056BLOCKBLOCKTiny model3.8 ms
test:457BLOCKBLOCKQwen761.2 ms
test:3779SAFESAFETiny model2.2 ms
test:169BLOCKBLOCKQwen1774.7 ms
test:770BLOCKSAFETiny model1.6 ms
test:4692SAFESAFEQwen470.7 ms
test:1479SAFESAFETiny model1.6 ms
test:4443BLOCKBLOCKTiny model1.8 ms
test:1326SAFESAFETiny model2.2 ms
test:4442SAFESAFETiny model1.6 ms
test:868SAFESAFETiny model1.7 ms
test:3897BLOCKBLOCKTiny model2.9 ms
test:196BLOCKSAFEQwen439.5 ms
test:60BLOCKBLOCKQwen606.4 ms
test:3797SAFESAFETiny model2.3 ms
test:5056BLOCKSAFETiny model1.8 ms
test:1076BLOCKBLOCKTiny model2.5 ms
test:4034BLOCKBLOCKQwen429.8 ms
test:170BLOCKBLOCKQwen422.9 ms
test:448SAFESAFEQwen397.5 ms
How this test was run

The tiny model trained on 384 messages; the older Qwen classifier trained on 1,400. Separate development examples selected the confidence thresholds. These 100 evaluation messages had already been inspected in earlier work.

Timing uses one warm CPU pass per method, rotating their order for each message. Input preparation and actual fallback work are included; loading is excluded. This is exploratory evidence, not validated moderation performance.

Source records

Exact sample IDs · Download all predictions · Download timings · Training and reproduction · Adversarial review