Research results
Check whether more layers are needed
Keep processing uncertain messages. Stop earlier on others.
Test results · Decision methods
Dataset: ToxicChat0124 · 100 previously inspected messages: 50 toxic, 50 benign. A small exploratory test, not live moderation.
- Read through a checkpoint
- Predict SAFE or BLOCK
- Stop, or continue to the next checkpoint
Where the learned checker stopped
On average, it used 15.9 of 24 layers, skipping 33.75% of transformer blocks. We verified that later blocks really did not execute.
| Same 100 messages | All 24 layers | Learned stopping |
|---|---|---|
| Correct | 77 | 79 |
| Toxic missed / 50 | 18 | 15 |
| Total request time | 63.6 s | 42.1 s |
It corrected three mistakes but introduced one new false block. This is promising exploration; it has not passed a fresh reliability test.
Three ways to decide when to stop
| Checker | Correct / 100 | Blocks skipped |
|---|---|---|
| Separate SAFE / BLOCK confidence limits | 78 | 17.25% |
| Confidence + agreement with the previous checkpoint | 79 | 16.50% |
| Learn when an answer may be wrong | 79 | 33.75% |
Only the learned-checker row above has this new measured runtime comparison. The other savings are block-count estimates.
What the checker learns, and what failed
Two logistic classifiers look at confidence, uncertainty and changes between checkpoints. They predict whether the current answer is wrong, and whether continuing would fix it. Full-model answers supply training targets only.
The same gate recipe reduced accuracy with a linear answer classifier: 77 to 75 correct. A separate REVIEW option answered only 64 messages, with 56 correct; the remaining 36 were unresolved.
Timing: one warmed, alternating paired pass on a CPU with four threads. Includes input preparation, classifier and checker; excludes model loading. Every exit trace matched the recorded policy.
All variants and training · Exact sample IDs · Runtime record · Adversarial review