Browser test results
Recorded browser comparisons
Saved BERT comparisons and the standalone tiny-classifier check.
These are saved runs of the BERT browser demo: 100 previously inspected ToxicChat messages, two reversed-order passes per comparison. Time includes input preparation and inference, but excludes loading and warmup. Each row belongs to its own paired run.
These are exploratory measurements on one computer, not fresh validation or a promise of the same speed on your device.
Fixed depth
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 1.82 s | 4.00 | 0 |
| Stop after layer 2 | 82 | 1.01 s | 2.00 | 5 |
Check before continuing
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 1.90 s | 4.00 | 0 |
| Check after layer 2 | 83 | 1.22 s | 2.56 | 0 |
Small model, then larger model
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 1.87 s | 4.00 | 0 |
| Tiny BERT, then larger BERT | 82 | 0.77 s | 2.84 | 3 |
The fallback runs two independent BERT models. Their blocks have different widths; the sum is not a compute-saving percentage. The 0.05/0.95 gate was not validated for this pairing.
Compare trained weights
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 1.98 s | 4.00 | 0 |
| Updated training | 81 | 1.88 s | 4.00 | 4 |
Process a queue together
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 2.09 s | 4.00 | 0 |
| Groups of four | 83 | 4.21 s | 4.00 | 0 |
Float32 versus INT8
| Path | Correct / 100 | Time / 100 | Mean blocks run | Changed answers |
|---|---|---|---|---|
| All 4 layers | 83 | 2.02 s | 4.00 | 0 |
| INT8 weights | 83 | 1.47 s | 4.00 | 0 |
The browser INT8 run matched these float32 decisions. Python INT8 differed on one message; this is not proof of lossless conversion across runtimes.
Standalone tiny classifier: separate export check
The two-layer classifier in the small-model demo got 82/100 correct, missed 13 toxic messages and falsely blocked 5 benign messages. The browser preserved all Python decisions. This is a separate single-pass functional check, not one of the paired comparisons above.
Both layers run. This later specialist uses a 0.4 BLOCK threshold; it is not the earlier 0.3-threshold model in the original Qwen cascade study.
Detailed evidence
Model, training and export checks · Every browser decision and timing