Early tests
Decisions without a written reply
Sometimes an app only needs a label, such as “support” or “spam”.
Run Qwen 0.5B on real messages and compare accuracy and time. All 24 layers run; this demo does not use the trained banking classifier.
An app may need a category or a yes/no decision. A classifier can return that result directly, without writing a sentence or JSON.
This changes the output. The model can still use every layer; stopping early is a separate optimization.
Run 100 real moderation messages in your browser · Watch recorded movement decisions
Try it on your device
The browser benchmark runs Qwen 0.5B on real ToxicChat messages. Compare a direct label decision with a written JSON answer, including mistakes and the time to finish. Both use all 24 layers.
Where this could be useful
When an app needs a category, score or action, a small output head can return it directly. Some cases have narrow prototypes; others are proposals.
- Live chat moderationReturn an allow, block or review decision.
- Route customer requestsSend a request to the right queue.
- Choose a receipt totalReturn the right amount from a short list.
- ECG labels without a written reportReturn structured predictions from a heart trace.
More use cases: search, actions, signals and image checks
Messages and documents
- Find text to redactMark spans instead of writing a new document.
- Rank search resultsPut useful matches near the top.
- Sort an inboxSeparate receipts, requests and newsletters.
Actions
- Choose a game moveReturn a direction on each turn.
- Commands on a small deviceUse a short path for familiar commands.
Medical signals and scans
- Check ECG recording qualityFlag an unusable trace for another recording.
- Mark regions in an MRIReturn a mask that a reviewer can inspect.
- EEG sleep-stage labelsClassify recorded sleep segments directly.
Language and image generation
- Stop an unsafe image before it finishesCheck during generation and cancel early.
What the research has shown
A separately trained Qwen classifier got 2,551 of 3,080 banking queries right across 77 topics. A word-based classifier got 2,464 right. The browser demo uses Qwen’s existing output scores; it does not reproduce those trained classifiers.
A separate changing-rule test failed. Following arbitrary instructions remains unproven, and shorter output does not guarantee equal accuracy.
Banking data and results · Matched one-token comparison · Changing-rule test · Other ways to reduce computation
Five more practical datasets to test
Processing several messages together
Reusing the fixed instructions: measured results
Try a specific task
Classify, or stop after one token?
A trained classifier returns scores directly. A one-token reply still uses the language output layer. See what each saves and which mistakes changed.
Short labels on a larger model
One-step labels matched ordinary decoding on 80 easy items. On a 24-item subset, skipping reasoning kept the same answers; median response time was about 0.09 seconds without reasoning versus 0.59 with it.