A label in one step
On easy decisions, skipping reasoning saved time; shortening an already short label added little.
The test used 80 easy yes/no, multiple-choice, sentiment and routing questions. A 24-question subset was also tested with thinking enabled.
| Answer path | Correct | Median time |
|---|---|---|
| One token, thinking off | 78 of 80 | 0.09 s |
| Ordinary label, thinking off | 78 of 80 | 0.09 s |
| Think, then label (subset) | 23 of 24 | 0.59 s |
The one-step and ordinary decoding answers matched on 80/80 items. Both got 78/80 right. On the 24 items with thinking enabled, one-step and thinking answers matched on 24/24; both got 23/24 right.
Median response time was about 0.09 seconds for either non-thinking path, versus 0.59 seconds with thinking. The roughly 6.5× ratio describes these easy-item timing samples; it is not a speedup over ordinary non-thinking decoding.
The allowed labels began with distinct tokens. Restricting the first token ensures a valid label choice; it does not guarantee the choice is correct. All model layers still run. Hard decisions where reasoning changes the answer need a separate accuracy comparison.
Removing earlier reasoning is a different experiment · Related: trained label outputs
Original measurements and method
Qwen3.8-27B on two B70 GPUs, one pass. The one-token path restricted output to the allowed labels.
Every item and timing · Original protocol and result · Method note