SkinDeepRESEARCHSteve Seguin

A label in one step

On easy decisions, skipping reasoning saved time; shortening an already short label added little.

The test used 80 easy yes/no, multiple-choice, sentiment and routing questions. A 24-question subset was also tested with thinking enabled.

Answer pathCorrectMedian time
One token, thinking off78 of 800.09 s
Ordinary label, thinking off78 of 800.09 s
Think, then label (subset)23 of 240.59 s

The one-step and ordinary decoding answers matched on 80/80 items. Both got 78/80 right. On the 24 items with thinking enabled, one-step and thinking answers matched on 24/24; both got 23/24 right.

Median response time was about 0.09 seconds for either non-thinking path, versus 0.59 seconds with thinking. The roughly 6.5× ratio describes these easy-item timing samples; it is not a speedup over ordinary non-thinking decoding.

The allowed labels began with distinct tokens. Restricting the first token ensures a valid label choice; it does not guarantee the choice is correct. All model layers still run. Hard decisions where reasoning changes the answer need a separate accuracy comparison.

Removing earlier reasoning is a different experiment · Related: trained label outputs

Original measurements and method

Qwen3.8-27B on two B70 GPUs, one pass. The one-token path restricted output to the allowed labels.

Every item and timing · Original protocol and result · Method note