Exploratory results
Reuse instructions and group messages
Test two execution improvements together.
Live q4 Qwen with vocabulary scores and a 64-token message limit; it differs from the trained Python benchmark below.
Back to decisions without text
Compute the fixed instructions once, then process several messages together. This tests the two optimizations together on the same queued workload.
| Execution | Correct / 50 | Time for 50 |
|---|---|---|
| One message at a time | 45 | 31.49 s |
| Groups of four | 45 | 28.86 s |
| Reuse instructions | 45 | 20.48 s |
| Reuse instructions + groups of four | 45 | 17.99 s |
42.9% less time with identical decisions on these 50 messages. Every message still uses all 24 layers. This combines fewer repeated instruction calculations with fewer separate model calls.
What was measured?
Two counterbalanced CPU passes, four threads, on 50 previously inspected ToxicChat messages. Input preparation, sorting, padding, cache construction and copying are included. All eight runs matched labels and the specified logit tolerance. The cache contains only the fixed instructions and is copied privately for each batch.
The workload is already queued. This is not a live-response latency estimate, a new accuracy test, or a guarantee of identical outputs on other data.