Is the 27B model enough?
Qwen3.8-27B was sufficient to solve this ledger with tools. It wrote code, maintained totals and answered all 24 questions correctly in the file-using runs. That does not make every memory strategy equally reliable.
Which model did each approach use?
The same 27-billion-parameter model ran the large-window, summary, original CLM, revised CLM and file-enabled agents. It used FP8 weights on two Intel Arc Pro B70 GPUs working together, not a separate model on each card. The experiments changed instructions, storage permissions and the agent software, without task-specific training.
The model chose actions and wrote programs; CPU code did the ledger arithmetic. Summaries came from another call to the same model, not a smaller helper. How the model and tools work together.
What did 27B handle well?
Both file agents answered 24/24 on the 121K and 478K streams. Revised CLM also got 24/24 on the 478K stream, with code reading 74 of its 76 batches without putting their raw text into the model’s prompt.
This is useful evidence for structured work with clear rules. It does not establish the same reliability for a long legal argument, ambiguous documents or questions that require connecting many passages.
Would a smarter model help?
It could write more reliable code, choose better excerpts or need fewer attempts. Those are plausible benefits, not measured improvements here: there is no matched larger-model ledger comparison.
A larger model can need more GPU memory and time per token; fewer retries might offset that cost. Compare total time to a correct answer.
One revised run made three mistakes after rewriting its delete-handling code. A better model might avoid that bug; retaining and checking the script can address it too. Neither greater size nor training can recover records that the application permanently discarded.
Would fine-tuning help?
These successful runs needed no additional training. Specialized training could teach more consistent saving, retrieval and editing habits.
There is evidence on a different task: the CLM paper’s reinforcement-learning experiment raised Qwen3.5-9B’s BrowseComp-Plus accuracy from 28.8% to 42.5%. That was a different model and a deep-research benchmark, not this B70 ledger. Training experiment in the paper.
How should I choose?
For structured records, start by keeping originals, using exact code and checking its results. For tasks requiring interpretation, compare models on representative questions under the same storage rules. Judge correct completed answers and total time, not parameter count or token speed alone.