01 · OCTOBER 9, 2026 · PILOT OBSERVATION
When a clear sentence gets an unclear answer
Our question was simple: as seven sentences gradually change meaning, do model judgments change in a sensible way?
On these 17 arcs, Ministral 3B met the strict 0/4 starting-point diagnostic on 8 arcs. Qwen 27B met it on 11. Stock Gemma 12B met it on 10, and stock Gemma 26B on 7. That already tells us “bigger always wins” is too simple.
The more striking pattern was clarity. Ministral’s starting-point clarity was below 4 in nine arcs. The three larger stock configurations returned starting-point clarity 4 throughout this batch. Some misses are about recognizing an ordinary meaning clearly, not about finding more concerning content.
The 4/4 endpoint was reached on 17/17 arcs by Qwen 27B and stock Gemma 26B, and 16/17 by Gemma 12B and Ministral 3B.
These are diagnostic counts from a small development corpus, not an IQ test or an isolated experiment on model size. The models differ in family and precision. The next useful step is to inspect counterexamples, then test frozen new families.
Inspect the result table ↗The question we keep
Can we predict which kinds of meaning a smaller model handles reliably, rather than treating its average score as a universal answer?