Semantic
Observatory
By Travis Somerville ↗

SOMERVILLE SEMANTIC DECISION BENCHMARK / RESEARCH PILOT / 09 OCT 2026

Which models follow the meaning?

Compare measured decisions. Inspect the sentences. Find the failure modes.

Loading measured results…

Latest results & model watch

Measured findings, separate from release news
MEASURED / STOCK MODELS

Qwen 27B keeps the order. Ministral 3B loses it more often.

Strength rises without reversal in 16 of 17 Qwen arcs, versus 12 of 17 for Ministral. Clarity dips appear in 15 versus 8. Same inputs; different configurations.

Inspect the side-by-side curves →
NEW MEASUREMENT / TYPESAFE JEV

Jev joins the comparison.

Testing the hosted Jev model on the same 119 inputs. Its API adapter is recorded separately from local native-tree scoring.

See results and provenance →
MODEL WATCH / NOT YET TESTED HERE

Qwen3.8-Flash-Next: a candidate for the next run.

The open-weight architecture preview expands the model-watch list. We haven’t measured its arc behavior, so it has no leaderboard score.

Read the official model card ↗

OPEN THE EVIDENCE

Every score has words behind it.

The strength × clarity matrix shows where models placed actual sentences. Select a cell, compare the readings, then inspect the full seven-position arc.

Explore the matrix ↗

Preview: four stock local configurations.