Semantic
Observatory
By Travis Somerville ↗

RESULTS / 17 ARCS · SAME TEST INPUTS

Compare behavior.
Keep the differences visible.

Five measured configurations. Endpoint, clarity and ordering diagnostics are shown separately; this pilot does not assign an overall intelligence score.

Loading model results…

What you can conclude

The smallest tested model has more ordinary-endpoint clarity failures in this batch. But bigger is not better on every measure. These comparisons vary model family, precision and sometimes training—not just parameter count.

The clarity-dip check is an exploratory audit flag. A flat curve can also expose a poorly chosen arc or an intentionally unambiguous progression. Inspect the language before treating the flag as a model error.

Read all definitions and limitations →