lv-llm-bench v0.1 · skabene · 2026-09-28

Latvian, measured — a small pre-registered LLM benchmark

160 synthetic Latvian items written for this benchmark (declension, formal register, reading a law, grammatical agreement; see Limitations), three Claude models, one run each, scored by a script against gold answers. No model judges another model.

Pre-registered in commit f280f3b on 2026-09-28. Items: 42953d3, pre-run verification: f972d97. Model calls: ….

Pre-registration here means the hypotheses, the prompts, the scoring rules and the rule for calling a difference "meaningful" were committed to git before any item was written and before any model was called. The commit hash is the timestamp. The results below are held to those rules, including the hypotheses that failed. Every later change is listed in DEVIATIONS. The full text is in PREREGISTRATION.

01Results

Pooled accuracy over all 160 items. Latency is the wall-clock time for one CLI call, including CLI start-up. Cost was not measured.

Bars show accuracy and whiskers show the 95% Wilson interval. At n = 40 an interval is roughly ±10pp wide, so overlapping whiskers mean the data do not separate the models.

02The pre-registered decision rule, applied

A difference counts as meaningful only if the 95% paired bootstrap interval excludes zero and the difference is at least 5 percentage points. Anything else is called tied. The bootstrap uses 10,000 resamples of items, seed 20260928, percentile method. The exact McNemar p is shown for reference but does not decide anything.

Pooled, all 160 items

Hypotheses (stated before the run)

Task labels from the pre-registration

TaskSaturated (all three ≥ 95%)Not discriminating (all within 5pp)

T1 by item tag

Items where all three models gave the same wrong answer. The pre-registration requires these to be re-checked: …

Per-task pairwise comparisons (exploratory; not a ranking)

03What surprised us / what didn't

What the pre-registered rule says

  • The three models are tied under the rule. The ordering was Opus 5.5 98.8% (158/160), Sonnet 5 96.9% (155/160) and Haiku 4.5 94.4% (151/160). No pair is a meaningful difference under the rule.
  • The closest call was Opus 5.5 vs Haiku 4.5. The gap is +4.4pp, and its 95% paired bootstrap interval [1.2, 8.1] excludes zero (McNemar p = 0.039). It is still under the pre-registered 5pp minimum, so it is reported as tied. Opus was right where Haiku was wrong on 8 items, and the reverse happened on 1. If the rule had been written after seeing the data, this would be the tempting place to move the line. The threshold stays where it was set, so H1 is not supported.
  • Opus 5.5 vs Sonnet 5 is +1.9pp [−0.6, 5.0], and Sonnet 5 vs Haiku 4.5 is +2.5pp [−1.9, 6.9]. Both are tied.

What didn't surprise us

  • Declension is the hard task (H2 supported). T1 is the lowest-scoring task for every model: Haiku 4.5 82.5%, Sonnet 5 92.5% and Opus 5.5 95.0%. It is also the only task not labelled saturated. Of the 16 wrong answers in the whole run (480 answers), 12 are T1 answers.
  • Legal reading and formal register are saturated (H4 and H5 supported). T2, T3 and T4 all meet the pre-registered "saturated" (all three ≥ 95%) and "not discriminating" labels. At this difficulty these three tasks cannot rank these three models. That is a limitation of the items as much as a result.
  • No formatting failures. All 480 calls succeeded on the first attempt, with 0 missing and 0 unparseable answers. Every model followed the "answer only" instruction, so the scores measure Latvian and not output format.

What surprised us

  • Irregular forms were less of a gap than expected (H3 not supported). Pooled across models, regular T1 items were right 26/27 times (96.3%) and irregular ones 61/69 times (88.4%). The gap is 7.9pp, under the 10pp threshold. Sonnet 5 and Opus 5.5 each got all 12 exception-noun items right. Of the 8 irregular misses, 6 were consonant alternation or plural-only nouns. With 9 regular items, this comparison is weak either way.
  • The same wrong forms came up across models. For decl-012 (sirds, gen. pl., gold siržu), both Haiku 4.5 and Opus 5.5 wrote sirdu, the form without the alternation. For decl-032 (brilles, gen. pl., gold briļļu), Sonnet 5 wrote brillu and Opus 5.5 wrote brilļu. No item had all three models agree on a non-gold answer, so the pre-registered post-run re-check was not triggered and no gold was changed. The gold forms themselves are still unreviewed by a native speaker.
  • Haiku 4.5 was the slowest per call in this run. Median latency was 10.15 s for Haiku 4.5 against 5.39 s for Sonnet 5 and 6.0 s for Opus 5.5. This is wall-clock time through the CLI, start-up included. The Haiku calls also ran first in the queue, so treat this as a property of this run, not of the model.

04Item explorer

All 160 items, each with its gold answer and every model's raw output, unedited. The marks show H = Haiku 4.5, S = Sonnet 5, O = Opus 5.5.

05Method

06Limitations

Rerun: python runner.py && python score.py && python build_site.py. Raw outputs are cached per model and item, so a rerun only fills in what is missing.