← Latvian, measured

Rendered copy of PREREGISTRATION.md. The original is committed at f280f3b (2026-09-28); the appended amendment section was added before any model call. Raw file: PREREGISTRATION.md.

Pre-registration: lv-llm-bench v0.1

Date: 2026-09-28. Author: skabene.

This file is committed on its own, before any benchmark item exists in the repo and before any model is called. The hash of that commit is the proof that the rules below were fixed first. Any change made after this commit goes into DEVIATIONS.md with a date and a reason, and this file is not edited.

What this is

A small benchmark of three Claude models on practical Latvian: noun and adjective declension, formal register, reading Latvian legal text, and grammatical agreement. All scoring is objective (exact match or a multiple-choice letter). No LLM judge is used.

It is a one-night, scaled-down version of a larger plan (about 340 items, 5 to 8 models, a second human rater, cost tracking). What was cut is listed at the end of this file.

Hypotheses

Stated before any run. Each one is published with a verdict (supported, not supported, or not testable) whatever the result.

IDHypothesisHow it is tested
H1Overall accuracy is ordered claude-opus-5-5 ≥ claude-sonnet-5 ≥ claude-haiku-4-5-20251001, and the Opus vs Haiku gap is a meaningful difference (definition below)Pooled accuracy over all items; paired comparison
H2Declension (T1) is the lowest-accuracy task for every modelPer-task accuracy, per model
H3Within T1, items tagged as irregular (alternation, exception, plurale_tantum) are answered correctly less often than items tagged regular, by at least 10pp, pooled across the three modelsPooled accuracy by tag group
H4Legal reading (T3) is saturated: every model scores ≥ 90%Per-task accuracy
H5Formal register recognition (T2) is ≥ 85% for every modelPer-task accuracy

Tasks and item counts

Target 40 items per task, 160 in total. Items are written after this commit and before any run. An item is dropped rather than kept if its gold answer is not certain; the floor is 30 items per task. Final counts are fixed in the items commit, which also comes before any run.

TaskID prefixWhat it testsFormatTarget items
T1 Declensiondecl-Give the form of a lemma for a given case and number (and, for adjective + noun phrases, gender and definite/indefinite ending). Mix of declensions 1 to 6, consonant alternation (brālis → brāļa), the exceptions (suns, ūdens, akmens, mēness, rudens, zibens, sāls), plural-only nouns, adjectives definite/indefiniteFree text, one word form or phrase40
T2 Registerreg-Pick the most appropriate formal rendering (official letter or customer service) among 4 options, or identify the option that breaks formal register (tu vs Jūs, slang, calques)Multiple choice A to D40
T3 Legal readinglaw-Short public excerpts of Latvian laws from likumi.lv (4 to 6 excerpts), questions answerable only from the excerptMultiple choice A to D40
T4 Agreementagr-A sentence with one agreement or case error; pick the corrected versionMultiple choice A to D40

Every item has: id, task, prompt, choices (multiple choice only), gold, rationale (one line). T1 items may have accept (other forms that count as correct, used only where standard Latvian allows two forms) and a tag. T3 items carry the source excerpt id. In multiple-choice tasks the position of the correct letter is balanced across A to D within each task (as close to equal as the item count allows).

Models

Model IDCalled as
claude-haiku-4-5-20251001claude -p --model claude-haiku-4-5-20251001
claude-sonnet-5claude -p --model claude-sonnet-5
claude-opus-5-5claude -p --model claude-opus-5-5

Models are called through the Claude Code CLI in headless mode (no API key is available for this run), with the system prompt replaced by --system-prompt, all tools disabled with --tools "", the prompt on stdin, and the working directory set to an empty temporary folder so no project context loads. This is not the same as a raw Messages API call; the CLI may add its own framing, and that is a limitation of the result.

Prompts (fixed)

The exact texts below are copied into prompts/ byte-for-byte. The item's prompt field is sent as the user message.

System prompt for T1 (prompts/system-declension.txt):

Tu esi latviešu valodas gramatikas eksperts. Atbildi tikai ar prasīto vārdformu vai vārdkopu mazajiem burtiem, bez paskaidrojumiem, bez pēdiņām un bez pieturzīmēm.

System prompt for T2, T3, T4 (prompts/system-mc.txt):

Tu esi latviešu valodas eksperts. Izlasi uzdevumu un atbildes variantus. Atbildi tikai ar pareizā varianta burtu (A, B, C vai D). Neraksti neko citu.

User message template for T1:

Vārds: {lemma}
Locījums: {case}
Skaitlis: {number}
Uzraksti šo formu.

For adjective + noun items the first line is Vārdkopa: {adjective} {noun} and one more line is added before the last: Galotne: noteiktā or Galotne: nenoteiktā. Case names are the Latvian names (nominatīvs, ģenitīvs, datīvs, akuzatīvs, instrumentālis, lokatīvs, vokatīvs); number is vienskaitlis or daudzskaitlis.

User message template for T2 and T4:

{question}

A) {choice A}
B) {choice B}
C) {choice C}
D) {choice D}

User message template for T3:

Likuma fragments:
"""
{excerpt}
"""

Jautājums: {question}
Atbildi, balstoties tikai uz šo fragmentu.

A) {choice A}
B) {choice B}
C) {choice C}
D) {choice D}

Scoring rules

Scoring is done by a script, never by hand, and never by a model.

Normalisation (both formats)

1. Unicode NFC. 2. Remove code fences (lines starting with three backticks), and the characters *, _, ` `, ", ', „, “, ”, «, »`. 3. Strip leading and trailing whitespace on every line; drop empty lines.

T1, free text

T2, T3, T4, multiple choice

Failures

Analysis and what counts as a meaningful difference

Item errors found after the run

Items are written by an LLM (see Limitations) and self-checked before the run; no native-speaker review happens before the run. After the run, every item where all three models give the same non-gold answer is re-checked against the source or a reference grammar. If the gold is wrong, the item is corrected or dropped, the change goes into DEVIATIONS.md, and both the original and the corrected scores are published. No other post-run changes to items are allowed.

What will be published regardless of the outcome

Limitations known in advance

Cut from the full plan, and why

CutWhy
About 340 items, 100 per taskOne night of work; items must be checked by hand
Open-weights models, TildeOpen, a local model, GPT, GeminiNo API access in this run; only the Claude CLI is available
Free-text register rewrite and VID summaries with an LLM judgeNeeds a judge and a human-agreement check; this version is objective-only
Translation taskStretch goal in the plan
Second human rater, kappaNo second rater available tonight
Cost per modelNot observable through the CLI; not estimated
Three runs per model, determinism checkTime; temperature cannot be set anyway
Contamination probe, canary stringsOut of scope for v0.1
Tēzaurs.lv-generated declension itemsNot downloaded tonight; items are hand-written instead

Amendments (before any model run)