Rendered copy of PREREGISTRATION.md. The original is committed at f280f3b (2026-09-28); the appended amendment section was added before any model call. Raw file: PREREGISTRATION.md.
Pre-registration: lv-llm-bench v0.1
Date: 2026-09-28. Author: skabene.
This file is committed on its own, before any benchmark item exists in the repo and before any model is called. The hash of that commit is the proof that the rules below were fixed first. Any change made after this commit goes into DEVIATIONS.md with a date and a reason, and this file is not edited.
What this is
A small benchmark of three Claude models on practical Latvian: noun and adjective declension, formal register, reading Latvian legal text, and grammatical agreement. All scoring is objective (exact match or a multiple-choice letter). No LLM judge is used.
It is a one-night, scaled-down version of a larger plan (about 340 items, 5 to 8 models, a second human rater, cost tracking). What was cut is listed at the end of this file.
Hypotheses
Stated before any run. Each one is published with a verdict (supported, not supported, or not testable) whatever the result.
| ID | Hypothesis | How it is tested |
|---|---|---|
| H1 | Overall accuracy is ordered claude-opus-5-5 ≥ claude-sonnet-5 ≥ claude-haiku-4-5-20251001, and the Opus vs Haiku gap is a meaningful difference (definition below) | Pooled accuracy over all items; paired comparison |
| H2 | Declension (T1) is the lowest-accuracy task for every model | Per-task accuracy, per model |
| H3 | Within T1, items tagged as irregular (alternation, exception, plurale_tantum) are answered correctly less often than items tagged regular, by at least 10pp, pooled across the three models | Pooled accuracy by tag group |
| H4 | Legal reading (T3) is saturated: every model scores ≥ 90% | Per-task accuracy |
| H5 | Formal register recognition (T2) is ≥ 85% for every model | Per-task accuracy |
Tasks and item counts
Target 40 items per task, 160 in total. Items are written after this commit and before any run. An item is dropped rather than kept if its gold answer is not certain; the floor is 30 items per task. Final counts are fixed in the items commit, which also comes before any run.
| Task | ID prefix | What it tests | Format | Target items |
|---|---|---|---|---|
| T1 Declension | decl- | Give the form of a lemma for a given case and number (and, for adjective + noun phrases, gender and definite/indefinite ending). Mix of declensions 1 to 6, consonant alternation (brālis → brāļa), the exceptions (suns, ūdens, akmens, mēness, rudens, zibens, sāls), plural-only nouns, adjectives definite/indefinite | Free text, one word form or phrase | 40 |
| T2 Register | reg- | Pick the most appropriate formal rendering (official letter or customer service) among 4 options, or identify the option that breaks formal register (tu vs Jūs, slang, calques) | Multiple choice A to D | 40 |
| T3 Legal reading | law- | Short public excerpts of Latvian laws from likumi.lv (4 to 6 excerpts), questions answerable only from the excerpt | Multiple choice A to D | 40 |
| T4 Agreement | agr- | A sentence with one agreement or case error; pick the corrected version | Multiple choice A to D | 40 |
Every item has: id, task, prompt, choices (multiple choice only), gold, rationale (one line). T1 items may have accept (other forms that count as correct, used only where standard Latvian allows two forms) and a tag. T3 items carry the source excerpt id. In multiple-choice tasks the position of the correct letter is balanced across A to D within each task (as close to equal as the item count allows).
Models
| Model ID | Called as |
|---|---|
| claude-haiku-4-5-20251001 | claude -p --model claude-haiku-4-5-20251001 |
| claude-sonnet-5 | claude -p --model claude-sonnet-5 |
| claude-opus-5-5 | claude -p --model claude-opus-5-5 |
Models are called through the Claude Code CLI in headless mode (no API key is available for this run), with the system prompt replaced by --system-prompt, all tools disabled with --tools "", the prompt on stdin, and the working directory set to an empty temporary folder so no project context loads. This is not the same as a raw Messages API call; the CLI may add its own framing, and that is a limitation of the result.
- One run per item per model. No retries on a wrong answer; retries (up to 2) only on a CLI failure or timeout (120 s).
- Temperature, top-p and seed cannot be set through the CLI. Outputs may not be deterministic, and determinism is not measured in this version.
- Latency (wall-clock per call, including CLI start-up) is recorded. Token counts and cost are not measured and are not estimated.
Prompts (fixed)
The exact texts below are copied into prompts/ byte-for-byte. The item's prompt field is sent as the user message.
System prompt for T1 (prompts/system-declension.txt):
Tu esi latviešu valodas gramatikas eksperts. Atbildi tikai ar prasīto vārdformu vai vārdkopu mazajiem burtiem, bez paskaidrojumiem, bez pēdiņām un bez pieturzīmēm.
System prompt for T2, T3, T4 (prompts/system-mc.txt):
Tu esi latviešu valodas eksperts. Izlasi uzdevumu un atbildes variantus. Atbildi tikai ar pareizā varianta burtu (A, B, C vai D). Neraksti neko citu.
User message template for T1:
Vārds: {lemma}
Locījums: {case}
Skaitlis: {number}
Uzraksti šo formu.
For adjective + noun items the first line is Vārdkopa: {adjective} {noun} and one more line is added before the last: Galotne: noteiktā or Galotne: nenoteiktā. Case names are the Latvian names (nominatīvs, ģenitīvs, datīvs, akuzatīvs, instrumentālis, lokatīvs, vokatīvs); number is vienskaitlis or daudzskaitlis.
User message template for T2 and T4:
{question}
A) {choice A}
B) {choice B}
C) {choice C}
D) {choice D}
User message template for T3:
Likuma fragments:
"""
{excerpt}
"""
Jautājums: {question}
Atbildi, balstoties tikai uz šo fragmentu.
A) {choice A}
B) {choice B}
C) {choice C}
D) {choice D}
Scoring rules
Scoring is done by a script, never by hand, and never by a model.
Normalisation (both formats)
1. Unicode NFC. 2. Remove code fences (lines starting with three backticks), and the characters *, _, ` `, ", ', „, “, ”, «, »`. 3. Strip leading and trailing whitespace on every line; drop empty lines.
T1, free text
- Take the last non-empty line after normalisation. Remove a leading
Atbilde:(case-insensitive) if present. Remove trailing.,,,!,;,:. Collapse internal whitespace to one space. Lowercase. - Correct if the result equals
goldor any string inaccept, compared after the same lowercase and NFC. Diacritics are significant (bralais wrong forbrāļa). - Empty output is wrong.
T2, T3, T4, multiple choice
- After normalisation, join lines with a space. The answer is the first match of the regular expression
(?<![A-Za-zĀ-ž])([ABCD])(?![A-Za-zĀ-ž])(a standalone capital A to D, not part of a word). - If there is no match the item is unparseable and scored wrong.
- Correct if the letter equals
gold.
Failures
- A call that still fails after 2 retries is recorded as
missing. In the primary table a missing answer counts as wrong. The number of missing and unparseable answers is reported per model.
Analysis and what counts as a meaningful difference
- Accuracy per model per task, with a 95% Wilson interval.
- Pooled accuracy per model over all items, with a 95% Wilson interval.
- Paired comparison for each pair of models on the pooled items: the difference in accuracy, a 95% paired bootstrap interval (10,000 resamples of items, seed 20260928, percentile method), and the exact McNemar p-value.
- A difference is called meaningful only if the 95% paired bootstrap interval excludes zero AND the difference is at least 5 percentage points. Anything else is reported as "tied", including differences that are statistically significant but under 5pp.
- Per-task comparisons use the same rule but are labelled exploratory: at n = 40 the standard error of one accuracy can be up to about 8pp, so the per-task tables are not a ranking.
- A task where all three models score ≥ 95% is reported as "saturated". A task where all three models are within 5pp of each other is reported as "not discriminating" among these three models.
Item errors found after the run
Items are written by an LLM (see Limitations) and self-checked before the run; no native-speaker review happens before the run. After the run, every item where all three models give the same non-gold answer is re-checked against the source or a reference grammar. If the gold is wrong, the item is corrected or dropped, the change goes into DEVIATIONS.md, and both the original and the corrected scores are published. No other post-run changes to items are allowed.
What will be published regardless of the outcome
- This file, its commit hash, and
DEVIATIONS.md. - All items with gold answers and rationales.
- Every raw model output, unedited, with its latency.
- Per-item scores, per-task and pooled accuracy tables, the paired comparisons, and a verdict for each of H1 to H5, including the ones that fail.
- The counts of missing and unparseable answers.
Limitations known in advance
- Items are written by an LLM (Claude Opus 5.5, in the same session that builds the harness), and Claude Opus 5.5 is one of the models tested. Items may be easier or harder for their own author's family in ways I can't measure here. All three models are from one vendor, so the benchmark cannot say anything about other vendors.
- No human review before the run. Gold answers are checked only by the authoring model against standard Latvian paradigms and, for T3, the excerpt text. A native-speaker review is planned after the run; any gold it changes goes into
DEVIATIONS.mdwith both scores published. There is no inter-rater agreement number. - T1 gold forms are written from the standard paradigms, not machine-extracted from Tēzaurs.lv.
- 40 items per task is small. Only large differences can be detected.
- One run per item, uncontrolled temperature: a rerun could change some answers.
- Calls go through the Claude Code CLI, not the raw API.
- Legal excerpts come from likumi.lv consolidated texts, which the portal marks as informative. Their reuse terms are not yet verified.
- Contamination is not tested. Once published, these items can end up in training data.
Cut from the full plan, and why
| Cut | Why |
|---|---|
| About 340 items, 100 per task | One night of work; items must be checked by hand |
| Open-weights models, TildeOpen, a local model, GPT, Gemini | No API access in this run; only the Claude CLI is available |
| Free-text register rewrite and VID summaries with an LLM judge | Needs a judge and a human-agreement check; this version is objective-only |
| Translation task | Stretch goal in the plan |
| Second human rater, kappa | No second rater available tonight |
| Cost per model | Not observable through the CLI; not estimated |
| Three runs per model, determinism check | Time; temperature cannot be set anyway |
| Contamination probe, canary strings | Out of scope for v0.1 |
| Tēzaurs.lv-generated declension items | Not downloaded tonight; items are hand-written instead |
Amendments (before any model run)
- 2026-09-28: Step-2 adversarial verification of all 160 items, done before any model call. 8 items had wording fixes, 0 were dropped, and no gold answer or gold letter changed. Per-item reasons are in
items/verification.mdand the fixes are applied bytools/apply_verification.py. The counts stay 40/40/40/40.