← Latvian, measured

Rendered copy of DEVIATIONS.md. Raw file: DEVIATIONS.md.

Deviations from PREREGISTRATION.md

Every change after the pre-registration commit, dated, with the reason. PREREGISTRATION.md itself is not edited.

DateWhatWhy
2026-09-28Clarification, not a change: two sentences in PREREGISTRATION.md overlap. "Item errors found after the run" allows only the all-models-agree re-check; "Limitations" also allows gold changes from a later native-speaker review. Both routes are kept, and any change from either is logged here with original and corrected scores published side by side.Remove ambiguity before any run
2026-09-28Final item counts: 40 / 40 / 40 / 40 (160). No task fell below the 30-item floor. T1 tags: 9 regular, 7 alternation, 12 exception, 4 plurale_tantum, 8 adjective. T3 uses 5 excerpts from 3 acts (Darba likums 46., 131., 149. pants; Patērētāju tiesību aizsardzības likums 30. pants; Dzīvojamo māju pārvaldīšanas likums 13. pants).Items written and committed before any model call, as pre-registered
2026-09-28H3 is tested on 9 regular vs 23 irregular T1 items per model (27 vs 69 pooled answers). This is small; the H3 verdict will be reported with that caveat.Item counts fixed by what could be written with certain gold in one night
2026-09-28Pre-run verification: wording of 8 items was fixed (reg-037, law-035, law-036, law-037, agr-001, agr-011, agr-027, agr-029). 0 items were dropped and no gold changed. PREREGISTRATION.md gained an appended "Amendments (before any model run)" section; its original sections are unchanged.Remove ungrammatical question wording, T4 sentences that had two errors where the prompt says one, and two distractors that could be argued correct. Details in items/verification.md
2026-09-28Run notes, no rule changed. (1) The first model calls were a 2-item smoke test (decl-001, decl-002 on all three models). Those outputs are the single pre-registered run of the two items and were not repeated. (2) Implementation details the pre-registration left open: the bootstrap percentile bounds are sorted[250] and sorted[9749] of the 10,000 resampled differences; p90 latency uses linear interpolation; "T1 lowest" (H2) allows ties. (3) All 480 calls succeeded on the first attempt, so no retries were used.Record choices made while writing score.py