# Translation generation comparison — September 24, 2026

Frozen source: translation-battles/2026-09-24-gold-v1.
65 generation attempts and 130 valid reviews; no new model calls made for this import.
Quizzes count as Articles. Five cases per model, Spanish and Japanese.
Quality is the source mean of ten 1–5 dimensions across ten reviews (two reviewers per case).
Speed = 60 × completed Articles / (attempted Articles × mean request seconds).
This retains failed/partial request time; it is a serial workload projection, not concurrency-4 wall throughput.
Value = completed Articles / catalog-estimated generation cost of all attempts.
Gemini Flash Lite has an accounting anomaly: cost is unknown and value is unranked.
Complete means generation finished, not publication-ready, structurally valid, or correct.
Review cost is excluded from generator value. Catalog estimates are not invoices.
Gold is frozen consensus reference text; generators receive English only. Reviewers are Sol and Opus at high effort; self-preference is possible.
Gold was preferred in 122/130 reviews; generated text in 6; 2 ties. Reviewers agreed on preference in 59/65 pairs.
Two trailing-comma JSON review parses were recovered without score edits. One Gemini raw-MDX normalization regression affects its reviewed artifact. See source report.

Source summary SHA256: `09af570aed797a9b981583f8f6b3a582801d7ac4e9ceccc39e6ced1a4c24f478`
