# Model release comparison assets

These assets accompany **Three Exclusive AI Benchmarks: Quality, Speed, and Value**.

## September 24: current translation-generation charts

The translation suite now measures generation: 13 models, five Articles (including a quiz), 65 attempts, 130 reviews. See `translation-generation-method.md` for denominators, source digest and limitations, and `translation-generation-source-report.md` for source findings. Frozen summary and cost receipts are included; raw responses are not republished.

Run `render-translation-generation.py`, then `render-efficiency.py` and `render-banner-data.py`. Build Astro to render the current ECharts SVGs. Translation quality uses a full /5 scale. Unknown Gemini cost is omitted from the value chart and shown as Unknown in the leaderboard. The other suites are unchanged.

Figures 09 and 11 and the 24k judging files below are historical, not part of the current article. Banner 12 is superseded by live HTML; screenshots are in docs/review/model-release.


## September 23 refresh (archived judging charts)

The live banner and vertical charts now use 13 extraction models, 13 models
from the matched 24k translation regression cohort, and 9 completed security
panels. The historical 8k reference-coverage and 16k lowest10 figures remain
separate, with their original data and methods. Historical generated banners
08 and 10 are superseded; 12 is the current browser-rendered banner.

- Extraction: `source.md` is the updated supplied comparison. `render.py`
  joins by model/run identity rather than table position; it includes Fable
  and Opus and preserves the pooled Luna denominator.
- Translation: `translation-24k-source-data.json` holds the 234 matched
  first attempts, fixture hash, source digest, and per-attempt accounting
  derived with the source repository's catalogCost/usageFromResult helpers.
  No raw model text or reasoning is republished. `render-translation-24k.py`
  validates the ten-item report's rounded costs and latencies, then exports
  the ten-Article rates (including quizzes). Failed-call costs are retained; speed uses
  parsed-call latency. DeepSeek's dropped-suggestion case remains flagged.
  This is regression on reused controls, not held-out judge accuracy.
- Security: `judge-followup-data.json` adds Fable and the strict-schema Opus
  run. Its cumulative study cost is $1.521706. `render-judge.py` includes nine
  completed panels, preserves incomplete runs separately, and marks the
  Opus output-contract mismatch in captions and data. No new calls occurred
  after the JSON-schema default change. The exact Opus median is 3.635 s.

Rebuild archived assets offline in this order: `render.py`, `render-judge.py`,
`render-translation-24k.py`, `render-efficiency.py`, `render-banner-data.py`.
Use `MPLCONFIGDIR=/tmp/model-chart-mpl` and Python with matplotlib, fonttools,
and brotli installed. `render-reference.py` intentionally reads the frozen
`historical-efficiency-original.csv`; never mix the new 24k rates into the
older three-language annotations. Shared chart data drives both article
versions. No model evaluation or calibration is part of rendering.
