New Models, 3 New Evals: OCR, Translation, and Security Verification Tests
GPT-6, Claude Opus 5.5 and Fable 5.1, Gemini, GLM 5.3, DeepSeek V4.1 and Qwen 3.8 on the handwriting, translation, and security-verifier tests behind 17th Street Labs’ own software.
Same eleven page scans. GPT-6 Astra Pro cost $2.71 to read them;
DeepSeek V4.1 Flash cost $0.014. For roughly 197 times the price, Astra
Pro got one more field right—out of 132. That is the kind of difference a
release-day leaderboard cannot tell you about your own workload.
17th Street Labs tests new models against software we build and maintain:
DocSlurp’s handwriting extraction, the translation engine behind Dan Levy’s
blog, and the Agent Security work behind Exploit Hunter. These tests ask models
to read messy inputs, preserve meaning, and verify evidence. The samples are
small—eleven documents, five articles, six security controls—so we say where
results clearly separate and call the rest ties.
Fable read the most handwriting fields correctly. Opus and Astra tied at the
top of translation, and Astra had the most publish-ready runs. Sol was the
cheapest and fastest security verifier to pass every control. Our picks for
each workload are at the end.
17th Street Labs · Three custom benchmarks
Choose the task. Compare the models.
Handwriting Recognition: model comparison
Quality
Field accuracy (%) · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View quality data
Handwriting Recognition: Field accuracy (%)
Model
Field accuracy (%)
Claude Fable 5.1
96.2%
Gemini 3.5 Flash Lite
92.4%
Claude Opus 5.5
92.4%
GLM 5.3 Flash
90.2%
GLM 5.3 FlashX
90.2%
Gemini 3.1 Flash Lite
89.4%
Gemini 3.8 Flash
89.4%
GPT-6 Astra
89.4%
GPT-6 Sol
88.6%
GPT-6 Astra Pro
88.6%
DeepSeek V4.1 Flash
87.9%
GPT-5.6 Luna
83.3%
GPT-6 Luna
66.9%
Zero baseline · Full score range
Speed
Page Scans per minute · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View speed data
Handwriting Recognition: Page Scans per minute
Model
Page Scans per minute
Gemini 3.5 Flash Lite
75.9
Gemini 3.1 Flash Lite
72.5
GLM 5.3 Flash
64.7
DeepSeek V4.1 Flash
62.9
GLM 5.3 FlashX
44.3
GPT-6 Sol
40.2
Gemini 3.8 Flash
39.8
GPT-6 Luna
27.8
GPT-5.6 Luna
24.0
Claude Opus 5.5
23.8
Claude Fable 5.1
22.6
GPT-6 Astra
21.3
GPT-6 Astra Pro
10.9
Zero baseline · Scaled to the highest rate
Value
Page Scans per $1 · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View value data
Handwriting Recognition: Page Scans per $1
Model
Page Scans per $1
DeepSeek V4.1 Flash
802
Gemini 3.1 Flash Lite
716
GLM 5.3 Flash
704
GPT-6 Luna
610
Gemini 3.5 Flash Lite
533
GPT-5.6 Luna
366
Gemini 3.8 Flash
323
GLM 5.3 FlashX
291
GPT-6 Sol
32.7
Claude Opus 5.5
29.5
Claude Fable 5.1
12.8
GPT-6 Astra
6.6
GPT-6 Astra Pro
4.1
Zero baseline · Scaled to the highest rate
Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.
11 documents, including a typed control. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.
Language Translation: model comparison
Quality
Translation quality /5 · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View quality data
Language Translation: Translation quality /5
Model
Translation quality /5
Claude Opus 5.5
4.37/5
GPT-6 Astra
4.32/5
GPT-6 Sol
4.06/5
GPT-5.6 Luna
3.95/5
Claude Fable 5.1
3.95/5
GPT-6 Luna
3.88/5
Gemini 3.8 Flash
3.81/5
GLM 5.3 Flash
3.80/5
DeepSeek V4.1 Flash
3.72/5
GLM 5.3 FlashX
3.67/5
Qwen 3.8 27B
3.28/5
Qwen 3.8 Flash
3.18/5
Zero baseline · Full score range
Speed
Articles per minute · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View speed data
Language Translation: Articles per minute
Model
Articles per minute
Qwen 3.8 27B
7.3
Gemini 3.8 Flash
6.1
DeepSeek V4.1 Flash
5.4
GLM 5.3 FlashX
5.1
GPT-5.6 Luna
4.5
GLM 5.3 Flash
4.1
GPT-6 Sol
4.1
GPT-6 Luna
3.9
Claude Opus 5.5
3.2
Qwen 3.8 Flash
2.7
GPT-6 Astra
2.5
Claude Fable 5.1
2.0
Zero baseline · Scaled to the highest rate
Value
Articles per $1 · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View value data
Language Translation: Articles per $1
Model
Articles per $1
DeepSeek V4.1 Flash
1,109
Qwen 3.8 Flash
987
GPT-6 Luna
957
GLM 5.3 Flash
872
GPT-5.6 Luna
419
GLM 5.3 FlashX
359
Qwen 3.8 27B
180
Gemini 3.8 Flash
129
GPT-6 Sol
47.8
Claude Opus 5.5
16.2
GPT-6 Astra
9.4
Claude Fable 5.1
6.5
Zero baseline · Scaled to the highest rate
Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.
Five Articles, including a quiz; Spanish and Japanese. Two model reviewers, 10 reviews per model. Speed counts completed Articles over all request time. Gemini 3.5 Flash Lite is excluded: a harness normalization regression affected its reviewed output and its cost is unknown; a rerun is pending. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.
Security Task Verification: model comparison
Quality
Controls passed · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View quality data
Security Task Verification: Controls passed
Model
Controls passed
GPT-6 Sol
18/18
Claude Opus 5.5 *
18/18
GPT-6 Astra
18/18
Claude Fable 5.1
18/18
DeepSeek V4.1 Flash
15/18
Gemini 3.8 Flash
15/18
GPT-6 Luna
13/18
GLM 5.3 Flash
12/18
Gemini 3.5 Flash Lite
12/18
Zero baseline · Full score range
Speed
Traces per minute · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View speed data
Security Task Verification: Traces per minute
Model
Traces per minute
Gemini 3.5 Flash Lite
49.3
DeepSeek V4.1 Flash
40.4
GLM 5.3 Flash
39.0
GPT-6 Luna
27.0
GPT-6 Sol
20.5
Gemini 3.8 Flash
17.6
Claude Opus 5.5 *
16.5
GPT-6 Astra
9.5
Claude Fable 5.1
8.4
Zero baseline · Scaled to the highest rate
Value
Traces per $1 · Higher is better
Scroll horizontally for all models. Focus the chart to use arrow keys.
View value data
Security Task Verification: Traces per $1
Model
Traces per $1
GPT-6 Luna
5,162
GLM 5.3 Flash
3,955
DeepSeek V4.1 Flash
3,660
Gemini 3.5 Flash Lite
1,044
Gemini 3.8 Flash
593
GPT-6 Sol
268
Claude Opus 5.5 *
54.9
GPT-6 Astra
44.6
Claude Fable 5.1
29.1
Zero baseline · Scaled to the highest rate
Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.
6 synthetic controls repeated 3 times. * Opus used strict JSON Schema; all other completed runs used JSON-object mode. Output settings are not matched. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.
Explore the sortable shortlists and all model results
Featured model families
gpt-6
opus-5.5
glm-5.3
deepseek-v4.1-flash
gemini-3.8-flash
qwen-3.8
fable-5.1
Three test suites
The shortlists
Choose quality, speed, or cost. Expand each list to see all models.
Relative score: LowerMiddleHigherThirds of each test’s observed range, not pass/fail grades.
Handwriting Recognition
Quality: field accuracy
Highest quality first. Ties: value, then speed.
1. Claude Fable 5.196.2%22.6/min12.8/$1
Quality96.2%Speed22.6/minValue12.8/$1
2. Gemini 3.5 Flash Lite92.4%75.9/min533/$1
Quality92.4%Speed75.9/minValue533/$1
3. Claude Opus 5.592.4%23.8/min29.5/$1
Quality92.4%Speed23.8/minValue29.5/$1
Show all 13 models
4. GLM 5.3 Flash90.2%64.7/min704/$1
Quality90.2%Speed64.7/minValue704/$1
5. GLM 5.3 FlashX90.2%44.3/min291/$1
Quality90.2%Speed44.3/minValue291/$1
6. Gemini 3.1 Flash Lite89.4%72.5/min716/$1
Quality89.4%Speed72.5/minValue716/$1
7. Gemini 3.8 Flash89.4%39.8/min323/$1
Quality89.4%Speed39.8/minValue323/$1
8. GPT-6 Astra89.4%21.3/min6.6/$1
Quality89.4%Speed21.3/minValue6.6/$1
9. GPT-6 Sol88.6%40.2/min32.7/$1
Quality88.6%Speed40.2/minValue32.7/$1
10. GPT-6 Astra Pro88.6%10.9/min4.1/$1
Quality88.6%Speed10.9/minValue4.1/$1
11. DeepSeek V4.1 Flash87.9%62.9/min802/$1
Quality87.9%Speed62.9/minValue802/$1
12. GPT-5.6 Luna83.3%24.0/min366/$1
Quality83.3%Speed24.0/minValue366/$1
13. GPT-6 Luna66.9%27.8/min610/$1
Quality66.9%Speed27.8/minValue610/$1
Reference contributors (not ranked)
These models helped build the issue reference. Their quality scores are not independent comparisons.
Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.
11 documents, including a typed control.
Language Translation
Quality: translation quality /5
Highest quality first. Ties: value, then speed.
1. Claude Opus 5.54.37/53.2/min16.2/$1
Quality4.37/5Speed3.2/minValue16.2/$1
2. GPT-6 Astra4.32/52.5/min9.4/$1
Quality4.32/5Speed2.5/minValue9.4/$1
3. GPT-6 Sol4.06/54.1/min47.8/$1
Quality4.06/5Speed4.1/minValue47.8/$1
Show all 12 models
4. GPT-5.6 Luna3.95/54.5/min419/$1
Quality3.95/5Speed4.5/minValue419/$1
5. Claude Fable 5.13.95/52.0/min6.5/$1
Quality3.95/5Speed2.0/minValue6.5/$1
6. GPT-6 Luna3.88/53.9/min957/$1
Quality3.88/5Speed3.9/minValue957/$1
7. Gemini 3.8 Flash3.81/56.1/min129/$1
Quality3.81/5Speed6.1/minValue129/$1
8. GLM 5.3 Flash3.80/54.1/min872/$1
Quality3.80/5Speed4.1/minValue872/$1
9. DeepSeek V4.1 Flash3.72/55.4/min1,109/$1
Quality3.72/5Speed5.4/minValue1,109/$1
10. GLM 5.3 FlashX3.67/55.1/min359/$1
Quality3.67/5Speed5.1/minValue359/$1
11. Qwen 3.8 27B3.28/57.3/min180/$1
Quality3.28/5Speed7.3/minValue180/$1
12. Qwen 3.8 Flash3.18/52.7/min987/$1
Quality3.18/5Speed2.7/minValue987/$1
Reference contributors (not ranked)
These models helped build the issue reference. Their quality scores are not independent comparisons.
Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.
Five Articles, including a quiz; Spanish and Japanese. Two model reviewers, 10 reviews per model. Speed counts completed Articles over all request time. Gemini 3.5 Flash Lite is excluded: a harness normalization regression affected its reviewed output and its cost is unknown; a rerun is pending.
These models helped build the issue reference. Their quality scores are not independent comparisons.
Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.
6 synthetic controls repeated 3 times. * Opus used strict JSON Schema; all other completed runs used JSON-object mode. Output settings are not matched.
Quality measures differ by task. Translation quality is model-assessed on five Articles, with two reviews per output. Rates count completed generations; unknown cost is unranked. Security includes an Opus strict-schema follow-up with different output settings.
Handwriting OCR: what the extra accuracy costs
Task
Phone photographs → structured fields. Direct vision extraction, before retrieval or the rest of the ingestion pipeline.
Measures
Field accuracy and exact-document rate: how often every required field is correct, scored against expected values.
Sample
Thirteen models, eleven independent documents including a typed control. A privately held-out corpus for DocSlurp, growing with iPhone and Android photographs.
Fable 5.1 extracted 96.2% of fields correctly and returned eight of eleven
page scans fully correct. Gemini 3.5 Flash Lite and Opus 5.5 tied at 92.4%
field accuracy, but Gemini returned seven exact documents; Opus returned five.
GLM Flash and FlashX tied at 90.2% of fields, with four and six exact documents
respectively.
With eleven documents, a single page moves the exact-document rate by nine
points. Fable’s lead over Gemini is one page and five fields. Treat the top
three as a shortlist, not a final order.
That distinction matters to an ingestion system. A mostly correct page can
still send the wrong date, amount, or name into a knowledge base. DocSlurp,
our document-ingestion and RAG project, uses this benchmark to measure that
first step. We plan to open source the project while keeping evaluation
material held out to reduce exposure of the answers.
Circles show field accuracy; diamonds show exact-document rate.
Model
Field accuracy
Exact documents
Page Scans/min
Page Scans/$1
Claude Fable 5.1
96.2%
72.7%
22.6
12.8
Gemini 3.5 Flash Lite
92.4%
63.6%
75.9
533
Claude Opus 5.5
92.4%
45.5%
23.8
29.5
GLM 5.3 Flash
90.2%
36.4%
64.7
704
GLM 5.3 FlashX
90.2%
54.5%
44.3
291
Gemini 3.1 Flash Lite
89.4%
54.5%
72.5
716
GPT-6 Astra
89.4%
45.5%
21.3
6.6
Gemini 3.8 Flash
89.4%
45.5%
39.8
323
GPT-6 Astra Pro
88.6%
45.5%
10.9
4.1
GPT-6 Sol
88.6%
54.5%
40.2
32.7
DeepSeek V4.1 Flash
87.9%
36.4%
62.9
802
GPT-5.6 Luna
83.3%
36.4%
24.0
366
GPT-6 Luna
66.9%
6.1%
27.8
610
Fable’s eleven-scan run cost $0.86, about 42 times Gemini Flash
Lite’s $0.021, for one additional exact document. Opus cost $0.37—about
18 times Gemini’s bill—with the same field accuracy and two fewer exact pages.
Whether that extra page is worth $0.84 depends on what a wrong page costs.
If every error means a person rebuilds the record by hand, Fable pays for
itself quickly. If downstream validation catches most mistakes, Gemini is the
better buy.
Gemini Flash Lite was also the fastest, at 75.9 Page Scans/minute. DeepSeek
delivered the most scans per dollar, 802, at 87.9% field accuracy, which
makes it a strong low-cost baseline for a larger document test.
i18n translation: quality gains and inference costs
Task
English Articles → Spanish or Japanese MDX, including a quiz. Preserve meaning, technical claims, code, links, and working components.
Measures
Ten quality dimensions, including faithfulness, technical accuracy, readability, authorial voice, and MDX integrity. Two frontier reviewers assess each output against a frozen reference.
Sample
Twelve ranked models, five source articles each (one is a quiz), two reviews per translation.
Opus 5.5 (4.37/5) and Astra (4.32) are effectively tied at the top; Sol
follows at 4.06. The two judges disagree on the order: Sol ranked Opus first
(4.42 to 4.28), and Opus ranked Astra first (4.36 to 4.32). Opus showed no
measurable self-preference. It scored its own translations 0.10 lower than Sol
did, in line with its stricter scoring overall. Sol scored its own output 0.36
higher than Opus did, against a typical gap of about 0.14—a possible mild
self-preference that, if anything, flatters Sol’s 4.06.
Quality scores are half the story. A fluent translation that drops a quiz
component still leaves work for an editor or developer, so each reviewer also
decided whether the output was ready to publish. Astra had the most
publish-ready runs, five. Opus had three. Sol and DeepSeek had none.
Model
Quality /5
Publish-ready runs
Articles/minute
Articles/$1
Claude Opus 5.5
4.37
3
3.2
16.2
GPT-6 Astra
4.32
5
2.5
9.4
GPT-6 Sol
4.06
0
4.1
47.8
GPT-5.6 Luna
3.95
2
4.5
419
Claude Fable 5.1
3.95
2
2.0
6.5
GPT-6 Luna
3.88
2
3.9
957
Gemini 3.8 Flash
3.81
1
6.1
129
GLM 5.3 Flash
3.80
1
4.1
872
DeepSeek V4.1 Flash
3.72
0
5.4
1,109
GLM 5.3 FlashX
3.67
1
5.1
359
Qwen 3.8 27B
3.28
0
7.3
180
Qwen 3.8 Flash
3.18
0
2.7
987
DeepSeek produced 1,109 Articles/$1, about 68 times as many as Opus, at
3.72/5. Those rates cover generation only. Neither judge marked any of
DeepSeek’s translations publish-ready, so the savings buy a first draft, not a
finished page.
Gemini 3.5 Flash Lite is left out of this ranking. Our import step broke one
of its valid outputs before review, and the provider returned no usable cost
for another. We will rerun it rather than rank a result our own harness damaged.
Building reusable golden datasets
The translation engine behind Dan Levy’s blog
has evolved through several generations. It demonstrates two ways to spend
frontier-model effort where it helps: generate competing translations and
judge/refine a selection, or build a reusable golden dataset through a more
thorough consensus process.
High-quality synthetic references are an engineering capability in their
own right. Our consensus workflow
starts with independent critiques, reconciles specific issues and exact edits,
and sends the revised article through separate audits. Source hashes,
structural checks, versioned references, and an issue ledger make the result
reusable and inspectable. The repository retains rejected proposals and
validation failures alongside the successful edits.
Creating the reference set has its own cost. The expanded collaboration
batch used an estimated $18.84 in model usage for five frozen reference
translations—about $3.77 per reference. That breaks down to $11.35 for
Opus, $4.01 for Sol, $2.04 for Fable, and $1.44 for Astra. The estimate
combines CLI-reported list-price costs with recorded token usage priced at
the run’s catalog rates, including cached-input discounts. It covers
collaborative refinement and audits; earlier pilots, starting translations,
engineering time, and human review are separate. The
cost evidence
records the calculation. Reference creation is also separate from generating
and judging benchmark candidates, and is excluded from the Articles/$1 chart.
This can be an economical approach when human-labeled gold is scarce or
expensive and the same cases will evaluate many models or prompt revisions.
Invest in the reference once, then reuse it to compare cheaper candidates.
The economics depend on reuse: reference creation and judging are separate
costs, not free additions to a generation call. Targeted human review remains
useful for important terminology and decisions that require domain expertise.
Here, generators received only the English source. Sol and Opus reviewed the
outputs against frozen consensus translations, with authorship hidden. The
reference was evaluated alongside each candidate rather than assigned a
perfect score. It was preferred in 122 of 130 reviews, and reviewers
agreed on preference for 59 of 65 pairs. These are model-assessed quality
scores; the small sample and reviewer preferences matter most when interpreting
close rankings.
Read a deployed article in French
or German to see the translation
system in use. These examples sit outside this panel.
Security verification: four models pass all six controls
Task
Read a task, response, and supporting context → return a structured verification judgment.
Measures
Whether the verifier accepts supported answers, rejects unsupported ones, and returns a judgment the application can parse.
Sample
Nine models. Six synthetic controls, each repeated three times: eighteen judgments per model. A gate check for the verifier, not a full security benchmark.
Sol, Opus, Astra, and Fable passed every control. Sol was the cheapest and
fastest of those four: 2.93-second median latency and $3.74 per 1,000
judgments.
Model
Thinking
Passed
Median latency
Cost / 1,000
GPT-6 Sol
none
18/18
2.93 s
$3.74
Claude Opus 5.5 *
low requested
18/18
3.64 s
$18.22
GPT-6 Astra
low
18/18
6.34 s
$22.41
Claude Fable 5.1
low
18/18
7.17 s
$34.42
DeepSeek V4.1 Flash
none
15/18
1.48 s
$0.27
Gemini 3.8 Flash
low
15/18
3.41 s
$1.69
GPT-6 Luna
none
13/18
2.22 s
$0.19
GLM 5.3 Flash
low
12/18
1.54 s
$0.25
Gemini 3.5 Flash Lite
minimal
12/18
1.22 s
$0.96
These controls come from the Agent Security work behind Exploit Hunter.
The engineering question is whether the verifier can distinguish completion
from a plausible claim and return a judgment the application can use. This
panel measures that verification layer, rather than an agent’s exploit-finding
or tool-use performance.
Every miss came from the same two of the six controls. All nine models
passed the other four in every repetition. That tells us exactly where to add
examples or route to a stronger verifier.
DeepSeek passed 15/18 at $0.27 per 1,000 judgments, about one
fourteenth of Sol’s cost. But a verifier that misses one case in six needs a
fallback before it gates anything.
Opus’s successful follow-up used strict JSON Schema through OpenRouter/Azure;
the other completed runs used JSON-object mode. Its original run stopped on
Markdown-fenced JSON. The output contract is part of the engineering result:
getting a sound judgment and getting reliably parseable output are both required.
A production selection should rerun the shortlist with matched settings and
representative task traces.
Handwriting: Gemini 3.5 Flash Lite as the default reader. It finished
within one page of the leader, ran three times faster, and cost about a
fortieth as much. Send pages that fail validation to Fable, the most accurate
reader here.
Translation: Astra when fewer editorial passes matter most; it drew the
most publish-ready runs. Opus matches its quality at about 40% lower cost
per article. Use Sol or DeepSeek only with an editor in the loop.
Security verification: Sol. It passed every control and was the cheapest
and fastest model that did. DeepSeek is worth a trial behind a stronger
fallback where its cost matters.
What we test next
Correction time for handwriting: which errors take a quick edit, and
which force someone to rebuild the record.
Human review of translation picks on real content, to check the judges’
publish-ready verdicts.
A clean rerun of Gemini 3.5 Flash Lite once the import fix lands.
Matched output settings for the security shortlist, and more cases than
six controls.
Our benchmarks grow with these projects. Some methods and cases are public;
others use private evaluation material. We are adding held-out cases to the
open suites so the published examples do not become the whole test. The goal
is a repeatable decision process that improves alongside the software.
Bring 17th Street Labs a workflow, an error budget, and a latency target.
We can help build an evaluation that shows which models meet your requirements,
where human review belongs, and what an acceptable result actually costs.
Results updated September 24, 2026 from the extraction comparison,
translation-generation benchmark, and security follow-ups. Each suite has its
own quality measure. Speed is projected workload throughput; value is inference
volume per $1, including imperfect outputs—not accepted work or total operating cost.
Page Scans: eleven independent documents, including a typed control;
the canary is excluded. Speed uses 660 / batch seconds. Value uses valid
document attempts divided by positive reported charges. Luna pools 33
attempts on the same eleven documents, with two exact results. The paired
bootstrap intervals in the evidence compare each model with GPT-6 Luna; we
did not compute intervals between the leaders, so close results are ties.
Articles: five attempts per model, a 24,000-token output cap, lowest
supported reasoning effort, and no retries. Speed uses completed Articles
divided by total request time; it is a serial
projection from a concurrency-four run. Value uses completed Articles per
all-attempt catalog generation cost, excluding reference and review costs.
Quality averages ten dimensions across two reviews per case. Sol and Opus
reviewed at high reasoning. Publish-ready runs are counted from the judges’ verdicts,
not human acceptance. Gemini 3.5 Flash Lite is excluded pending a
rerun (harness normalization regression, unknown cost). The source report
records normalization, component, and review-parsing issues.
Traces: speed is 60 / median request seconds; value is 18 / run cost,
retaining rubric failures. Six unique controls are repeated three times.
Opus used strict schema; the other completed runs used JSON-object mode.
Azure reported zero reasoning tokens for Opus despite low being requested.
Ranking uses source precision; displayed values are rounded. Equal quality
scores share tooltip ranks, with value and speed breaking display-order ties.
Recorded costs and catalog estimates follow each source’s accounting method.