17th Street LabsStart a project
← Blog

New Models, 3 New Evals: OCR, Translation, and Security Verification Tests

GPT-6, Claude Opus 5.5 and Fable 5.1, Gemini, GLM 5.3, DeepSeek V4.1 and Qwen 3.8 on the handwriting, translation, and security-verifier tests behind 17th Street Labs’ own software.

By Dan Levy, 17th Street Labs

Three evaluation tasks: handwriting becomes structured text, speech bubbles exchange translations, and a shield carries a verification check.

Same eleven page scans. GPT-6 Astra Pro cost $2.71 to read them; DeepSeek V4.1 Flash cost $0.014. For roughly 197 times the price, Astra Pro got one more field right—out of 132. That is the kind of difference a release-day leaderboard cannot tell you about your own workload.

17th Street Labs tests new models against software we build and maintain: DocSlurp’s handwriting extraction, the translation engine behind Dan Levy’s blog, and the Agent Security work behind Exploit Hunter. These tests ask models to read messy inputs, preserve meaning, and verify evidence. The samples are small—eleven documents, five articles, six security controls—so we say where results clearly separate and call the rest ties.

Fable read the most handwriting fields correctly. Opus and Astra tied at the top of translation, and Astra had the most publish-ready runs. Sol was the cheapest and fastest security verifier to pass every control. Our picks for each workload are at the end.

17th Street Labs · Three custom benchmarks

Choose the task. Compare the models.

Handwriting Recognition: model comparison

Quality

Field accuracy (%) · Higher is better

Claude Fable 5.1 Gemini 3.5 Flash Lite Claude Opus 5.5 GLM 5.3 Flash GLM 5.3 FlashX Gemini 3.1 Flash Lite Gemini 3.8 Flash GPT-6 Astra GPT-6 Sol GPT-6 Astra Pro DeepSeek V4.1 Flash GPT-5.6 Luna GPT-6 Luna 96.2 92.4 92.4 90.2 90.2 89.4 89.4 89.4 88.6 88.6 87.9 83.3 66.9 Claude Fable 5.1 Quality: 96.2% · #1 of 13 Speed: 22.6 Page Scans/min · #11 of 13 Value: 12.8 Page Scans/$1 · #11 of 13 Higher is better · Ranks use unrounded values.Gemini 3.5 Flash Lite Quality: 92.4% · Tied #2 of 13 Speed: 75.9 Page Scans/min · #1 of 13 Value: 533 Page Scans/$1 · #5 of 13 Higher is better · Ranks use unrounded values.Claude Opus 5.5 Quality: 92.4% · Tied #2 of 13 Speed: 23.8 Page Scans/min · #10 of 13 Value: 29.5 Page Scans/$1 · #10 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 90.2% · Tied #4 of 13 Speed: 64.7 Page Scans/min · #3 of 13 Value: 704 Page Scans/$1 · #3 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 FlashX Quality: 90.2% · Tied #4 of 13 Speed: 44.3 Page Scans/min · #5 of 13 Value: 291 Page Scans/$1 · #8 of 13 Higher is better · Ranks use unrounded values.Gemini 3.1 Flash Lite Quality: 89.4% · Tied #6 of 13 Speed: 72.5 Page Scans/min · #2 of 13 Value: 716 Page Scans/$1 · #2 of 13 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 89.4% · Tied #6 of 13 Speed: 39.8 Page Scans/min · #7 of 13 Value: 323 Page Scans/$1 · #7 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Quality: 89.4% · Tied #6 of 13 Speed: 21.3 Page Scans/min · #12 of 13 Value: 6.6 Page Scans/$1 · #12 of 13 Higher is better · Ranks use unrounded values.GPT-6 Sol Quality: 88.6% · Tied #9 of 13 Speed: 40.2 Page Scans/min · #6 of 13 Value: 32.7 Page Scans/$1 · #9 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Pro Quality: 88.6% · Tied #9 of 13 Speed: 10.9 Page Scans/min · #13 of 13 Value: 4.1 Page Scans/$1 · #13 of 13 Higher is better · Ranks use unrounded values.DeepSeek V4.1 Flash Quality: 87.9% · #11 of 13 Speed: 62.9 Page Scans/min · #4 of 13 Value: 802 Page Scans/$1 · #1 of 13 Higher is better · Ranks use unrounded values.GPT-5.6 Luna Quality: 83.3% · #12 of 13 Speed: 24.0 Page Scans/min · #9 of 13 Value: 366 Page Scans/$1 · #6 of 13 Higher is better · Ranks use unrounded values.GPT-6 Luna Quality: 66.9% · #13 of 13 Speed: 27.8 Page Scans/min · #8 of 13 Value: 610 Page Scans/$1 · #4 of 13 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View quality data
Handwriting Recognition: Field accuracy (%)
ModelField accuracy (%)
Claude Fable 5.196.2%
Gemini 3.5 Flash Lite92.4%
Claude Opus 5.592.4%
GLM 5.3 Flash90.2%
GLM 5.3 FlashX90.2%
Gemini 3.1 Flash Lite89.4%
Gemini 3.8 Flash89.4%
GPT-6 Astra89.4%
GPT-6 Sol88.6%
GPT-6 Astra Pro88.6%
DeepSeek V4.1 Flash87.9%
GPT-5.6 Luna83.3%
GPT-6 Luna66.9%

Zero baseline · Full score range

Speed

Page Scans per minute · Higher is better

Gemini 3.5 Flash Lite Gemini 3.1 Flash Lite GLM 5.3 Flash DeepSeek V4.1 Flash GLM 5.3 FlashX GPT-6 Sol Gemini 3.8 Flash GPT-6 Luna GPT-5.6 Luna Claude Opus 5.5 Claude Fable 5.1 GPT-6 Astra GPT-6 Astra Pro 75.9 72.5 64.7 62.9 44.3 40.2 39.8 27.8 24.0 23.8 22.6 21.3 10.9 Gemini 3.5 Flash Lite Quality: 92.4% · Tied #2 of 13 Speed: 75.9 Page Scans/min · #1 of 13 Value: 533 Page Scans/$1 · #5 of 13 Higher is better · Ranks use unrounded values.Gemini 3.1 Flash Lite Quality: 89.4% · Tied #6 of 13 Speed: 72.5 Page Scans/min · #2 of 13 Value: 716 Page Scans/$1 · #2 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 90.2% · Tied #4 of 13 Speed: 64.7 Page Scans/min · #3 of 13 Value: 704 Page Scans/$1 · #3 of 13 Higher is better · Ranks use unrounded values.DeepSeek V4.1 Flash Quality: 87.9% · #11 of 13 Speed: 62.9 Page Scans/min · #4 of 13 Value: 802 Page Scans/$1 · #1 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 FlashX Quality: 90.2% · Tied #4 of 13 Speed: 44.3 Page Scans/min · #5 of 13 Value: 291 Page Scans/$1 · #8 of 13 Higher is better · Ranks use unrounded values.GPT-6 Sol Quality: 88.6% · Tied #9 of 13 Speed: 40.2 Page Scans/min · #6 of 13 Value: 32.7 Page Scans/$1 · #9 of 13 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 89.4% · Tied #6 of 13 Speed: 39.8 Page Scans/min · #7 of 13 Value: 323 Page Scans/$1 · #7 of 13 Higher is better · Ranks use unrounded values.GPT-6 Luna Quality: 66.9% · #13 of 13 Speed: 27.8 Page Scans/min · #8 of 13 Value: 610 Page Scans/$1 · #4 of 13 Higher is better · Ranks use unrounded values.GPT-5.6 Luna Quality: 83.3% · #12 of 13 Speed: 24.0 Page Scans/min · #9 of 13 Value: 366 Page Scans/$1 · #6 of 13 Higher is better · Ranks use unrounded values.Claude Opus 5.5 Quality: 92.4% · Tied #2 of 13 Speed: 23.8 Page Scans/min · #10 of 13 Value: 29.5 Page Scans/$1 · #10 of 13 Higher is better · Ranks use unrounded values.Claude Fable 5.1 Quality: 96.2% · #1 of 13 Speed: 22.6 Page Scans/min · #11 of 13 Value: 12.8 Page Scans/$1 · #11 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Quality: 89.4% · Tied #6 of 13 Speed: 21.3 Page Scans/min · #12 of 13 Value: 6.6 Page Scans/$1 · #12 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Pro Quality: 88.6% · Tied #9 of 13 Speed: 10.9 Page Scans/min · #13 of 13 Value: 4.1 Page Scans/$1 · #13 of 13 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View speed data
Handwriting Recognition: Page Scans per minute
ModelPage Scans per minute
Gemini 3.5 Flash Lite75.9
Gemini 3.1 Flash Lite72.5
GLM 5.3 Flash64.7
DeepSeek V4.1 Flash62.9
GLM 5.3 FlashX44.3
GPT-6 Sol40.2
Gemini 3.8 Flash39.8
GPT-6 Luna27.8
GPT-5.6 Luna24.0
Claude Opus 5.523.8
Claude Fable 5.122.6
GPT-6 Astra21.3
GPT-6 Astra Pro10.9

Zero baseline · Scaled to the highest rate

Value

Page Scans per $1 · Higher is better

DeepSeek V4.1 Flash Gemini 3.1 Flash Lite GLM 5.3 Flash GPT-6 Luna Gemini 3.5 Flash Lite GPT-5.6 Luna Gemini 3.8 Flash GLM 5.3 FlashX GPT-6 Sol Claude Opus 5.5 Claude Fable 5.1 GPT-6 Astra GPT-6 Astra Pro 802 716 704 610 533 366 323 291 32.7 29.5 12.8 6.6 4.1 DeepSeek V4.1 Flash Quality: 87.9% · #11 of 13 Speed: 62.9 Page Scans/min · #4 of 13 Value: 802 Page Scans/$1 · #1 of 13 Higher is better · Ranks use unrounded values.Gemini 3.1 Flash Lite Quality: 89.4% · Tied #6 of 13 Speed: 72.5 Page Scans/min · #2 of 13 Value: 716 Page Scans/$1 · #2 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 90.2% · Tied #4 of 13 Speed: 64.7 Page Scans/min · #3 of 13 Value: 704 Page Scans/$1 · #3 of 13 Higher is better · Ranks use unrounded values.GPT-6 Luna Quality: 66.9% · #13 of 13 Speed: 27.8 Page Scans/min · #8 of 13 Value: 610 Page Scans/$1 · #4 of 13 Higher is better · Ranks use unrounded values.Gemini 3.5 Flash Lite Quality: 92.4% · Tied #2 of 13 Speed: 75.9 Page Scans/min · #1 of 13 Value: 533 Page Scans/$1 · #5 of 13 Higher is better · Ranks use unrounded values.GPT-5.6 Luna Quality: 83.3% · #12 of 13 Speed: 24.0 Page Scans/min · #9 of 13 Value: 366 Page Scans/$1 · #6 of 13 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 89.4% · Tied #6 of 13 Speed: 39.8 Page Scans/min · #7 of 13 Value: 323 Page Scans/$1 · #7 of 13 Higher is better · Ranks use unrounded values.GLM 5.3 FlashX Quality: 90.2% · Tied #4 of 13 Speed: 44.3 Page Scans/min · #5 of 13 Value: 291 Page Scans/$1 · #8 of 13 Higher is better · Ranks use unrounded values.GPT-6 Sol Quality: 88.6% · Tied #9 of 13 Speed: 40.2 Page Scans/min · #6 of 13 Value: 32.7 Page Scans/$1 · #9 of 13 Higher is better · Ranks use unrounded values.Claude Opus 5.5 Quality: 92.4% · Tied #2 of 13 Speed: 23.8 Page Scans/min · #10 of 13 Value: 29.5 Page Scans/$1 · #10 of 13 Higher is better · Ranks use unrounded values.Claude Fable 5.1 Quality: 96.2% · #1 of 13 Speed: 22.6 Page Scans/min · #11 of 13 Value: 12.8 Page Scans/$1 · #11 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Quality: 89.4% · Tied #6 of 13 Speed: 21.3 Page Scans/min · #12 of 13 Value: 6.6 Page Scans/$1 · #12 of 13 Higher is better · Ranks use unrounded values.GPT-6 Astra Pro Quality: 88.6% · Tied #9 of 13 Speed: 10.9 Page Scans/min · #13 of 13 Value: 4.1 Page Scans/$1 · #13 of 13 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View value data
Handwriting Recognition: Page Scans per $1
ModelPage Scans per $1
DeepSeek V4.1 Flash802
Gemini 3.1 Flash Lite716
GLM 5.3 Flash704
GPT-6 Luna610
Gemini 3.5 Flash Lite533
GPT-5.6 Luna366
Gemini 3.8 Flash323
GLM 5.3 FlashX291
GPT-6 Sol32.7
Claude Opus 5.529.5
Claude Fable 5.112.8
GPT-6 Astra6.6
GPT-6 Astra Pro4.1

Zero baseline · Scaled to the highest rate

Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.

11 documents, including a typed control. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.

Language Translation: model comparison

Quality

Translation quality /5 · Higher is better

Claude Opus 5.5 GPT-6 Astra GPT-6 Sol GPT-5.6 Luna Claude Fable 5.1 GPT-6 Luna Gemini 3.8 Flash GLM 5.3 Flash DeepSeek V4.1 Flash GLM 5.3 FlashX Qwen 3.8 27B Qwen 3.8 Flash 4.37 4.32 4.06 3.95 3.95 3.88 3.81 3.80 3.72 3.67 3.28 3.18 Claude Opus 5.5 Quality: 4.37/5 · #1 of 12 Speed: 3.2 Articles/min · #9 of 12 Value: 16.2 Articles/$1 · #10 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Astra Quality: 4.32/5 · #2 of 12 Speed: 2.5 Articles/min · #11 of 12 Value: 9.4 Articles/$1 · #11 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Sol Quality: 4.06/5 · #3 of 12 Speed: 4.1 Articles/min · #7 of 12 Value: 47.8 Articles/$1 · #9 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-5.6 Luna Quality: 3.95/5 · Tied #4 of 12 Speed: 4.5 Articles/min · #5 of 12 Value: 419 Articles/$1 · #5 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsClaude Fable 5.1 Quality: 3.95/5 · Tied #4 of 12 Speed: 2.0 Articles/min · #12 of 12 Value: 6.5 Articles/$1 · #12 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Luna Quality: 3.88/5 · #6 of 12 Speed: 3.9 Articles/min · #8 of 12 Value: 957 Articles/$1 · #3 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGemini 3.8 Flash Quality: 3.81/5 · #7 of 12 Speed: 6.1 Articles/min · #2 of 12 Value: 129 Articles/$1 · #8 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 Flash Quality: 3.80/5 · #8 of 12 Speed: 4.1 Articles/min · #6 of 12 Value: 872 Articles/$1 · #4 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsDeepSeek V4.1 Flash Quality: 3.72/5 · #9 of 12 Speed: 5.4 Articles/min · #3 of 12 Value: 1,109 Articles/$1 · #1 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 FlashX Quality: 3.67/5 · #10 of 12 Speed: 5.1 Articles/min · #4 of 12 Value: 359 Articles/$1 · #6 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsQwen 3.8 27B Quality: 3.28/5 · #11 of 12 Speed: 7.3 Articles/min · #1 of 12 Value: 180 Articles/$1 · #7 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsQwen 3.8 Flash Quality: 3.18/5 · #12 of 12 Speed: 2.7 Articles/min · #10 of 12 Value: 987 Articles/$1 · #2 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviews

Scroll horizontally for all models. Focus the chart to use arrow keys.

View quality data
Language Translation: Translation quality /5
ModelTranslation quality /5
Claude Opus 5.54.37/5
GPT-6 Astra4.32/5
GPT-6 Sol4.06/5
GPT-5.6 Luna3.95/5
Claude Fable 5.13.95/5
GPT-6 Luna3.88/5
Gemini 3.8 Flash3.81/5
GLM 5.3 Flash3.80/5
DeepSeek V4.1 Flash3.72/5
GLM 5.3 FlashX3.67/5
Qwen 3.8 27B3.28/5
Qwen 3.8 Flash3.18/5

Zero baseline · Full score range

Speed

Articles per minute · Higher is better

Qwen 3.8 27B Gemini 3.8 Flash DeepSeek V4.1 Flash GLM 5.3 FlashX GPT-5.6 Luna GLM 5.3 Flash GPT-6 Sol GPT-6 Luna Claude Opus 5.5 Qwen 3.8 Flash GPT-6 Astra Claude Fable 5.1 7.3 6.1 5.4 5.1 4.5 4.1 4.1 3.9 3.2 2.7 2.5 2.0 Qwen 3.8 27B Quality: 3.28/5 · #11 of 12 Speed: 7.3 Articles/min · #1 of 12 Value: 180 Articles/$1 · #7 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGemini 3.8 Flash Quality: 3.81/5 · #7 of 12 Speed: 6.1 Articles/min · #2 of 12 Value: 129 Articles/$1 · #8 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsDeepSeek V4.1 Flash Quality: 3.72/5 · #9 of 12 Speed: 5.4 Articles/min · #3 of 12 Value: 1,109 Articles/$1 · #1 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 FlashX Quality: 3.67/5 · #10 of 12 Speed: 5.1 Articles/min · #4 of 12 Value: 359 Articles/$1 · #6 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-5.6 Luna Quality: 3.95/5 · Tied #4 of 12 Speed: 4.5 Articles/min · #5 of 12 Value: 419 Articles/$1 · #5 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 Flash Quality: 3.80/5 · #8 of 12 Speed: 4.1 Articles/min · #6 of 12 Value: 872 Articles/$1 · #4 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Sol Quality: 4.06/5 · #3 of 12 Speed: 4.1 Articles/min · #7 of 12 Value: 47.8 Articles/$1 · #9 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Luna Quality: 3.88/5 · #6 of 12 Speed: 3.9 Articles/min · #8 of 12 Value: 957 Articles/$1 · #3 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsClaude Opus 5.5 Quality: 4.37/5 · #1 of 12 Speed: 3.2 Articles/min · #9 of 12 Value: 16.2 Articles/$1 · #10 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsQwen 3.8 Flash Quality: 3.18/5 · #12 of 12 Speed: 2.7 Articles/min · #10 of 12 Value: 987 Articles/$1 · #2 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Astra Quality: 4.32/5 · #2 of 12 Speed: 2.5 Articles/min · #11 of 12 Value: 9.4 Articles/$1 · #11 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsClaude Fable 5.1 Quality: 3.95/5 · Tied #4 of 12 Speed: 2.0 Articles/min · #12 of 12 Value: 6.5 Articles/$1 · #12 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviews

Scroll horizontally for all models. Focus the chart to use arrow keys.

View speed data
Language Translation: Articles per minute
ModelArticles per minute
Qwen 3.8 27B7.3
Gemini 3.8 Flash6.1
DeepSeek V4.1 Flash5.4
GLM 5.3 FlashX5.1
GPT-5.6 Luna4.5
GLM 5.3 Flash4.1
GPT-6 Sol4.1
GPT-6 Luna3.9
Claude Opus 5.53.2
Qwen 3.8 Flash2.7
GPT-6 Astra2.5
Claude Fable 5.12.0

Zero baseline · Scaled to the highest rate

Value

Articles per $1 · Higher is better

DeepSeek V4.1 Flash Qwen 3.8 Flash GPT-6 Luna GLM 5.3 Flash GPT-5.6 Luna GLM 5.3 FlashX Qwen 3.8 27B Gemini 3.8 Flash GPT-6 Sol Claude Opus 5.5 GPT-6 Astra Claude Fable 5.1 1,109 987 957 872 419 359 180 129 47.8 16.2 9.4 6.5 DeepSeek V4.1 Flash Quality: 3.72/5 · #9 of 12 Speed: 5.4 Articles/min · #3 of 12 Value: 1,109 Articles/$1 · #1 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsQwen 3.8 Flash Quality: 3.18/5 · #12 of 12 Speed: 2.7 Articles/min · #10 of 12 Value: 987 Articles/$1 · #2 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Luna Quality: 3.88/5 · #6 of 12 Speed: 3.9 Articles/min · #8 of 12 Value: 957 Articles/$1 · #3 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 Flash Quality: 3.80/5 · #8 of 12 Speed: 4.1 Articles/min · #6 of 12 Value: 872 Articles/$1 · #4 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-5.6 Luna Quality: 3.95/5 · Tied #4 of 12 Speed: 4.5 Articles/min · #5 of 12 Value: 419 Articles/$1 · #5 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGLM 5.3 FlashX Quality: 3.67/5 · #10 of 12 Speed: 5.1 Articles/min · #4 of 12 Value: 359 Articles/$1 · #6 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsQwen 3.8 27B Quality: 3.28/5 · #11 of 12 Speed: 7.3 Articles/min · #1 of 12 Value: 180 Articles/$1 · #7 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGemini 3.8 Flash Quality: 3.81/5 · #7 of 12 Speed: 6.1 Articles/min · #2 of 12 Value: 129 Articles/$1 · #8 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Sol Quality: 4.06/5 · #3 of 12 Speed: 4.1 Articles/min · #7 of 12 Value: 47.8 Articles/$1 · #9 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsClaude Opus 5.5 Quality: 4.37/5 · #1 of 12 Speed: 3.2 Articles/min · #9 of 12 Value: 16.2 Articles/$1 · #10 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsGPT-6 Astra Quality: 4.32/5 · #2 of 12 Speed: 2.5 Articles/min · #11 of 12 Value: 9.4 Articles/$1 · #11 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviewsClaude Fable 5.1 Quality: 3.95/5 · Tied #4 of 12 Speed: 2.0 Articles/min · #12 of 12 Value: 6.5 Articles/$1 · #12 of 12 Higher is better · Ranks use unrounded values. 5/5 Articles complete · 10 reviews

Scroll horizontally for all models. Focus the chart to use arrow keys.

View value data
Language Translation: Articles per $1
ModelArticles per $1
DeepSeek V4.1 Flash1,109
Qwen 3.8 Flash987
GPT-6 Luna957
GLM 5.3 Flash872
GPT-5.6 Luna419
GLM 5.3 FlashX359
Qwen 3.8 27B180
Gemini 3.8 Flash129
GPT-6 Sol47.8
Claude Opus 5.516.2
GPT-6 Astra9.4
Claude Fable 5.16.5

Zero baseline · Scaled to the highest rate

Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.

Five Articles, including a quiz; Spanish and Japanese. Two model reviewers, 10 reviews per model. Speed counts completed Articles over all request time. Gemini 3.5 Flash Lite is excluded: a harness normalization regression affected its reviewed output and its cost is unknown; a rerun is pending. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.

Security Task Verification: model comparison

Quality

Controls passed · Higher is better

GPT-6 Sol Claude Opus 5.5 * GPT-6 Astra Claude Fable 5.1 DeepSeek V4.1 Flash Gemini 3.8 Flash GPT-6 Luna GLM 5.3 Flash Gemini 3.5 Flash Lite 18/18 18/18 18/18 18/18 15/18 15/18 13/18 12/18 12/18 GPT-6 Sol Quality: 18/18 · Tied #1 of 9 Speed: 20.5 Traces/min · #5 of 9 Value: 268 Traces/$1 · #6 of 9 Higher is better · Ranks use unrounded values.Claude Opus 5.5 * Quality: 18/18 · Tied #1 of 9 Speed: 16.5 Traces/min · #7 of 9 Value: 54.9 Traces/$1 · #7 of 9 Higher is better · Ranks use unrounded values. Strict JSON Schema · low requested; 0 reasoning tokensGPT-6 Astra Quality: 18/18 · Tied #1 of 9 Speed: 9.5 Traces/min · #8 of 9 Value: 44.6 Traces/$1 · #8 of 9 Higher is better · Ranks use unrounded values.Claude Fable 5.1 Quality: 18/18 · Tied #1 of 9 Speed: 8.4 Traces/min · #9 of 9 Value: 29.1 Traces/$1 · #9 of 9 Higher is better · Ranks use unrounded values.DeepSeek V4.1 Flash Quality: 15/18 · Tied #5 of 9 Speed: 40.4 Traces/min · #2 of 9 Value: 3,660 Traces/$1 · #3 of 9 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 15/18 · Tied #5 of 9 Speed: 17.6 Traces/min · #6 of 9 Value: 593 Traces/$1 · #5 of 9 Higher is better · Ranks use unrounded values.GPT-6 Luna Quality: 13/18 · #7 of 9 Speed: 27.0 Traces/min · #4 of 9 Value: 5,162 Traces/$1 · #1 of 9 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 12/18 · Tied #8 of 9 Speed: 39.0 Traces/min · #3 of 9 Value: 3,955 Traces/$1 · #2 of 9 Higher is better · Ranks use unrounded values.Gemini 3.5 Flash Lite Quality: 12/18 · Tied #8 of 9 Speed: 49.3 Traces/min · #1 of 9 Value: 1,044 Traces/$1 · #4 of 9 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View quality data
Security Task Verification: Controls passed
ModelControls passed
GPT-6 Sol18/18
Claude Opus 5.5 *18/18
GPT-6 Astra18/18
Claude Fable 5.118/18
DeepSeek V4.1 Flash15/18
Gemini 3.8 Flash15/18
GPT-6 Luna13/18
GLM 5.3 Flash12/18
Gemini 3.5 Flash Lite12/18

Zero baseline · Full score range

Speed

Traces per minute · Higher is better

Gemini 3.5 Flash Lite DeepSeek V4.1 Flash GLM 5.3 Flash GPT-6 Luna GPT-6 Sol Gemini 3.8 Flash Claude Opus 5.5 * GPT-6 Astra Claude Fable 5.1 49.3 40.4 39.0 27.0 20.5 17.6 16.5 9.5 8.4 Gemini 3.5 Flash Lite Quality: 12/18 · Tied #8 of 9 Speed: 49.3 Traces/min · #1 of 9 Value: 1,044 Traces/$1 · #4 of 9 Higher is better · Ranks use unrounded values.DeepSeek V4.1 Flash Quality: 15/18 · Tied #5 of 9 Speed: 40.4 Traces/min · #2 of 9 Value: 3,660 Traces/$1 · #3 of 9 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 12/18 · Tied #8 of 9 Speed: 39.0 Traces/min · #3 of 9 Value: 3,955 Traces/$1 · #2 of 9 Higher is better · Ranks use unrounded values.GPT-6 Luna Quality: 13/18 · #7 of 9 Speed: 27.0 Traces/min · #4 of 9 Value: 5,162 Traces/$1 · #1 of 9 Higher is better · Ranks use unrounded values.GPT-6 Sol Quality: 18/18 · Tied #1 of 9 Speed: 20.5 Traces/min · #5 of 9 Value: 268 Traces/$1 · #6 of 9 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 15/18 · Tied #5 of 9 Speed: 17.6 Traces/min · #6 of 9 Value: 593 Traces/$1 · #5 of 9 Higher is better · Ranks use unrounded values.Claude Opus 5.5 * Quality: 18/18 · Tied #1 of 9 Speed: 16.5 Traces/min · #7 of 9 Value: 54.9 Traces/$1 · #7 of 9 Higher is better · Ranks use unrounded values. Strict JSON Schema · low requested; 0 reasoning tokensGPT-6 Astra Quality: 18/18 · Tied #1 of 9 Speed: 9.5 Traces/min · #8 of 9 Value: 44.6 Traces/$1 · #8 of 9 Higher is better · Ranks use unrounded values.Claude Fable 5.1 Quality: 18/18 · Tied #1 of 9 Speed: 8.4 Traces/min · #9 of 9 Value: 29.1 Traces/$1 · #9 of 9 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View speed data
Security Task Verification: Traces per minute
ModelTraces per minute
Gemini 3.5 Flash Lite49.3
DeepSeek V4.1 Flash40.4
GLM 5.3 Flash39.0
GPT-6 Luna27.0
GPT-6 Sol20.5
Gemini 3.8 Flash17.6
Claude Opus 5.5 *16.5
GPT-6 Astra9.5
Claude Fable 5.18.4

Zero baseline · Scaled to the highest rate

Value

Traces per $1 · Higher is better

GPT-6 Luna GLM 5.3 Flash DeepSeek V4.1 Flash Gemini 3.5 Flash Lite Gemini 3.8 Flash GPT-6 Sol Claude Opus 5.5 * GPT-6 Astra Claude Fable 5.1 5,162 3,955 3,660 1,044 593 268 54.9 44.6 29.1 GPT-6 Luna Quality: 13/18 · #7 of 9 Speed: 27.0 Traces/min · #4 of 9 Value: 5,162 Traces/$1 · #1 of 9 Higher is better · Ranks use unrounded values.GLM 5.3 Flash Quality: 12/18 · Tied #8 of 9 Speed: 39.0 Traces/min · #3 of 9 Value: 3,955 Traces/$1 · #2 of 9 Higher is better · Ranks use unrounded values.DeepSeek V4.1 Flash Quality: 15/18 · Tied #5 of 9 Speed: 40.4 Traces/min · #2 of 9 Value: 3,660 Traces/$1 · #3 of 9 Higher is better · Ranks use unrounded values.Gemini 3.5 Flash Lite Quality: 12/18 · Tied #8 of 9 Speed: 49.3 Traces/min · #1 of 9 Value: 1,044 Traces/$1 · #4 of 9 Higher is better · Ranks use unrounded values.Gemini 3.8 Flash Quality: 15/18 · Tied #5 of 9 Speed: 17.6 Traces/min · #6 of 9 Value: 593 Traces/$1 · #5 of 9 Higher is better · Ranks use unrounded values.GPT-6 Sol Quality: 18/18 · Tied #1 of 9 Speed: 20.5 Traces/min · #5 of 9 Value: 268 Traces/$1 · #6 of 9 Higher is better · Ranks use unrounded values.Claude Opus 5.5 * Quality: 18/18 · Tied #1 of 9 Speed: 16.5 Traces/min · #7 of 9 Value: 54.9 Traces/$1 · #7 of 9 Higher is better · Ranks use unrounded values. Strict JSON Schema · low requested; 0 reasoning tokensGPT-6 Astra Quality: 18/18 · Tied #1 of 9 Speed: 9.5 Traces/min · #8 of 9 Value: 44.6 Traces/$1 · #8 of 9 Higher is better · Ranks use unrounded values.Claude Fable 5.1 Quality: 18/18 · Tied #1 of 9 Speed: 8.4 Traces/min · #9 of 9 Value: 29.1 Traces/$1 · #9 of 9 Higher is better · Ranks use unrounded values.

Scroll horizontally for all models. Focus the chart to use arrow keys.

View value data
Security Task Verification: Traces per $1
ModelTraces per $1
GPT-6 Luna5,162
GLM 5.3 Flash3,955
DeepSeek V4.1 Flash3,660
Gemini 3.5 Flash Lite1,044
Gemini 3.8 Flash593
GPT-6 Sol268
Claude Opus 5.5 *54.9
GPT-6 Astra44.6
Claude Fable 5.129.1

Zero baseline · Scaled to the highest rate

Value measures inference volume per $1, including imperfect outputs. Speed shows units per minute. Hover or focus a bar for details. Click or tap a bar to highlight that model across all three charts.

6 synthetic controls repeated 3 times. * Opus used strict JSON Schema; all other completed runs used JSON-object mode. Output settings are not matched. Quality measures differ by suite. Rates include imperfect outputs; compare within each suite only.

Explore the sortable shortlists and all model results

Featured model families

  • gpt-6
  • opus-5.5
  • glm-5.3
  • deepseek-v4.1-flash
  • gemini-3.8-flash
  • qwen-3.8
  • fable-5.1

Three test suites

The shortlists

Choose quality, speed, or cost. Expand each list to see all models.

Relative score: LowerMiddleHigherThirds of each test’s observed range, not pass/fail grades.

Handwriting Recognition

Quality: field accuracy

Highest quality first. Ties: value, then speed.

  1. 1. Claude Fable 5.196.2%22.6/min12.8/$1

    Quality96.2%Speed22.6/minValue12.8/$1

  2. 2. Gemini 3.5 Flash Lite92.4%75.9/min533/$1

    Quality92.4%Speed75.9/minValue533/$1

  3. 3. Claude Opus 5.592.4%23.8/min29.5/$1

    Quality92.4%Speed23.8/minValue29.5/$1

Show all 13 models
  1. 4. GLM 5.3 Flash90.2%64.7/min704/$1

    Quality90.2%Speed64.7/minValue704/$1

  2. 5. GLM 5.3 FlashX90.2%44.3/min291/$1

    Quality90.2%Speed44.3/minValue291/$1

  3. 6. Gemini 3.1 Flash Lite89.4%72.5/min716/$1

    Quality89.4%Speed72.5/minValue716/$1

  4. 7. Gemini 3.8 Flash89.4%39.8/min323/$1

    Quality89.4%Speed39.8/minValue323/$1

  5. 8. GPT-6 Astra89.4%21.3/min6.6/$1

    Quality89.4%Speed21.3/minValue6.6/$1

  6. 9. GPT-6 Sol88.6%40.2/min32.7/$1

    Quality88.6%Speed40.2/minValue32.7/$1

  7. 10. GPT-6 Astra Pro88.6%10.9/min4.1/$1

    Quality88.6%Speed10.9/minValue4.1/$1

  8. 11. DeepSeek V4.1 Flash87.9%62.9/min802/$1

    Quality87.9%Speed62.9/minValue802/$1

  9. 12. GPT-5.6 Luna83.3%24.0/min366/$1

    Quality83.3%Speed24.0/minValue366/$1

  10. 13. GPT-6 Luna66.9%27.8/min610/$1

    Quality66.9%Speed27.8/minValue610/$1

Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.

11 documents, including a typed control.

Language Translation

Quality: translation quality /5

Highest quality first. Ties: value, then speed.

  1. 1. Claude Opus 5.54.37/53.2/min16.2/$1

    Quality4.37/5Speed3.2/minValue16.2/$1

  2. 2. GPT-6 Astra4.32/52.5/min9.4/$1

    Quality4.32/5Speed2.5/minValue9.4/$1

  3. 3. GPT-6 Sol4.06/54.1/min47.8/$1

    Quality4.06/5Speed4.1/minValue47.8/$1

Show all 12 models
  1. 4. GPT-5.6 Luna3.95/54.5/min419/$1

    Quality3.95/5Speed4.5/minValue419/$1

  2. 5. Claude Fable 5.13.95/52.0/min6.5/$1

    Quality3.95/5Speed2.0/minValue6.5/$1

  3. 6. GPT-6 Luna3.88/53.9/min957/$1

    Quality3.88/5Speed3.9/minValue957/$1

  4. 7. Gemini 3.8 Flash3.81/56.1/min129/$1

    Quality3.81/5Speed6.1/minValue129/$1

  5. 8. GLM 5.3 Flash3.80/54.1/min872/$1

    Quality3.80/5Speed4.1/minValue872/$1

  6. 9. DeepSeek V4.1 Flash3.72/55.4/min1,109/$1

    Quality3.72/5Speed5.4/minValue1,109/$1

  7. 10. GLM 5.3 FlashX3.67/55.1/min359/$1

    Quality3.67/5Speed5.1/minValue359/$1

  8. 11. Qwen 3.8 27B3.28/57.3/min180/$1

    Quality3.28/5Speed7.3/minValue180/$1

  9. 12. Qwen 3.8 Flash3.18/52.7/min987/$1

    Quality3.18/5Speed2.7/minValue987/$1

Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.

Five Articles, including a quiz; Spanish and Japanese. Two model reviewers, 10 reviews per model. Speed counts completed Articles over all request time. Gemini 3.5 Flash Lite is excluded: a harness normalization regression affected its reviewed output and its cost is unknown; a rerun is pending.

Security Task Verification

Quality: controls passed

Highest quality first. Ties: value, then speed.

  1. 1. GPT-6 Sol18/1820.5/min268/$1

    Quality18/18Speed20.5/minValue268/$1

  2. 2. Claude Opus 5.5 *18/1816.5/min54.9/$1
    Strict JSON Schema · low requested; 0 reasoning tokens

    Quality18/18Speed16.5/minValue54.9/$1

  3. 3. GPT-6 Astra18/189.5/min44.6/$1

    Quality18/18Speed9.5/minValue44.6/$1

Show all 9 models
  1. 4. Claude Fable 5.118/188.4/min29.1/$1

    Quality18/18Speed8.4/minValue29.1/$1

  2. 5. DeepSeek V4.1 Flash15/1840.4/min3,660/$1

    Quality15/18Speed40.4/minValue3,660/$1

  3. 6. Gemini 3.8 Flash15/1817.6/min593/$1

    Quality15/18Speed17.6/minValue593/$1

  4. 7. GPT-6 Luna13/1827.0/min5,162/$1

    Quality13/18Speed27.0/minValue5,162/$1

  5. 8. GLM 5.3 Flash12/1839.0/min3,955/$1

    Quality12/18Speed39.0/minValue3,955/$1

  6. 9. Gemini 3.5 Flash Lite12/1849.3/min1,044/$1

    Quality12/18Speed49.3/minValue1,044/$1

Bars start at zero. Quality uses the full score range; speed and value use the highest rate in this suite. Higher is better.

6 synthetic controls repeated 3 times. * Opus used strict JSON Schema; all other completed runs used JSON-object mode. Output settings are not matched.

Quality measures differ by task. Translation quality is model-assessed on five Articles, with two reviews per output. Rates count completed generations; unknown cost is unranked. Security includes an Opus strict-schema follow-up with different output settings.

Handwriting OCR: what the extra accuracy costs

Task
Phone photographs → structured fields. Direct vision extraction, before retrieval or the rest of the ingestion pipeline.
Measures
Field accuracy and exact-document rate: how often every required field is correct, scored against expected values.
Sample
Thirteen models, eleven independent documents including a typed control. A privately held-out corpus for DocSlurp, growing with iPhone and Android photographs.

Fable 5.1 extracted 96.2% of fields correctly and returned eight of eleven page scans fully correct. Gemini 3.5 Flash Lite and Opus 5.5 tied at 92.4% field accuracy, but Gemini returned seven exact documents; Opus returned five. GLM Flash and FlashX tied at 90.2% of fields, with four and six exact documents respectively.

With eleven documents, a single page moves the exact-document rate by nine points. Fable’s lead over Gemini is one page and five fields. Treat the top three as a shortlist, not a final order.

That distinction matters to an ingestion system. A mostly correct page can still send the wrong date, amount, or name into a knowledge base. DocSlurp, our document-ingestion and RAG project, uses this benchmark to measure that first step. We plan to open source the project while keeping evaluation material held out to reduce exposure of the answers.

Updated extraction results: Fable leads at 96.2% field accuracy and 72.7% exact documents; Gemini Flash Lite and Opus tie at 92.4% of fields.

Circles show field accuracy; diamonds show exact-document rate.

Model Field accuracy Exact documents Page Scans/min Page Scans/$1
Claude Fable 5.1 96.2% 72.7% 22.6 12.8
Gemini 3.5 Flash Lite 92.4% 63.6% 75.9 533
Claude Opus 5.5 92.4% 45.5% 23.8 29.5
GLM 5.3 Flash 90.2% 36.4% 64.7 704
GLM 5.3 FlashX 90.2% 54.5% 44.3 291
Gemini 3.1 Flash Lite 89.4% 54.5% 72.5 716
GPT-6 Astra 89.4% 45.5% 21.3 6.6
Gemini 3.8 Flash 89.4% 45.5% 39.8 323
GPT-6 Astra Pro 88.6% 45.5% 10.9 4.1
GPT-6 Sol 88.6% 54.5% 40.2 32.7
DeepSeek V4.1 Flash 87.9% 36.4% 62.9 802
GPT-5.6 Luna 83.3% 36.4% 24.0 366
GPT-6 Luna 66.9% 6.1% 27.8 610

Fable’s eleven-scan run cost $0.86, about 42 times Gemini Flash Lite’s $0.021, for one additional exact document. Opus cost $0.37—about 18 times Gemini’s bill—with the same field accuracy and two fewer exact pages.

Whether that extra page is worth $0.84 depends on what a wrong page costs. If every error means a person rebuilds the record by hand, Fable pays for itself quickly. If downstream validation catches most mistakes, Gemini is the better buy.

Gemini Flash Lite was also the fastest, at 75.9 Page Scans/minute. DeepSeek delivered the most scans per dollar, 802, at 87.9% field accuracy, which makes it a strong low-cost baseline for a larger document test.

Extraction evidence · Download the data.

i18n translation: quality gains and inference costs

Task
English Articles → Spanish or Japanese MDX, including a quiz. Preserve meaning, technical claims, code, links, and working components.
Measures
Ten quality dimensions, including faithfulness, technical accuracy, readability, authorial voice, and MDX integrity. Two frontier reviewers assess each output against a frozen reference.
Sample
Twelve ranked models, five source articles each (one is a quiz), two reviews per translation.

Opus 5.5 (4.37/5) and Astra (4.32) are effectively tied at the top; Sol follows at 4.06. The two judges disagree on the order: Sol ranked Opus first (4.42 to 4.28), and Opus ranked Astra first (4.36 to 4.32). Opus showed no measurable self-preference. It scored its own translations 0.10 lower than Sol did, in line with its stricter scoring overall. Sol scored its own output 0.36 higher than Opus did, against a typical gap of about 0.14—a possible mild self-preference that, if anything, flatters Sol’s 4.06.

Quality scores are half the story. A fluent translation that drops a quiz component still leaves work for an editor or developer, so each reviewer also decided whether the output was ready to publish. Astra had the most publish-ready runs, five. Opus had three. Sol and DeepSeek had none.

Model Quality /5 Publish-ready runs Articles/minute Articles/$1
Claude Opus 5.5 4.37 3 3.2 16.2
GPT-6 Astra 4.32 5 2.5 9.4
GPT-6 Sol 4.06 0 4.1 47.8
GPT-5.6 Luna 3.95 2 4.5 419
Claude Fable 5.1 3.95 2 2.0 6.5
GPT-6 Luna 3.88 2 3.9 957
Gemini 3.8 Flash 3.81 1 6.1 129
GLM 5.3 Flash 3.80 1 4.1 872
DeepSeek V4.1 Flash 3.72 0 5.4 1,109
GLM 5.3 FlashX 3.67 1 5.1 359
Qwen 3.8 27B 3.28 0 7.3 180
Qwen 3.8 Flash 3.18 0 2.7 987

DeepSeek produced 1,109 Articles/$1, about 68 times as many as Opus, at 3.72/5. Those rates cover generation only. Neither judge marked any of DeepSeek’s translations publish-ready, so the savings buy a first draft, not a finished page.

Gemini 3.5 Flash Lite is left out of this ranking. Our import step broke one of its valid outputs before review, and the provider returned no usable cost for another. We will rerun it rather than rank a result our own harness damaged.

Building reusable golden datasets

The translation engine behind Dan Levy’s blog has evolved through several generations. It demonstrates two ways to spend frontier-model effort where it helps: generate competing translations and judge/refine a selection, or build a reusable golden dataset through a more thorough consensus process.

High-quality synthetic references are an engineering capability in their own right. Our consensus workflow starts with independent critiques, reconciles specific issues and exact edits, and sends the revised article through separate audits. Source hashes, structural checks, versioned references, and an issue ledger make the result reusable and inspectable. The repository retains rejected proposals and validation failures alongside the successful edits.

Creating the reference set has its own cost. The expanded collaboration batch used an estimated $18.84 in model usage for five frozen reference translations—about $3.77 per reference. That breaks down to $11.35 for Opus, $4.01 for Sol, $2.04 for Fable, and $1.44 for Astra. The estimate combines CLI-reported list-price costs with recorded token usage priced at the run’s catalog rates, including cached-input discounts. It covers collaborative refinement and audits; earlier pilots, starting translations, engineering time, and human review are separate. The cost evidence records the calculation. Reference creation is also separate from generating and judging benchmark candidates, and is excluded from the Articles/$1 chart.

This can be an economical approach when human-labeled gold is scarce or expensive and the same cases will evaluate many models or prompt revisions. Invest in the reference once, then reuse it to compare cheaper candidates. The economics depend on reuse: reference creation and judging are separate costs, not free additions to a generation call. Targeted human review remains useful for important terminology and decisions that require domain expertise.

Here, generators received only the English source. Sol and Opus reviewed the outputs against frozen consensus translations, with authorship hidden. The reference was evaluated alongside each candidate rather than assigned a perfect score. It was preferred in 122 of 130 reviews, and reviewers agreed on preference for 59 of 65 pairs. These are model-assessed quality scores; the small sample and reviewer preferences matter most when interpreting close rankings.

Read a deployed article in French or German to see the translation system in use. These examples sit outside this panel.

Generation report · Scoring and rate definitions.

Security verification: four models pass all six controls

Task
Read a task, response, and supporting context → return a structured verification judgment.
Measures
Whether the verifier accepts supported answers, rejects unsupported ones, and returns a judgment the application can parse.
Sample
Nine models. Six synthetic controls, each repeated three times: eighteen judgments per model. A gate check for the verifier, not a full security benchmark.

Sol, Opus, Astra, and Fable passed every control. Sol was the cheapest and fastest of those four: 2.93-second median latency and $3.74 per 1,000 judgments.

Model Thinking Passed Median latency Cost / 1,000
GPT-6 Sol none 18/18 2.93 s $3.74
Claude Opus 5.5 * low requested 18/18 3.64 s $18.22
GPT-6 Astra low 18/18 6.34 s $22.41
Claude Fable 5.1 low 18/18 7.17 s $34.42
DeepSeek V4.1 Flash none 15/18 1.48 s $0.27
Gemini 3.8 Flash low 15/18 3.41 s $1.69
GPT-6 Luna none 13/18 2.22 s $0.19
GLM 5.3 Flash low 12/18 1.54 s $0.25
Gemini 3.5 Flash Lite minimal 12/18 1.22 s $0.96

These controls come from the Agent Security work behind Exploit Hunter. The engineering question is whether the verifier can distinguish completion from a plausible claim and return a judgment the application can use. This panel measures that verification layer, rather than an agent’s exploit-finding or tool-use performance.

Every miss came from the same two of the six controls. All nine models passed the other four in every repetition. That tells us exactly where to add examples or route to a stronger verifier.

DeepSeek passed 15/18 at $0.27 per 1,000 judgments, about one fourteenth of Sol’s cost. But a verifier that misses one case in six needs a fallback before it gates anything.

Opus’s successful follow-up used strict JSON Schema through OpenRouter/Azure; the other completed runs used JSON-object mode. Its original run stopped on Markdown-fenced JSON. The output contract is part of the engineering result: getting a sound judgment and getting reliably parseable output are both required. A production selection should rerun the shortlist with matched settings and representative task traces.

Security evidence · Opus schema follow-up.

What we would pick today

Handwriting: Gemini 3.5 Flash Lite as the default reader. It finished within one page of the leader, ran three times faster, and cost about a fortieth as much. Send pages that fail validation to Fable, the most accurate reader here.

Translation: Astra when fewer editorial passes matter most; it drew the most publish-ready runs. Opus matches its quality at about 40% lower cost per article. Use Sol or DeepSeek only with an editor in the loop.

Security verification: Sol. It passed every control and was the cheapest and fastest model that did. DeepSeek is worth a trial behind a stronger fallback where its cost matters.

What we test next

Our benchmarks grow with these projects. Some methods and cases are public; others use private evaluation material. We are adding held-out cases to the open suites so the published examples do not become the whole test. The goal is a repeatable decision process that improves alongside the software.

Bring 17th Street Labs a workflow, an error budget, and a latency target. We can help build an evaluation that shows which models meet your requirements, where human review belongs, and what an acceptable result actually costs.

Talk to 17th Street Labs about evaluating your AI workflow.

Measurement details and source data

Results updated September 24, 2026 from the extraction comparison, translation-generation benchmark, and security follow-ups. Each suite has its own quality measure. Speed is projected workload throughput; value is inference volume per $1, including imperfect outputs—not accepted work or total operating cost.

  • Page Scans: eleven independent documents, including a typed control; the canary is excluded. Speed uses 660 / batch seconds. Value uses valid document attempts divided by positive reported charges. Luna pools 33 attempts on the same eleven documents, with two exact results. The paired bootstrap intervals in the evidence compare each model with GPT-6 Luna; we did not compute intervals between the leaders, so close results are ties.
  • Articles: five attempts per model, a 24,000-token output cap, lowest supported reasoning effort, and no retries. Speed uses completed Articles divided by total request time; it is a serial projection from a concurrency-four run. Value uses completed Articles per all-attempt catalog generation cost, excluding reference and review costs. Quality averages ten dimensions across two reviews per case. Sol and Opus reviewed at high reasoning. Publish-ready runs are counted from the judges’ verdicts, not human acceptance. Gemini 3.5 Flash Lite is excluded pending a rerun (harness normalization regression, unknown cost). The source report records normalization, component, and review-parsing issues.
  • Traces: speed is 60 / median request seconds; value is 18 / run cost, retaining rubric failures. Six unique controls are repeated three times. Opus used strict schema; the other completed runs used JSON-object mode. Azure reported zero reasoning tokens for Opus despite low being requested.

Ranking uses source precision; displayed values are rounded. Equal quality scores share tooltip ranks, with value and speed breaking display-order ties. Recorded costs and catalog estimates follow each source’s accounting method.

Reproduction notes · Recognition data · Translation data · Security data · Efficiency data.

This is part of our work on AI Evals & Reliability.

Blog

There’s more where
that came from.

Enjoying the useful bits? Pop in your email and explore the Blog.

Free access, remembered in this browser. We use your email to register your access—not to sign you up for marketing.

Tell us what you are building.

A few sentences is plenty. An engineer reads every message.

Goes straight to the team. Reply within a business day.