# Measured extraction comparison

Direct model extraction, not pipeline or LLM judge performance. Eleven independent documents. Efficiency uses positive reported charges, never list estimates or zero accounting values. Results are grouped by model regardless of serving backend. Billing verification stays in the raw call records. Confidence calibration uses valid calls only; failed attempts are listed separately.

Docs/$1 = valid document attempts / reported cost. Cost and time gaps are `(value / lowest value - 1) × 100`, compared per document among complete, positively billed model results. GPT-6 Luna aggregates three rounds by total cost and documents; displayed cost/time are means per 11-document batch. These are observed workload projections, not guaranteed throughput or correctly extracted documents per dollar.

| Model | Field accuracy | Docs / $1 | Cost / 11 docs (above cheapest) | Wall time (slower than fastest) |
| --- | ---: | ---: | ---: | ---: |
| anthropic/claude-fable-5.1 | 96.2% | 12.8 | $0.86181 (+6,180.3%) | 29.2s (+237.3%) |
| google/gemini-3.5-flash-lite | 92.4% | 533 | $0.02062 (+50.3%) | 8.7s (+0.0%) |
| anthropic/claude-opus-5.5 | 92.4% | 29.5 | $0.37246 (+2,614.2%) | 27.7s (+220.1%) |
| z-ai/glm-5.3-flash | 90.2% | 704 | $0.01563 (+13.9%) | 10.2s (+17.8%) |
| z-ai/glm-5.3-flashx | 90.2% | 291 | $0.03783 (+175.7%) | 14.9s (+71.7%) |
| google/gemini-3.1-flash-lite | 89.4% | 716 | $0.01537 (+12.0%) | 9.1s (+4.9%) |
| openai/gpt-6-astra | 89.4% | 6.6 | $1.66424 (+12,027.8%) | 31.0s (+257.4%) |
| google/gemini-3.8-flash | 89.4% | 323 | $0.03409 (+148.4%) | 16.6s (+91.1%) |
| openai/gpt-6-astra-pro | 88.6% | 4.1 | $2.70819 (+19,635.4%) | 60.5s (+598.9%) |
| openai/gpt-6-sol | 88.6% | 32.7 | $0.33683 (+2,354.6%) | 16.4s (+88.9%) |
| deepseek/deepseek-v4.1-flash | 87.9% | 802 | $0.01372 (+0.0%) | 10.5s (+21.5%) |
| openai/gpt-5.6-luna | 83.3% | 366 | $0.03009 (+119.3%) | 27.5s (+217.0%) |
| openai/gpt-6-luna | 66.9% | 610 | $0.01802 (+31.3%) | 23.7s (+174.0%) |

## Individual rounds

| Model / round | Valid | Field accuracy | Exact documents | Reported cost / list estimate | Median / p95 request | Wall time | Auto-accept error |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| anthropic/claude-fable-5.1 (fable51) | 11/11 | 96.2% | 72.7% | $0.86181 / $0.86181 | 9.0s / 10.7s | 29.2s | 1.8% |
| google/gemini-3.5-flash-lite (gemini35flashlite) | 11/11 | 92.4% | 63.6% | $0.02062 / $0.02062 | 2.9s / 3.3s | 8.7s | 7.0% |
| anthropic/claude-opus-5.5 (opus55) | 11/11 | 92.4% | 45.5% | $0.37246 / $0.37246 | 8.4s / 12.8s | 27.7s | 2.1% |
| z-ai/glm-5.3-flash (glm53flash) | 11/11 | 90.2% | 36.4% | $0.01563 / $0.01563 | 2.7s / 10.1s | 10.2s | 3.6% |
| z-ai/glm-5.3-flashx (glm53flashx) | 11/11 | 90.2% | 54.5% | $0.03783 / $0.03859 | 4.5s / 7.9s | 14.9s | 4.9% |
| google/gemini-3.1-flash-lite (gemini31flashlite) | 11/11 | 89.4% | 54.5% | $0.01537 / $0.01537 | 2.9s / 3.5s | 9.1s | 7.1% |
| openai/gpt-6-astra (astra6) | 11/11 | 89.4% | 45.5% | $1.66424 / $1.55057 | 10.1s / 12.3s | 31.0s | 7.9% |
| google/gemini-3.8-flash (gemini38flash) | 11/11 | 89.4% | 45.5% | $0.03409 / $0.03409 | 4.6s / 13.9s | 16.6s | 9.4% |
| openai/gpt-6-astra-pro (astra6-pro) | 11/11 | 88.6% | 45.5% | $2.70819 / $3.86414 | 14.4s / 32.0s | 60.5s | 9.4% |
| openai/gpt-6-sol (sol6) | 11/11 | 88.6% | 54.5% | $0.33683 / $0.31373 | 5.2s / 6.0s | 16.4s | 9.6% |
| deepseek/deepseek-v4.1-flash (deepseek41flash) | 11/11 | 87.9% | 36.4% | $0.01372 / $0.00833 | 3.5s / 5.7s | 10.5s | 7.1% |
| openai/gpt-5.6-luna (luna56) | 11/11 | 83.3% | 36.4% | $0.03009 / $0.02811 | 7.6s / 14.0s | 27.5s | 11.9% |
| openai/gpt-6-luna (luna6-r3) | 11/11 | 67.4% | 0.0% | $0.01804 / $0.01678 | 5.0s / 11.0s | 21.0s | 21.1% |
| openai/gpt-6-luna (luna6-r1) | 11/11 | 66.7% | 9.1% | $0.01815 / $0.01687 | 10.2s / 15.2s | 30.3s | 24.2% |
| openai/gpt-6-luna (luna6-r2) | 11/11 | 66.7% | 9.1% | $0.01788 / $0.01663 | 4.8s / 10.3s | 19.9s | 24.1% |

## Paired differences against GPT-6 Luna round 1

Bootstrap intervals resample documents, not fields. These are exploratory comparisons without multiplicity correction. A positive delta favors the candidate.

| Candidate | Delta (percentage points) | 95% interval |
| --- | ---: | ---: |
| shared-openrouter/astra6-pro.json | 21.97 | [12.12, 32.58] |
| shared-openrouter/astra6.json | 22.73 | [12.12, 34.09] |
| deepseek41flash.json | 21.21 | [9.09, 34.09] |
| shared-openrouter/fable51.json | 29.55 | [18.18, 41.67] |
| gemini31flashlite.json | 22.73 | [11.36, 34.85] |
| gemini35flashlite.json | 25.76 | [15.15, 37.88] |
| gemini38flash.json | 22.73 | [12.12, 34.85] |
| glm53flash.json | 23.48 | [14.39, 33.33] |
| glm53flashx.json | 23.48 | [14.39, 33.33] |
| shared-openrouter/luna56.json | 16.67 | [6.06, 28.03] |
| shared-openrouter/luna6-r2.json | 0.00 | [-6.06, 5.30] |
| shared-openrouter/luna6-r3.json | 0.76 | [-4.55, 6.06] |
| shared-openrouter/opus55.json | 25.76 | [15.15, 37.88] |
| shared-openrouter/sol6.json | 21.97 | [9.85, 34.09] |

## GPT-6 Luna repeatability

```json
{
  "rounds": 3,
  "independentDocuments": 11,
  "perDocument": [
    {
      "id": "digital_typed_01",
      "accuracies": [
        1,
        1,
        0.9166666666666666
      ],
      "correctnessFlips": 1,
      "rawValueFlips": 1
    },
    {
      "id": "handwritten_messy_01",
      "accuracies": [
        0.3333333333333333,
        0.3333333333333333,
        0.5
      ],
      "correctnessFlips": 2,
      "rawValueFlips": 4
    },
    {
      "id": "handwritten_messy_02",
      "accuracies": [
        0.9166666666666666,
        0.9166666666666666,
        0.9166666666666666
      ],
      "correctnessFlips": 0,
      "rawValueFlips": 1
    },
    {
      "id": "handwritten_messy_03",
      "accuracies": [
        0.5833333333333334,
        0.75,
        0.5
      ],
      "correctnessFlips": 3,
      "rawValueFlips": 5
    },
    {
      "id": "handwritten_neat_01",
      "accuracies": [
        0.5833333333333334,
        0.5833333333333334,
        0.6666666666666666
      ],
      "correctnessFlips": 1,
      "rawValueFlips": 4
    },
    {
      "id": "handwritten_neat_02",
      "accuracies": [
        0.75,
        0.8333333333333334,
        0.8333333333333334
      ],
      "correctnessFlips": 3,
      "rawValueFlips": 4
    },
    {
      "id": "handwritten_neat_03",
      "accuracies": [
        0.3333333333333333,
        0.3333333333333333,
        0.3333333333333333
      ],
      "correctnessFlips": 0,
      "rawValueFlips": 3
    },
    {
      "id": "handwritten_neat_04_lowlight",
      "accuracies": [
        0.6666666666666666,
        0.6666666666666666,
        0.6666666666666666
      ],
      "correctnessFlips": 0,
      "rawValueFlips": 1
    },
    {
      "id": "handwritten_neat_05",
      "accuracies": [
        0.5833333333333334,
        0.5833333333333334,
        0.6666666666666666
      ],
      "correctnessFlips": 1,
      "rawValueFlips": 3
    },
    {
      "id": "handwritten_neat_06",
      "accuracies": [
        0.9166666666666666,
        0.9166666666666666,
        0.9166666666666666
      ],
      "correctnessFlips": 0,
      "rawValueFlips": 0
    },
    {
      "id": "handwritten_neat_07",
      "accuracies": [
        0.6666666666666666,
        0.4166666666666667,
        0.5
      ],
      "correctnessFlips": 3,
      "rawValueFlips": 5
    }
  ]
}
```

Raw per-call responses, effective models, billing provenance, usage, per-field outcomes, input hashes, and settings are retained in the individual run JSON files. The canary is excluded.
