AI Benchmark

Compare leading AI models across performance, speed, cost, and capability

Updated September 30, 2026 · 11 models · 15 benchmarks

11

Models evaluated

15

Benchmarks covered

69.7

Highest overall · Astra

$0.15/1M in

Lowest cost · DS V4.1, Qwen Next

60.3

Average benchmark score

Overall benchmark score

Mean of each model's recorded % scores. Ranked best first — hover any bar for details.

Benchmark breakdown

Click + beside a model to add it to the comparison below (max 4), or click a column header to sort (missing values always sort last). O OpenAI table · A Anthropic table · I independent

Model comparison

Pick 2–4 models. Cards, chart, and share link update together.

Shared benchmarks, side by side

Performance vs cost

Overall score against list input price (log scale). Cheaper is right-to-left; stronger is bottom-to-top. Hover any point.

All models

GPT-6 Astra #1

OpenAI · Sep 3, 2026

69.7

Best result
ExploitBench 100%
Scored
14 tests
In / 1M
$10
Context
1.05M

Claude Opus 5.5 #3

Anthropic · Sep 22, 2026

61

Best result
OSWorld 2.0 (partial) 81.8%
Scored
9 tests
In / 1M
$4
Context
1M

Claude Fable 5.1 #2

Anthropic · Sep 1, 2026

61.8

Best result
GPQA Diamond 93.7%
Scored
11 tests
In / 1M
$10
Context
1M

Gemini 3.8 Flash #4

Google · Sep 2, 2026

58

Best result
GPQA Diamond 95.3%
Scored
4 tests
In / 1M
$0.75
Context
1M

GPT-6 Sol #6

OpenAI · Sep 22, 2026

51.7

Best result
DeepSWE v1.1 68.8%
Scored
7 tests
In / 1M
$2
Context
1.05M

GPT-6.1 Sol soon

OpenAI · Sep 29, 2026

—

Best result
—
Scored
0 tests
In / 1M
~$2
Context
—

Claude Opus 5 #5

Anthropic · Jul 2026

54

Best result
GPQA Diamond 93.7%
Scored
13 tests
In / 1M
$5
Context
—

DeepSeek V4.1 Flash —

DeepSeek · Sep 10, 2026

—

Best result
DeepSWE v1.1 74.2%
Scored
2 tests
In / 1M
$0.15–0.30
Context
1M

Muse Spark 1.3 —

Meta · Sep 2, 2026

—

Best result
DeepSWE v1.1 75.4%
Scored
1 tests
In / 1M
$1.25
Context
—

Grok 4.7 —

xAI · Sep 21, 2026

—

Best result
—
Scored
1 tests
In / 1M
$2
Context
500K

Qwen3.8-Flash-Next soon

Alibaba · Aug 2026 (preview)

—

Best result
—
Scored
0 tests
In / 1M
$0.15
Context
262K–1M
Methodology & data
  • What: 11 frontier models × 15 benchmarks, last updated September 30, 2026.
  • Sources: OpenAI's GPT-6 Astra launch table (Sep 3, 2026), Anthropic's Claude Opus 5.5 launch table (Sep 22, 2026), and independent measurements (ARC Prize, Artificial Analysis snapshots, tracked aggregators). Every table cell is tagged O, A, or I.
  • Overall score: the mean of a model's recorded % scores, shown only with 3+ scored tests (a 1-test mean is not comparable to a 12-test mean). Elo (GDPval) and index points (AA Index) are different scales and are excluded — they still appear as their own columns.
  • Rank: sorted by overall, best first. Models below the 3-test bar or with no published scores (GPT-6.1 Sol, Qwen3.8-Flash-Next) are tracked but unranked.
  • Cost: list input price per million tokens. "~" means the source figure was approximate; ranges use the midpoint. "—" means unpublished.
  • Key limitation: OpenAI and Anthropic test on different harnesses with different effort settings — the same model scores 37.3% and 40% on Terminal-Bench 4.0 across the two tables. Vendor cells are each lab's best case.
  • Not tracked: inference speed (no verified per-model figures in the dataset) and historical trends (no time series exists) — both are omitted rather than estimated.
  • Corrections: same rule as the archive — real sources only, no invented numbers.