Skip to content

Instantly share code, notes, and snippets.

@cybersiddhu
Last active August 18, 2026 17:27
Show Gist options
  • Select an option

  • Save cybersiddhu/2e365a8557a0e55745d84eb41e6a08c7 to your computer and use it in GitHub Desktop.

Select an option

Save cybersiddhu/2e365a8557a0e55745d84eb41e6a08c7 to your computer and use it in GitHub Desktop.
Artificial Analysis Deep-Dive Leaderboards (AA-Briefcase, AA-Omniscience, GDPval-AA v2) — extracted 2026-08-04

Artificial Analysis Deep-Dive Leaderboards (AA-Briefcase, AA-Omniscience, GDPval-AA v2) — extracted 2026-08-18

Artificial Analysis Deep-Dive Leaderboards (Non-Terminal-Bench)

Extracted from Artificial Analysis page HTML on 2026-08-18. Source: Artificial Analysis. Attribution required: https://artificialanalysis.ai/.

Scope note (refresh 2026-08-18). Artificial Analysis removed full deep-dive telemetry from the public page between 2026-08-07 and 2026-08-18. What remains public vs Pro-gated:

  • Full leaderboards (public): AA-Briefcase Elo (62 models), GDPval-AA v2 Elo (203 models).
  • Truncated (public, selector subset only): AA-Omniscience index (31 models; was 461), and all cost/token/turn telemetry (20/31/29 models). Full sets require the Pro API.

This file covers evaluations NOT in the Terminal-Bench v2.1 gist:

  • AA-Briefcase — Agentic knowledge work benchmark (91 tasks)
  • AA-Omniscience — Knowledge & hallucination benchmark (~6,000 questions)
  • GDPval-AA v2 — Real-world economically valuable tasks (Elo anchored to human baseline of 1,000)

AA-Briefcase

AA-Briefcase is a frontier agentic evaluation for long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos. 91 tasks, Terminus 2 agent harness.

AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo. Higher is better.

Table 1: AA-Briefcase Elo Leaderboard

Rank Model Creator Elo CI Release Date
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic 1,714 −11/+12 Jul 2026
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,689 −11/+12 Jul 2026
3 Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic 1,607 −11/+12 Jul 2026
4 Grok 4.6 (high) SpaceXAI 1,578 −11/+12 Aug 2026
5 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic 1,574 −10/+10 Jun 2026
6 Kimi K3 (max) Kimi 1,543 −9/+10 Jul 2026
7 GPT-5.6 Sol (max) OpenAI 1,503 −10/+11 Jul 2026
8 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,469 −10/+10 Jul 2026
9 Qwen3.8 Max Alibaba 1,421 −10/+12 Aug 2026
10 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Anthropic 1,384 −9/+10 Jun 2026
11 Muse Spark 1.2 (xhigh) Meta 1,358 −11/+12 Aug 2026
12 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic 1,340 −9/+8 May 2026
13 Grok 4.5 (high) SpaceXAI 1,313 −9/+10 Jul 2026
14 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,292 −10/+9 Jun 2026
15 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) DeepSeek 1,285 −9/+10 Jul 2026
16 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic 1,276 −8/+9 Apr 2026
17 GLM-5.2 (max) Z AI 1,251 −9/+9 Jun 2026
18 Claude Opus 5 (Adaptive Reasoning, Low Effort) Anthropic 1,225 −10/+10 Jul 2026
19 Claude Sonnet 5 (Adaptive Reasoning, High Effort) Anthropic 1,193 −9/+9 Jun 2026
20 GPT-5.5 (xhigh) OpenAI 1,150 −8/+8 Apr 2026
21 Gemini 3.7 Flash (high) Google 1,131 −11/+11 Aug 2026
22 MiniMax-M3 MiniMax 1,107 −7/+9 Jun 2026
23 GPT-5.5 (high) OpenAI 1,099 −8/+8 Apr 2026
24 Claude Opus 4.7 (Non-reasoning, High Effort) Anthropic 1,083 −8/+8 Apr 2026
25 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) Anthropic 1,075 −8/+8 Feb 2026
26 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,057 −9/+8 Jun 2026
27 GPT-5.5 (medium) OpenAI 1,000 −0/+0 Apr 2026
28 GLM-5.1 (Reasoning) Z AI 971 −8/+8 Apr 2026
29 Gemini 3.6 Flash (high) Google 962 −9/+9 Jul 2026
30 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) Anthropic 931 −9/+8 Jun 2026
31 DeepSeek V4 Pro (Reasoning, Max Effort) DeepSeek 930 −8/+8 Apr 2026
32 Inkling Small Thinking Machines 917 −11/+10 Jul 2026
33 Qwen3.7 Max Alibaba 914 −8/+8 May 2026
34 MiMo-V2.5-Pro Xiaomi 880 −8/+8 Apr 2026
35 Nemotron 3 Ultra 550B A55B (Reasoning) NVIDIA 873 −8/+9 Jun 2026
36 Gemini 3.5 Flash (high) Google 872 −8/+8 May 2026
37 Gemini 3.5 Flash (medium) Google 871 −9/+8 May 2026
38 GPT-5.3 Codex (xhigh) OpenAI 870 −8/+8 Feb 2026
39 Muse Spark 1.1 (xhigh) Meta 869 −10/+11 Jul 2026
40 Inkling (xhigh) Thinking Machines 841 −9/+10 Jul 2026
41 DeepSeek V4 Flash (Reasoning, Max Effort) DeepSeek 833 −9/+8 Apr 2026
42 Kimi K2.6 Kimi 819 −9/+8 Apr 2026
43 Qwen3.6 27B (Reasoning) Alibaba 810 −10/+9 Apr 2026
44 Grok 4.3 (high) SpaceXAI 760 −9/+8 Apr 2026
45 GPT-5.4 mini (xhigh) OpenAI 717 −9/+9 Mar 2026
46 Muse Spark Meta 643 −10/+10 Apr 2026
47 Gemini 3.5 Flash-Lite Google 635 −12/+11 Jul 2026
48 Claude 4.5 Haiku (Reasoning) Anthropic 612 −10/+9 Oct 2025
49 KAT-Coder-Pro V1 KwaiKAT 599 −11/+11 Nov 2025
50 Qwen3.5 397B A17B (Reasoning) Alibaba 554 −11/+10 Feb 2026
51 Mistral Medium 3.5 Mistral 517 −10/+9 Apr 2026
52 Gemini 3.1 Pro Preview Google 458 −12/+10 Feb 2026
53 Gemma 4 31B (Reasoning) Google 374 −12/+11 Apr 2026
54 Command A+ Cohere 369 −15/+13 May 2026
55 North Mini Code Cohere 239 −15/+14 Jun 2026
56 Gemini 3.1 Flash-Lite Google 231 −14/+12 Mar 2026
57 Solar Pro 3 Upstage 138 −14/+12 Apr 2026
58 K2 Think V2 MBZUAI Institute of Foundation Models 60 −15/+13 Dec 2025
59 gpt-oss-120b (high) OpenAI 8 −8/+13 Aug 2025
60 Nemotron 3 Super 120B A12B (Reasoning) NVIDIA 0 −0/+0 Mar 2026
61 gpt-oss-20b (high) OpenAI 0 −0/+0 Aug 2025
62 Llama 4 Maverick Meta 0 −0/+0 Apr 2025

Table 2: AA-Briefcase Cost per Task (USD)

Subset of 20 models (the page's selector). Approximate: computed from `canonicalEvalTokenCounts` × list pricing with cache split, ÷ 91 tasks. Source per-eval cost records (`briefcaseCost.total`) are now Pro-only.

Model Creator Cost/Task
Command A+ Cohere $0.00
Gemini 3.5 Flash-Lite Google $0.12
Claude 4.5 Haiku Anthropic $0.34
gpt-oss-120b (high) OpenAI $0.57
Mistral Medium 3.5 Mistral $0.80
MiniMax-M3 MiniMax $0.94
Inkling Thinking Machines $1.43
Nemotron 3 Ultra NVIDIA $1.61
Muse Spark 1.2 (xhigh) Meta $1.90
Gemini 3.7 Flash (high) Google $2.15
GLM-5.2 (max) Z AI $2.28
Grok 4.6 (high) SpaceXAI $4.42
GPT-5.6 Sol (max) OpenAI $5.32
Qwen3.8 Max Alibaba $5.50
Nemotron 3 Super NVIDIA $6.14
Kimi K3 (max) Kimi $6.73
Claude Opus 5 (high) Anthropic $10.41
Claude Opus 5 (xhigh) Anthropic $14.26
Claude Opus 5 (max) Anthropic $17.79
Claude Fable 5 (with fallback) Anthropic $22.30

Table 3: AA-Briefcase Output Tokens per Task

Subset of 20 models.

Model Creator Reasoning Tokens Answer Tokens Total
Gemini 3.5 Flash-Lite Google 8k 15k 23k
Command A+ Cohere 2k 25k 27k
Claude 4.5 Haiku Anthropic 0k 29k 29k
Mistral Medium 3.5 Mistral 2k 32k 34k
Nemotron 3 Ultra NVIDIA 4k 43k 47k
Inkling Thinking Machines 12k 41k 53k
GPT-5.6 Sol (max) OpenAI 23k 35k 58k
Grok 4.6 (high) SpaceXAI 20k 61k 81k
MiniMax-M3 MiniMax 0k 82k 82k
Muse Spark 1.2 (xhigh) Meta 9k 81k 90k
gpt-oss-120b (high) OpenAI 28k 70k 98k
Qwen3.8 Max Alibaba 51k 49k 100k
Claude Opus 5 (high) Anthropic 37k 77k 114k
GLM-5.2 (max) Z AI 75k 40k 115k
Kimi K3 (max) Kimi 56k 63k 119k
Claude Fable 5 (with fallback) Anthropic 65k 75k 140k
Claude Opus 5 (xhigh) Anthropic 51k 93k 144k
Gemini 3.7 Flash (high) Google 40k 121k 161k
Claude Opus 5 (max) Anthropic 62k 108k 170k
Nemotron 3 Super NVIDIA 336k 268k 604k

Table 4: AA-Briefcase Time per Task (minutes)

Subset of 20 models (only those with median output speed).

Model Creator Min/Task
Gemini 3.5 Flash-Lite Google 2
Mistral Medium 3.5 Mistral 5
Command A+ Cohere 5
Claude 4.5 Haiku Anthropic 6
Nemotron 3 Ultra NVIDIA 7
gpt-oss-120b (high) OpenAI 11
Gemini 3.7 Flash (high) Google 12
Inkling Thinking Machines 14
GLM-5.2 (max) Z AI 16
MiniMax-M3 MiniMax 16
GPT-5.6 Sol (max) OpenAI 19
Claude Fable 5 (with fallback) Anthropic 24
Claude Opus 5 (high) Anthropic 28
Grok 4.6 (high) SpaceXAI 31
Qwen3.8 Max Alibaba 33
Claude Opus 5 (xhigh) Anthropic 34
Claude Opus 5 (max) Anthropic 40
Kimi K3 (max) Kimi 53
Nemotron 3 Super NVIDIA 70

Table 5: AA-Briefcase Mean Turns per Task

Subset of 20 models.

Model Creator Turns
Gemini 3.5 Flash-Lite Google 22
Claude 4.5 Haiku Anthropic 28
Command A+ Cohere 29
Mistral Medium 3.5 Mistral 37
Nemotron 3 Ultra NVIDIA 42
GPT-5.6 Sol (max) OpenAI 50
Grok 4.6 (high) SpaceXAI 53
GLM-5.2 (max) Z AI 56
Muse Spark 1.2 (xhigh) Meta 60
Claude Fable 5 (with fallback) Anthropic 67
Gemini 3.7 Flash (high) Google 69
Qwen3.8 Max Alibaba 70
Claude Opus 5 (high) Anthropic 76
Inkling Thinking Machines 81
Kimi K3 (max) Kimi 83
Claude Opus 5 (xhigh) Anthropic 91
MiniMax-M3 MiniMax 101
Claude Opus 5 (max) Anthropic 103
gpt-oss-120b (high) OpenAI 113
Nemotron 3 Super NVIDIA 414

AA-Omniscience

AA-Omniscience is a knowledge and hallucination benchmark (~6,000 questions) that rewards accuracy, punishes bad guesses, and provides a comprehensive view of which models produce factually reliable outputs across different domains. Scores range from -100 to 100.

Truncated: public page now ships only 31 models' omniscience scores (selector subset); the full ~461-model leaderboard is Pro-only.

Table 6: AA-Omniscience Index Leaderboard

Rank Model Creator Index Accuracy Hallucination Rate
1 Claude Fable 5 (with fallback) Anthropic 43 65% 64%
2 Claude Opus 5 (max) Anthropic 37 61% 61%
3 Claude Opus 5 (xhigh) Anthropic 35 60% 60%
4 Claude Opus 5 (high) Anthropic 34 59% 61%
5 Gemini 3.1 Pro Preview Google 32 55% 51%
6 Grok 4.6 (high) SpaceXAI 30 48% 34%
7 Muse Spark 1.2 (xhigh) Meta 27 45% 33%
8 Gemini 3.7 Flash (high) Google 26 55% 65%
9 GPT-5.6 Sol (max) OpenAI 22 59% 92%
10 Kimi K3 (max) Kimi 20 48% 53%
11 Motif 3 Motif Technologies 10 30% 28%
12 Gemini 3.5 Flash-Lite Google 5 29% 34%
13 GLM-5.2 (max) Z AI 4 24% 26%
14 Qwen3.8 Max Alibaba 3 32% 42%
15 Inkling Thinking Machines 2 42% 68%
16 MiniMax-M3 MiniMax 1 17% 18%
17 DeepSeek V4 Pro 0813 (max) DeepSeek 1 49% 95%
18 GPT-5.6 Terra (max) OpenAI 0 47% 88%
19 Nemotron 3 Ultra NVIDIA 0 23% 30%
20 Solar Open2 250B Upstage -2 19% 25%
21 Command A+ Cohere -4 9% 14%
22 Claude 4.5 Haiku Anthropic -4 18% 27%
23 K-EXAONE 2.0 LG AI Research -7 13% 23%
24 A.X-K2 SK Telecom -8 19% 33%
25 Qwen3.8 27B Alibaba -10 16% 30%
26 GPT-5.6 Luna (max) OpenAI -10 43% 93%
27 Nemotron 3.5 Lightning NVIDIA -18 14% 38%
28 Muse Glimmer (high) Meta -33 27% 82%
29 Mistral Medium 3.5 Mistral -37 25% 82%
30 Nemotron 3 Super NVIDIA -41 24% 87%
31 gpt-oss-120b (high) OpenAI -49 22% 91%

Table 7: AA-Omniscience Domains

Benchmark covers 6 domains with subdomains:

Domain Subdomains
Business Accounting (200), Business & Management, Corporate & Markets (200), Economics (200), Financial Institutions (150), Investments (150)
Health Basic Clinical Science (200), Clinical Medicine (200), Health Systems & Services (150), Medical Ethics & Professionalism (100), Public & Global Health (100)
Humanities & Social Sciences Art & Architecture (100), History (200), Languages & Linguistics (200), Media & Communication (100), Philosophy (200), Political Science (200), Sociology & Anthropology (200)
Law Civil Law (200), Commercial Law (200), Common Law (200), Constitutional Law (150), Criminal Law (200), International Law (150), Legal Systems (100)
Science, Engineering & Mathematics Astronomy & Physics (200), Biology (200), Chemistry (200), Engineering (200), Mathematics (200)
Software Engineering (SWE) Algorithms & Data Structures (200), Architecture & Design Patterns (200), Cloud & Distributed Systems (200), Databases & Information Systems (200), DevOps, Tooling & Build Systems (200), Programming Language Semantics (200), Security & Cryptography (200), Systems Programming & OS (200), Web, Mobile & Interface (200)

Table 8: AA-Omniscience Cost Breakdown (Total USD)

Subset of 31 models. Approximate: `canonicalEvalTokenCounts.omniscience` × list pricing (no cache discount).

Model Creator Total Cost
Qwen3.8 27B Alibaba $0.00
Motif 3 Motif Technologies $0.00
Solar Open2 250B Upstage $0.00
A.X-K2 SK Telecom $0.00
K-EXAONE 2.0 LG AI Research $0.00
Command A+ Cohere $0.00
Nemotron 3.5 Lightning NVIDIA $3.97
gpt-oss-120b (high) OpenAI $7.00
Nemotron 3 Super NVIDIA $11.06
Muse Glimmer (high) Meta $17.06
Gemini 3.5 Flash-Lite Google $18.60
Nemotron 3 Ultra NVIDIA $20.21
MiniMax-M3 MiniMax $20.97
Gemini 3.7 Flash (high) Google $23.55
Claude 4.5 Haiku Anthropic $27.13
GLM-5.2 (max) Z AI $51.54
Grok 4.6 (high) SpaceXAI $51.99
Claude Opus 5 (high) Anthropic $56.40
Muse Spark 1.2 (xhigh) Meta $62.73
GPT-5.6 Luna (max) OpenAI $65.83
DeepSeek V4 Pro 0813 (max) DeepSeek $75.03
Claude Opus 5 (xhigh) Anthropic $85.74
Mistral Medium 3.5 Mistral $86.92
Gemini 3.1 Pro Preview Google $106.03
Inkling Thinking Machines $126.58
Claude Opus 5 (max) Anthropic $155.28
Qwen3.8 Max Alibaba $172.12
GPT-5.6 Terra (max) OpenAI $422.02
Claude Fable 5 (with fallback) Anthropic $427.98
GPT-5.6 Sol (max) OpenAI $767.80
Kimi K3 (max) Kimi $796.95

GDPval-AA v2

GDPval-AA v2 evaluates AI models on real-world, economically valuable tasks across a wide range of occupations. Elo ratings are anchored to a human baseline of 1,000. Higher is better.

Table 9: GDPval-AA v2 Full Leaderboard

Rank Model Creator Elo CI Release Date
1 Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic 1,845 −21/+21 Jul 2026
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,814 −20/+20 Jul 2026
3 Grok 4.6 (high) SpaceXAI 1,747 −20/+20 Aug 2026
4 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic 1,738 −16/+16 Jun 2026
5 Qwen3.8 Max Alibaba 1,735 −17/+17 Aug 2026
6 Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic 1,733 −19/+19 Jul 2026
7 GPT-5.6 Sol (max) OpenAI 1,723 −16/+16 Jul 2026
8 Qwen3.8 2.4T A95B Alibaba 1,720 −22/+22 Aug 2026
9 Kimi K3 (max) Kimi 1,681 −19/+19 Jul 2026
10 GPT-5.6 Sol (xhigh) OpenAI 1,679 −17/+17 Jul 2026
11 Muse Spark 1.2 (xhigh) Meta 1,628 −21/+20 Aug 2026
12 GPT-5.6 Sol (high) OpenAI 1,621 −16/+16 Jul 2026
13 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,620 −19/+19 Jul 2026
14 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Anthropic 1,595 −16/+16 Jun 2026
15 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) DeepSeek 1,590 −21/+21 Aug 2026
16 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic 1,584 −15/+15 May 2026
17 GPT-5.6 Luna (max) OpenAI 1,578 −16/+16 Jul 2026
18 GPT-5.6 Terra (max) OpenAI 1,576 −16/+16 Jul 2026
19 GPT-5.6 Terra (xhigh) OpenAI 1,572 −16/+16 Jul 2026
20 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) DeepSeek 1,559 −19/+19 Jul 2026
21 GPT-5.6 Sol (medium) OpenAI 1,551 −16/+16 Jul 2026
22 Qwen3.8 27B Alibaba 1,546 −22/+22 Aug 2026
23 Gemini 3.7 Flash (high) Google 1,532 −19/+19 Aug 2026
24 GPT-5.6 Luna (xhigh) OpenAI 1,526 −17/+17 Jul 2026
25 Grok 4.5 (high) SpaceXAI 1,524 −18/+18 Jul 2026
26 GPT-5.6 Terra (high) OpenAI 1,512 −16/+16 Jul 2026
27 GLM-5.2 (max) Z AI 1,505 −15/+15 Jun 2026
28 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,502 −16/+16 Jun 2026
29 Gemini 3.7 Flash (medium) Google 1,498 −21/+21 Aug 2026
30 GPT-5.5 (xhigh) OpenAI 1,490 −15/+15 Apr 2026
31 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic 1,490 −15/+15 Apr 2026
32 GPT-5.5 (high) OpenAI 1,466 −15/+15 Apr 2026
33 GPT-5.6 Luna (high) OpenAI 1,466 −16/+16 Jul 2026
34 Claude Opus 5 (Adaptive Reasoning, Low Effort) Anthropic 1,457 −19/+19 Jul 2026
35 Gemini 3.7 Flash (low) Google 1,452 −19/+22 Aug 2026
36 GPT-5.6 Sol (low) OpenAI 1,442 −16/+16 Jul 2026
37 Gemini 3.6 Flash (high) Google 1,422 −17/+17 Jul 2026
38 GPT-5.6 Terra (medium) OpenAI 1,405 −16/+16 Jul 2026
39 Claude Sonnet 5 (Adaptive Reasoning, High Effort) Anthropic 1,401 −16/+16 Jun 2026
40 GPT-5.4 (xhigh) OpenAI 1,394 −15/+15 Mar 2026
41 GLM-5.2 (Non-reasoning) Z AI 1,393 −20/+20 Jun 2026
42 MiniMax-M3 MiniMax 1,387 −15/+15 Jun 2026
43 GPT-5.6 Sol (Non-reasoning) OpenAI 1,379 −17/+17 Jul 2026
44 GPT-5.5 (medium) OpenAI 1,375 −16/+16 Apr 2026
45 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) Anthropic 1,374 −15/+15 Feb 2026
46 Muse Spark 1.1 (xhigh) Meta 1,374 −18/+18 Jul 2026
47 Claude Sonnet 5 (Non-reasoning, High Effort) Anthropic 1,372 −17/+17 Jun 2026
48 Gemini 3.5 Flash (high) Google 1,343 −15/+15 May 2026
49 DeepSeek V4 Pro (Reasoning, Max Effort) DeepSeek 1,306 −15/+15 Apr 2026
50 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,304 −16/+16 Jun 2026
51 DeepSeek V4 Pro (Reasoning, High Effort) DeepSeek 1,291 −20/+20 Apr 2026
52 Solar Pro 4 Upstage 1,276 −20/+20 Aug 2026
53 GPT-5.6 Luna (medium) OpenAI 1,275 −16/+16 Jul 2026
54 Motif 3 Motif Technologies 1,275 −20/+20 Aug 2026
55 Qwen3.7 Max Alibaba 1,271 −15/+15 May 2026
56 Kimi K3 (low) Kimi 1,270 −20/+20 Jul 2026
57 Inkling Small Thinking Machines 1,268 −20/+20 Jul 2026
58 MiMo-V2.5-Pro Xiaomi 1,266 −15/+15 Apr 2026
59 JT-4.1 Flash 236B A21B China Mobile 1,263 −19/+19 Jul 2026
60 GLM-5.1 (Reasoning) Z AI 1,258 −15/+15 Apr 2026
61 Motif 3 (Beta) Motif Technologies 1,257 −19/+19 Jul 2026
62 GPT-5.6 Terra (low) OpenAI 1,255 −16/+16 Jul 2026
63 Nex-N2-Pro Nex AGI 1,249 −17/+17 Jun 2026
64 GPT-5.6 Terra (Non-reasoning) OpenAI 1,245 −17/+17 Jul 2026
65 Inkling (xhigh) Thinking Machines 1,239 −19/+19 Jul 2026
66 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) Anthropic 1,219 −16/+16 Jun 2026
67 Grok Build 0.1 0616 SpaceXAI 1,215 −16/+16 Jun 2026
68 Hy3 Tencent 1,214 −19/+19 Jul 2026
69 G9v3-39A5B AI9Stars 1,197 −20/+20 Aug 2026
70 Kimi K2.6 Kimi 1,190 −15/+15 Apr 2026
71 DeepSeek V4 Flash (Reasoning, Max Effort) DeepSeek 1,190 −16/+16 Apr 2026
72 Kimi K2.7 Code Kimi 1,190 −15/+15 Jun 2026
73 GPT-5.5 (low) OpenAI 1,190 −16/+16 Apr 2026
74 Agnes 2.5 Pro Alpha Sapiens AI 1,176 −20/+20 Jul 2026
75 GPT-5.4 mini (xhigh) OpenAI 1,172 −15/+15 Mar 2026
76 GLM-4.7 (Reasoning) Z AI 1,169 −17/+17 Dec 2025
77 Nemotron 3 Ultra 550B A55B (Reasoning) NVIDIA 1,163 −15/+15 Jun 2026
78 MiniMax-M2.7 MiniMax 1,160 −15/+15 Mar 2026
79 GPT-5.6 Luna (low) OpenAI 1,156 −16/+16 Jul 2026
80 DeepSeek V4 Flash (Reasoning, High Effort) DeepSeek 1,155 −19/+19 Apr 2026
81 MiMo-V2.5 Xiaomi 1,150 −19/+19 Apr 2026
82 Muse Spark Meta 1,146 −15/+15 Apr 2026
83 Qwen3.6 Plus Alibaba 1,140 −15/+15 Apr 2026
84 Qwen3.6 27B (Reasoning) Alibaba 1,140 −15/+15 Apr 2026
85 Gemini 3.5 Flash-Lite Google 1,140 −18/+18 Jul 2026
86 GPT-5.5 (Non-reasoning) OpenAI 1,124 −15/+15 Apr 2026
87 Solar Open2 250B Upstage 1,122 −21/+21 Aug 2026
88 Qwen3.6 27B (Non-reasoning) Alibaba 1,112 −17/+17 Apr 2026
89 A.X-K2 SK Telecom 1,109 −20/+20 Aug 2026
90 Ling 3.0 Flash InclusionAI 1,107 −20/+20 Aug 2026
91 GPT-5.4 nano (xhigh) OpenAI 1,105 −15/+15 Mar 2026
92 Grok 4.3 (Non-reasoning) SpaceXAI 1,099 −15/+15 Apr 2026
93 Grok 4.3 (high) SpaceXAI 1,088 −15/+15 Apr 2026
94 GPT-5 (high) OpenAI 1,083 −18/+18 Aug 2025
95 GPT-5.6 Luna (Non-reasoning) OpenAI 1,074 −17/+17 Jul 2026
96 Claude 4.5 Sonnet (Reasoning) Anthropic 1,056 −18/+18 Sep 2025
97 Qwen3.6 35B A3B (Reasoning) Alibaba 1,055 −15/+15 Apr 2026
98 LongCat 2.0 LongCat 1,030 −18/+18 Jun 2026
99 Qwen3.6 35B A3B (Non-reasoning) Alibaba 1,019 −20/+20 Apr 2026
100 Step 3.7 Flash StepFun 1,018 −15/+15 May 2026
101 Kimi K2.5 (Reasoning) Kimi 1,006 −19/+19 Jan 2026
102 GPT-5.1 (high) OpenAI 993 −18/+18 Nov 2025
103 Qwen3.5 122B A10B (Reasoning) Alibaba 988 −15/+15 Feb 2026
104 K-EXAONE 2.0 0803 LG AI Research 979 −21/+21 Aug 2026
105 Gemini 3.1 Pro Preview Google 965 −16/+16 Feb 2026
106 Qwen3.5 397B A17B (Reasoning) Alibaba 964 −16/+16 Feb 2026
107 Muse Glimmer (high) Meta 953 −22/+20 Aug 2026
108 Qwen3.7 Plus Alibaba 945 −16/+16 Jun 2026
109 GPT-5 mini (high) OpenAI 935 −18/+18 Aug 2025
110 GLM-4.6 (Reasoning) Z AI 934 −18/+18 Sep 2025
111 Mistral Medium 3.5 Mistral 934 −16/+16 Apr 2026
112 Ring-2.6-1T InclusionAI 922 −16/+16 May 2026
113 Claude 4.5 Haiku (Reasoning) Anthropic 913 −16/+16 Oct 2025
114 KAT Coder Pro V2 KwaiKAT 906 −19/+19 Mar 2026
115 KAT-Coder-Pro V1 KwaiKAT 905 −17/+17 Nov 2025
116 Qwen3.5 122B A10B (Non-reasoning) Alibaba 890 −19/+19 Feb 2026
117 DeepSeek V3.1 Terminus (Reasoning) DeepSeek 883 −20/+20 Sep 2025
118 Claude 4 Sonnet (Reasoning) Anthropic 874 −18/+18 May 2025
119 DeepSeek V3.2 (Reasoning) DeepSeek 866 −20/+20 Dec 2025
120 G9v3-3B AI9Stars 865 −27/+27 Jul 2026
121 MiMo-V2-Flash (Non-reasoning) Xiaomi 839 −20/+20 Dec 2025
122 Nemotron 3.5 Lightning NVIDIA 824 −20/+20 Aug 2026
123 Gemma 4 31B (Reasoning) Google 811 −16/+16 Apr 2026
124 gpt-oss-120b (high) OpenAI 800 −16/+16 Aug 2025
125 Qwen3.5 35B A3B (Non-reasoning) Alibaba 796 −21/+21 Feb 2026
126 GPT-5.4 mini (Non-Reasoning) OpenAI 788 −17/+17 Mar 2026
127 Ling 3.0 Tiny InclusionAI 772 −22/+22 Aug 2026
128 Gemma 4 26B A4B (Reasoning) Google 769 −16/+16 Apr 2026
129 Gemma 4 31B (Non-reasoning) Google 747 −19/+19 Apr 2026
130 Devstral 2 Mistral 744 −17/+17 Dec 2025
131 Devstral Small 2 Mistral 732 −17/+17 Dec 2025
132 Command A+ Cohere 717 −19/+19 May 2026
133 Qwen3 Coder Next Alibaba 717 −19/+19 Feb 2026
134 GPT-5.5 Instant (June 2026) OpenAI 715 −18/+18 Jun 2026
135 Nemotron 3 Super 120B A12B (Reasoning) NVIDIA 698 −16/+16 Mar 2026
136 Mercury 2 Inception 698 −20/+20 Feb 2026
137 Nova 2.0 Pro Preview (medium) Amazon 682 −16/+16 Nov 2025
138 EXAONE 4.5 33B LG AI Research 681 −21/+21 Apr 2026
139 Gemini 2.5 Pro Google 669 −18/+18 Jun 2025
140 HyperNova 60B 2605 Multiverse Computing 658 −20/+20 May 2026
141 Nova 2.0 Pro Preview (low) Amazon 653 −17/+17 Nov 2025
142 Gemini 3.1 Flash-Lite Google 648 −16/+16 Mar 2026
143 Gemma 4 12B (Reasoning) Google 645 −21/+21 Jun 2026
144 Qwen3.5 9B (Reasoning) Alibaba 644 −19/+19 Mar 2026
145 Mistral Large 3 Mistral 639 −17/+17 Dec 2025
146 Mistral Medium 3.1 Mistral 610 −18/+18 Aug 2025
147 Mistral Small 3.1 Mistral 602 −18/+18 Mar 2025
148 K-EXAONE (Reasoning) LG AI Research 593 −21/+21 Dec 2025
149 Mistral Small 4 (Reasoning) Mistral 590 −19/+19 Mar 2026
150 Nova 2.0 Lite (high) Amazon 589 −18/+18 Oct 2025
151 gpt-oss-20b (high) OpenAI 567 −17/+17 Aug 2025
152 Trinity Large Thinking Arcee AI 563 −21/+21 Apr 2026
153 Nova 2.0 Pro Preview (Non-reasoning) Amazon 561 −17/+17 Nov 2025
154 DiffusionGemma 26B A4B Google 553 −18/+18 Jun 2026
155 Ling 2.6 Flash InclusionAI 544 −20/+20 Apr 2026
156 Qwen3 235B A22B 2507 (Reasoning) Alibaba 544 −20/+20 Jul 2025
157 North Mini Code Cohere 543 −19/+19 Jun 2026
158 Celeris-1 Celeris 535 −21/+21 Jul 2026
159 DeepSeek R1 (Jan '25) DeepSeek 530 −19/+19 Jan 2025
160 GPT-4.1 mini OpenAI 505 −19/+19 Apr 2025
161 Nemotron Cascade 2 30B A3B NVIDIA 501 −21/+21 Mar 2026
162 Solar Pro 3 Upstage 499 −17/+17 Apr 2026
163 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) NVIDIA 490 −17/+17 Dec 2025
164 Ministral 3 14B Mistral 484 −17/+17 Dec 2025
165 o3-mini (high) OpenAI 471 −20/+20 Jan 2025
166 Nemotron 3 Nano Omni 30B A3B Reasoning NVIDIA 465 −21/+21 Apr 2026
167 Claude 3.5 Haiku Anthropic 457 −18/+18 Oct 2024
168 Ministral 3 8B Mistral 453 −18/+18 Dec 2025
169 Granite 4.1 30B IBM 431 −17/+17 Apr 2026
170 Magistral Medium 1.2 Mistral 414 −19/+19 Sep 2025
171 gpt-oss-120b (low) OpenAI 410 −22/+22 Aug 2025
172 K2 Think V2 MBZUAI Institute of Foundation Models 379 −19/+19 Dec 2025
173 Qwen3 Next 80B A3B (Reasoning) Alibaba 377 −21/+21 Sep 2025
174 DeepSeek V3 0324 DeepSeek 325 −19/+19 Mar 2025
175 Qwen3 30B A3B 2507 (Reasoning) Alibaba 324 −21/+21 Jul 2025
176 Qwen3 32B (Reasoning) Alibaba 289 −19/+19 Apr 2025
177 Ministral 3 3B Mistral 284 −18/+18 Dec 2025
178 Magistral Small 1.2 Mistral 262 −19/+19 Sep 2025
179 GPT-4 OpenAI 239 −19/+19 Mar 2023
180 GPT-4o mini OpenAI 238 −20/+20 Jul 2024
181 Qwen3 14B (Reasoning) Alibaba 236 −19/+19 Apr 2025
182 DeepSeek V3 (Dec '24) DeepSeek 231 −20/+20 Dec 2024
183 Gemma 4 E4B (Reasoning) Google 227 −23/+23 Apr 2026
184 Qwen3.5 2B (Reasoning) Alibaba 217 −22/+22 Mar 2026
185 Qwen3 8B (Reasoning) Alibaba 215 −20/+20 Apr 2025
186 NVIDIA Nemotron 3 Nano 4B NVIDIA 204 −22/+22 Mar 2026
187 Granite 4.1 3B IBM 123 −21/+21 Apr 2026
188 Llama 4 Scout Meta 109 −18/+18 Apr 2025
189 Llama 3.3 Instruct 70B Meta 97 −20/+20 Dec 2024
190 Gemma 4 E2B (Reasoning) Google 84 −22/+22 Apr 2026
191 Mistral Small 3.2 Mistral 78 −20/+20 Jun 2025
192 GPT-4.1 nano OpenAI 61 −20/+20 Apr 2025
193 Llama 4 Maverick Meta 4 −17/+17 Apr 2025
194 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) NVIDIA -72 −17/+17 Dec 2025
195 Qwen3.5 2B (Non-reasoning) Alibaba -81 −19/+19 Mar 2026
196 Qwen3.5 0.8B (Non-reasoning) Alibaba -83 −19/+19 Mar 2026
197 MiniCPM-V 4.6 1.3B OpenBMB -86 −17/+17 May 2026
198 Llama 3.1 Instruct 8B Meta -102 −17/+17 Jul 2024
199 Qwen3.5 0.8B (Reasoning) Alibaba -103 −18/+18 Mar 2026
200 Nanbeige4.1-3B Nanbeige -117 −18/+18 Feb 2026
201 Gemma 3 12B Instruct Google -121 −17/+17 Mar 2025
202 Gemma 3 27B Instruct Google -122 −17/+17 Mar 2025
203 Phi-4 Mini Instruct Microsoft -122 −17/+17 Feb 2024

Table 10: GDPval-AA v2 Cost per Task

Subset of 29 models. Approximate: `canonicalEvalTokenCounts.gdpval` × list pricing with cache split, ÷ 220 tasks.

Model Creator Cost/Task
Qwen3.8 27B Alibaba $0.00
Motif 3 Motif Technologies $0.00
Solar Open2 250B Upstage $0.00
A.X-K2 SK Telecom $0.00
K-EXAONE 2.0 LG AI Research $0.00
Command A+ Cohere $0.00
Muse Glimmer (high) Meta $0.06
Nemotron 3.5 Lightning NVIDIA $0.09
GPT-5.6 Luna (max) OpenAI $0.10
Gemini 3.5 Flash-Lite Google $0.10
gpt-oss-120b (high) OpenAI $0.11
Claude 4.5 Haiku Anthropic $0.19
MiniMax-M3 MiniMax $0.35
Nemotron 3 Super NVIDIA $0.37
Mistral Medium 3.5 Mistral $0.41
Inkling Thinking Machines $0.42
Nemotron 3 Ultra NVIDIA $0.51
DeepSeek V4 Pro 0813 (max) DeepSeek $0.54
Muse Spark 1.2 (xhigh) Meta $0.88
GLM-5.2 (max) Z AI $1.15
GPT-5.6 Terra (max) OpenAI $1.15
Gemini 3.7 Flash (high) Google $1.25
Grok 4.6 (high) SpaceXAI $2.09
Kimi K3 (max) Kimi $2.15
GPT-5.6 Sol (max) OpenAI $3.49
Qwen3.8 Max Alibaba $3.51
Claude Opus 5 (xhigh) Anthropic $4.96
Claude Opus 5 (max) Anthropic $6.77
Claude Fable 5 (with fallback) Anthropic $7.88

Table 11: GDPval-AA v2 Average Turns per Task

Subset of 29 models.

Model Creator Turns
Claude 4.5 Haiku Anthropic 15
Gemini 3.5 Flash-Lite Google 21
Solar Open2 250B Upstage 21
Nemotron 3 Ultra NVIDIA 22
Command A+ Cohere 24
Muse Glimmer (high) Meta 26
K-EXAONE 2.0 LG AI Research 26
Mistral Medium 3.5 Mistral 27
Nemotron 3.5 Lightning NVIDIA 27
gpt-oss-120b (high) OpenAI 29
GLM-5.2 (max) Z AI 31
Claude Fable 5 (with fallback) Anthropic 33
GPT-5.6 Luna (max) OpenAI 35
Grok 4.6 (high) SpaceXAI 37
GPT-5.6 Terra (max) OpenAI 37
Kimi K3 (max) Kimi 38
DeepSeek V4 Pro 0813 (max) DeepSeek 39
Motif 3 Motif Technologies 39
A.X-K2 SK Telecom 42
Inkling Thinking Machines 43
Muse Spark 1.2 (xhigh) Meta 44
GPT-5.6 Sol (max) OpenAI 45
Claude Opus 5 (xhigh) Anthropic 51
Nemotron 3 Super NVIDIA 51
Gemini 3.7 Flash (high) Google 52
MiniMax-M3 MiniMax 54
Qwen3.8 27B Alibaba 58
Claude Opus 5 (max) Anthropic 60
Qwen3.8 Max Alibaba 64

Cross-Evaluation Summary: Top Models Across All Three Benchmarks

Omniscience column is limited to the 31-model public subset; "N/A" means the model's omniscience score is Pro-only.

Model Creator AA-Briefcase Elo Omniscience Index GDPval-AA v2 Elo
Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic 1,714 (#1) 37 (#2) 1,845 (#1)
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,689 (#2) 35 (#3) 1,814 (#2)
Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic 1,607 (#3) 34 (#4) 1,733 (#6)
Grok 4.6 (high) SpaceXAI 1,578 (#4) 30 (#6) 1,747 (#3)
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic 1,574 (#5) 43 (#1) 1,738 (#4)
Kimi K3 (max) Kimi 1,543 (#6) 20 (#10) 1,681 (#9)
GPT-5.6 Sol (max) OpenAI 1,503 (#7) 22 (#9) 1,723 (#7)
Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,469 (#8) N/A 1,620 (#13)
Qwen3.8 Max Alibaba 1,421 (#9) 3 (#14) 1,735 (#5)
Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Anthropic 1,384 (#10) N/A 1,595 (#14)
Muse Spark 1.2 (xhigh) Meta 1,358 (#11) 27 (#7) 1,628 (#11)
Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic 1,340 (#12) N/A 1,584 (#16)
Grok 4.5 (high) SpaceXAI 1,313 (#13) N/A 1,524 (#25)
Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) Anthropic 1,292 (#14) N/A 1,502 (#28)
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) DeepSeek 1,285 (#15) N/A 1,559 (#20)
Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic 1,276 (#16) N/A 1,490 (#31)
GLM-5.2 (max) Z AI 1,251 (#17) 4 (#13) 1,505 (#27)
Claude Opus 5 (Adaptive Reasoning, Low Effort) Anthropic 1,225 (#18) N/A 1,457 (#34)
Claude Sonnet 5 (Adaptive Reasoning, High Effort) Anthropic 1,193 (#19) N/A 1,401 (#39)
GPT-5.5 (xhigh) OpenAI 1,150 (#20) N/A 1,490 (#30)
Gemini 3.7 Flash (high) Google 1,131 (#21) 26 (#8) 1,532 (#23)
MiniMax-M3 MiniMax 1,107 (#22) 1 (#16) 1,387 (#42)
GPT-5.5 (high) OpenAI 1,099 (#23) N/A 1,466 (#32)
Claude Opus 4.7 (Non-reasoning, High Effort) Anthropic 1,083 (#24) N/A N/A
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) Anthropic 1,075 (#25) N/A 1,374 (#45)
Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) Anthropic 1,057 (#26) N/A 1,304 (#50)
GPT-5.5 (medium) OpenAI 1,000 (#27) N/A 1,375 (#44)
GLM-5.1 (Reasoning) Z AI 971 (#28) N/A 1,258 (#60)
Gemini 3.6 Flash (high) Google 962 (#29) N/A 1,422 (#37)
Claude Sonnet 5 (Adaptive Reasoning, Low Effort) Anthropic 931 (#30) N/A 1,219 (#66)

Extracted 2026-08-18 from Artificial Analysis public page HTML. Full AA-Omniscience leaderboard and per-eval cost/token records are Pro-tier only.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment