Skip to content

Instantly share code, notes, and snippets.

@cybersiddhu
Last active August 18, 2026 16:43
Show Gist options
  • Select an option

  • Save cybersiddhu/28c47a2280c81eade339a634d1b2156f to your computer and use it in GitHub Desktop.

Select an option

Save cybersiddhu/28c47a2280c81eade339a634d1b2156f to your computer and use it in GitHub Desktop.
LiveBench Contamination-Free Benchmark Leaderboard (2026-06-25 Release) — extracted 2026-08-04

LiveBench Contamination-Free Benchmark Leaderboard (2026-06-25 Release) — extracted 2026-08-18

LiveBench Deep-Dive Leaderboard (2026-06-25 Release)

Extracted from LiveBench.ai on 2026-08-18. Release version: LiveBench-2026-06-25.

LiveBench is a contamination-free LLM benchmark featuring 23 objective tasks across 7 categories, with questions updated from recent sources (math competitions, arXiv papers, Kaggle/Socrata datasets, news, movie synopses) and scored automatically against objective ground-truth without LLM judges.


LiveBench 7-Category Overall Leaderboard

Scores are 0–100 percentages per category. Cost per successful task = (Σ cost ÷ Σ questions ÷ score) × 100 over selected scope.

Rank Model Creator Overall Reasoning Coding Agentic Coding Math Data Analysis Language Instruction Following Cost/Task ($)
1 Claude Fable 5 Max Effort Anthropic 83.0 89.7 86.0 62.2 96.0 80.5 90.7 75.8 $1.439
2 GPT-5.6 Sol Max Effort OpenAI 81.0 91.7 83.9 56.2 96.2 79.8 87.7 71.8 $0.515
3 GPT-5.5 Thinking xHigh Effort OpenAI 80.2 89.7 82.1 54.0 95.9 81.6 87.4 70.7 $0.435
4 Claude 5 Opus Thinking Max Effort Anthropic 80.1 91.2 81.4 65.2 95.7 74.6 88.7 63.8 $0.699
5 Smaug-Agentic 🔓 Abacus.AI 79.5 90.3 82.5 64.6 83.9 79.9 84.4 71.0 $0.329
6 Kimi K3 🔓 Moonshot AI 79.2 90.7 81.4 62.2 84.4 78.7 85.5 71.4 $0.348
7 Gemini 3.7 Flash High Google 78.8 87.8 78.9 58.3 93.5 68.0 85.5 79.9 $0.157
8 Qwen 3.8 Max 🔓 Alibaba 78.5 88.2 72.9 64.6 91.3 78.4 79.7 74.1 $0.275
9 Grok 4.6 xAI 78.0 90.5 76.8 57.0 92.6 73.9 83.7 71.9 $0.207
10 GPT-5.4 Thinking xHigh Effort OpenAI 78.0 88.1 77.5 53.8 94.1 79.3 82.6 70.2 $0.387
11 Muse Spark 1.2 xHigh Effort Meta 78.0 90.0 77.5 57.6 91.2 76.5 78.6 74.3 $0.375
12 GPT-5.6 Terra Max Effort OpenAI 77.9 90.6 78.2 54.9 94.9 79.3 82.9 64.6 $0.352
13 DeepSeek V4 Pro 0813 🔓 DeepSeek 77.4 85.8 77.2 54.9 95.1 79.2 82.1 67.7 $0.044
14 Gemini 3.1 Pro Preview High Google 77.0 84.0 76.5 44.1 91.0 78.5 85.4 79.1 $0.286
15 Claude 4.7 Opus Thinking xHigh Effort Anthropic 76.5 87.2 82.1 50.7 92.9 78.3 77.9 66.7 $0.528
16 Claude 4.8 Opus Thinking Max Effort Anthropic 76.2 89.2 81.8 50.5 94.3 66.0 79.7 72.0 $0.983
17 Claude Sonnet 5 xHigh Effort Anthropic 76.0 88.7 80.7 59.4 92.9 71.7 75.0 63.9 $0.505
18 Grok 4.5 xAI 75.8 87.2 68.6 56.5 90.8 73.0 82.8 71.5 $0.131
19 Muse Spark 1.1 xHigh Effort Meta 75.3 87.7 77.2 58.5 87.1 72.5 74.3 69.6 $0.198
20 Gemini 3.5 Flash High Google 74.6 82.0 78.2 49.0 88.2 64.9 84.6 75.6 $0.249
21 GPT-5.2 High OpenAI 74.6 83.2 76.1 50.3 93.2 78.2 79.8 61.8 $0.234
22 Claude 4.6 Opus Thinking High Effort Anthropic 74.5 88.7 78.2 49.0 89.3 69.9 83.3 63.3 $0.404
23 DeepSeek V4 Flash 0731 🔓 DeepSeek 74.2 86.6 75.0 46.8 86.8 79.3 79.2 65.5 $0.060
24 GPT-5.2 Codex OpenAI 74.0 77.7 83.6 49.4 88.8 78.2 73.7 66.4 $0.187
25 Gemini 3.6 Flash High Google 73.6 85.1 77.9 43.4 86.4 63.0 83.9 75.4 $0.235
26 GPT-5.6 Luna Max Effort OpenAI 73.6 85.6 82.9 48.4 87.2 78.0 72.6 60.1 $0.169
27 GLM-5.2 🔓 Z.AI 73.2 78.6 79.7 51.8 89.8 73.7 76.2 62.3 $0.225
28 Qwen 3.7 Max Alibaba 73.1 83.3 74.2 43.6 85.2 71.8 79.7 74.0 $0.182
29 Claude 4.6 Sonnet Thinking Medium Effort Anthropic 73.0 84.8 79.3 42.6 87.0 77.9 76.1 63.2 $0.306
30 Claude 4.5 Opus Thinking High Effort Anthropic 72.6 80.1 79.7 39.7 90.4 74.4 81.3 62.5 $0.610
31 Inkling xHigh Effort Thinking Machines 71.9 78.3 71.0 49.4 88.4 72.8 73.5 70.1 $0.310
32 DeepSeek V4 Pro 🔓 DeepSeek 71.6 82.7 70.0 42.6 90.7 74.5 78.1 62.4 $0.050
33 Kimi K2.6 Thinking 🔓 Moonshot AI 70.5 79.4 78.6 46.9 84.3 65.1 75.1 64.4 $0.169
34 GPT-5.4 Nano xHigh OpenAI 69.6 81.1 70.8 46.8 91.0 67.6 62.5 67.2 $0.091
35 Qwen 3.6 Plus Alibaba 68.9 75.8 78.2 41.4 83.7 69.9 75.0 58.3 $0.227
36 Kimi K2.7 Code 🔓 Moonshot AI 68.4 82.8 74.0 45.7 79.6 62.7 77.9 56.3 $0.100
37 Grok Build 0.1 xAI 67.8 76.4 65.4 45.8 78.4 70.8 72.5 65.2 $0.024
38 Minimax M3 Minimax 67.3 74.5 68.2 40.7 76.9 76.2 76.8 57.5 $0.060
39 GPT-5.4 Mini xHigh OpenAI 66.4 71.3 71.6 41.7 78.5 70.8 71.0 59.8 $0.334
40 DeepSeek V4 Flash 🔓 DeepSeek 65.5 70.6 69.2 37.6 79.6 68.0 70.1 63.1 $0.016
41 Qwen 3.6 27B 🔓 Alibaba 64.0 70.3 71.8 39.3 79.9 70.4 63.3 53.2 $0.202
42 Gemini 3.5 Flash-Lite High Google 63.9 60.2 76.1 45.3 73.7 53.2 71.8 67.2 $0.069
43 Grok 4.3 xAI 62.3 70.8 69.9 18.5 84.3 55.8 73.6 62.8 $0.061
44 Qwen3.8 27B 🔓 Alibaba 59.0 65.8 65.3 61.4 72.8 42.7 48.1 57.0 $0.109

Category Leaders

  • Reasoning: GPT-5.6 Sol Max Effort (91.7%) / Claude 5 Opus Thinking Max Effort (91.2%)
  • Coding: Claude Fable 5 Max Effort (86.0%) / GPT-5.6 Sol Max Effort (83.9%)
  • Agentic Coding: Claude 5 Opus Thinking Max Effort (65.2%) / Smaug-Agentic (64.6%) / Qwen 3.8 Max (64.6%)
  • Mathematics: GPT-5.6 Sol Max Effort (96.2%) / Claude Fable 5 Max Effort (96.0%)
  • Data Analysis: GPT-5.5 Thinking xHigh Effort (81.6%) / Claude Fable 5 Max Effort (80.5%)
  • Language: Claude Fable 5 Max Effort (90.7%) / Claude 5 Opus Thinking Max Effort (88.7%)
  • Instruction Following: Gemini 3.7 Flash High (79.9%) / Gemini 3.1 Pro Preview High (79.1%)

Efficiency & Cost Rankings

Top budget-to-performance frontier models:

  • Ultra-Budget ($0.016 – $0.050 / task):
    • DeepSeek V4 Flash ($0.016 / task, 65.5% overall)
    • Grok Build 0.1 ($0.024 / task, 67.8% overall)
    • DeepSeek V4 Pro 0813 ($0.044 / task, 77.4% overall)
    • DeepSeek V4 Pro ($0.050 / task, 71.6% overall)
  • Budget ($0.051 – $0.100 / task):
    • DeepSeek V4 Flash 0731 ($0.060 / task, 74.2% overall)
    • Minimax M3 ($0.060 / task, 67.3% overall)
    • Grok 4.3 ($0.061 / task, 62.3% overall)
    • Gemini 3.5 Flash-Lite High ($0.069 / task, 63.9% overall)
    • GPT-5.4 Nano xHigh ($0.091 / task, 69.6% overall)
    • Kimi K2.7 Code ($0.100 / task, 68.4% overall)
  • Mid-Tier Value ($0.101 – $0.200 / task):
    • Qwen3.8 27B ($0.109 / task, 59.0% overall)
    • Grok 4.5 ($0.131 / task, 75.8% overall)
    • Gemini 3.7 Flash High ($0.157 / task, 78.8% overall)
    • GPT-5.6 Luna Max Effort ($0.169 / task, 73.6% overall)
    • Kimi K2.6 Thinking ($0.169 / task, 70.5% overall)
    • Qwen 3.7 Max ($0.182 / task, 73.1% overall)
    • GPT-5.2 Codex ($0.187 / task, 74.0% overall)
    • Muse Spark 1.1 xHigh Effort ($0.198 / task, 75.3% overall)

Extracted 2026-08-18 from LiveBench.ai.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment