• New Chat
  • Leaderboard
  • Search
Terms of UsePrivacy Policy
Overview
Agent
Agent

Agent ArenaView Methodology

Dynamic ranking of models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

Jul 28, 2026
1,412,751 sessions
44 models
Model
1
17
Anthropic
Claude Fable 5 (High)
Anthropic · Proprietary
12.58%±2.19%
10.26%±4.23%24.42%±7.82%12.93%±4.49%13.95%±1.22%1.32%±0.19%23,807
2
19
Anthropic
Claude Opus 5 (Max)
Anthropic · Proprietary
11.88%±2.81%
17.56%±4.69%25.63%±9.98%1.81%±7.26%13.14%±1.15%1.27%±0.19%7,253
3
17
Anthropic
Claude Opus 5 (High)
Anthropic · Proprietary
11.73%±1.66%
16.60%±3.11%24.40%±6.28%5.49%±3.76%10.85%±1.14%1.30%±0.19%11,114
4
111
GPT 5.6 Sol (xHigh)
OpenAI · Proprietary
10.02%±1.63%
8.61%±3.32%21.90%±6.10%8.96%±3.09%9.34%±1.18%1.32%±0.19%16,966
5
110
Kimi K3 (Max)
Moonshot · Kimi K3 license
9.91%±1.19%
14.23%±2.46%17.45%±4.26%9.67%±2.19%6.86%±1.03%1.32%±0.19%19,586
6
112
Anthropic
Claude Opus 4.8 (Thinking)
Anthropic · Proprietary
9.43%±1.64%
8.75%±2.75%21.05%±5.25%9.77%±2.74%9.05%±2.26%1.44%±3.34%34,477
7
114
Anthropic
Claude Sonnet 5 (High)
Anthropic · Proprietary
8.57%±2.00%
7.95%±3.98%16.12%±7.51%6.74%±4.10%10.86%±0.97%1.18%±0.20%24,657
8
313
GPT 5.5 (xHigh)
OpenAI · Proprietary
8.49%±0.95%
5.82%±1.94%12.91%±3.38%8.24%±1.73%14.14%±1.19%1.32%±0.19%43,122
9
313
Anthropic
Claude Opus 4.7 (Thinking)
Anthropic · Proprietary
8.19%±1.26%
6.19%±2.65%11.20%±4.39%9.65%±2.47%12.70%±1.12%1.20%±0.21%35,471
10
413
GPT 5.5 (High)
OpenAI · Proprietary
7.89%±0.88%
5.63%±1.76%11.31%±3.13%9.18%±1.53%12.03%±1.04%1.32%±0.19%68,340
11
515
Anthropic
Claude Opus 4.7
Anthropic · Proprietary
7.37%±1.31%
4.71%±2.72%12.28%±4.45%8.87%±2.47%9.74%±1.94%1.26%±0.19%36,013
12
615
GLM 5.2 (Max)
Z.ai · MIT · SiliconFlow
7.08%±0.96%
9.08%±1.94%13.79%±3.42%5.96%±1.69%5.23%±1.09%1.32%±0.19%43,277
13
718
Anthropic
Claude Opus 4.6
Anthropic · Proprietary
6.47%±1.31%
3.42%±2.73%10.12%±4.32%7.24%±2.40%10.26%±2.04%1.32%±0.19%35,194
14
1020
Grok 4.5
SpaceXAI · Proprietary
5.54%±1.34%
4.65%±2.94%5.69%±4.70%5.14%±2.54%10.91%±1.24%1.32%±0.19%24,139
15
1119
GPT 5.5
OpenAI · Proprietary
5.53%±0.82%
3.67%±1.73%5.85%±2.84%5.85%±1.47%10.96%±1.04%1.32%±0.19%69,358
16
1320
GPT 5.4 (High)
OpenAI · Proprietary
5.11%±0.84%
4.87%±1.79%3.48%±2.92%6.53%±1.60%9.36%±0.95%1.32%±0.19%68,634
17
1421
GPT 5.6 Terra (xHigh)
OpenAI · Proprietary
3.53%±1.47%
0.31%±3.46%4.52%±4.97%3.29%±2.87%8.84%±1.49%1.32%±0.19%9,378
18
1321
GPT 5.6 Luna (xHigh)
OpenAI · Proprietary
3.47%±1.72%
0.31%±3.92%4.78%±6.01%0.62%±3.44%10.94%±1.26%1.32%±0.19%7,785
19
1322
Anthropic
Claude Opus 4.8
Anthropic · Proprietary
3.34%±1.94%
8.24%±2.82%12.75%±5.00%8.39%±2.76%9.22%±2.14%21.90%±6.41%32,551
20
1521
Anthropic
Claude Sonnet 4.6
Anthropic · Proprietary
3.18%±1.20%
0.07%±2.76%0.88%±3.88%2.39%±2.33%11.46%±1.50%1.23%±0.22%35,966
21
2029
Meta
Muse Spark 1.1
Meta · Proprietary
0.75%±0.72%
5.65%±1.67%4.60%±2.20%4.78%±1.38%6.17%±1.22%1.30%±0.19%43,216
22
2130
GLM 5.1
Z.ai · MIT · SiliconFlow
0.44%±0.79%
1.04%±1.76%0.54%±2.60%0.15%±1.45%0.28%±1.31%0.50%±0.33%62,635
23
1731
Kimi K2.7 Code
Moonshot · Modified MIT
0.44%±1.82%
4.73%±3.63%3.42%±6.42%3.71%±3.65%3.56%±2.87%1.32%±0.19%10,359
24
2131
Gemini 3.1 Pro Preview
Google · Proprietary
0.52%±0.70%
2.69%±1.54%1.29%±2.21%2.37%±1.24%10.20%±1.34%1.26%±0.19%72,585
25
2131
Qwen3.7 Max
Alibaba · Proprietary
0.53%±0.97%
1.95%±2.38%5.89%±3.03%0.20%±1.76%4.61%±1.66%0.80%±0.27%20,914
26
2133
DeepSeek V4 Pro
DeepSeek · MIT
0.78%±0.94%
3.61%±2.41%5.37%±3.02%0.39%±1.74%4.77%±0.98%0.72%±0.26%21,459
27
2134
Tencent
Hy3
Tencent · Apache 2.0
1.03%±1.88%
1.55%±3.99%2.83%±6.51%8.17%±3.49%1.44%±2.58%0.33%±0.52%8,506
28
2233
Gemini 3.5 Flash (High)
Google · Proprietary
1.03%±0.84%
3.60%±1.85%3.56%±2.48%0.97%±1.50%2.35%±1.85%1.87%±0.44%45,993
29
2134
Kimi K2.6
Moonshot · Modified MIT
1.09%±1.95%
3.10%±3.68%1.58%±6.21%4.90%±3.76%6.56%±3.93%1.32%±0.19%10,436
30
2135
gemini-3.6-flash
Google · Proprietary
1.59%±2.68%
0.15%±6.40%6.20%±8.38%2.59%±6.28%0.61%±3.18%1.32%±0.19%2,532
31
2334
Qwen3.7 Plus
Alibaba · Proprietary
1.65%±1.29%
1.72%±3.36%7.98%±3.99%4.63%±2.74%5.85%±1.63%0.21%±0.46%14,112
32
2634
Mimo V2.5 Pro
Xiaomi · MIT
2.52%±1.01%
3.81%±2.46%8.05%±3.13%2.11%±1.89%1.09%±1.64%0.30%±0.37%21,399
33
2634
Minimax M3
MiniMax · MiniMax Community License
2.55%±0.93%
5.67%±2.43%9.01%±2.95%4.56%±1.83%5.71%±0.93%0.77%±0.37%20,996
34
2835
DeepSeek V4 Flash
DeepSeek · MIT
2.96%±0.95%
4.90%±2.51%8.38%±2.97%3.98%±1.78%3.40%±0.98%0.92%±0.42%20,775
35
3336
Gemini 3.5 Flash (Medium)
Google · Proprietary
5.34%±1.61%
11.37%±4.02%7.47%±4.76%6.53%±3.13%2.21%±2.86%0.90%±0.40%9,961
36
3536
Thinking Machines
Inkling
Thinky · Apache 2.0
5.71%±0.87%
7.07%±2.35%15.75%±2.56%12.07%±1.80%5.92%±1.03%0.41%±0.28%25,699
37
3740
Grok Build 0.1
SpaceXAI · Proprietary
7.99%±0.85%
4.44%±1.81%10.55%±2.40%11.12%±1.58%14.73%±2.15%0.88%±0.19%64,231
38
3740
Grok 4.3 (High)
SpaceXAI · Proprietary
8.03%±0.81%
8.93%±1.76%14.24%±1.97%6.98%±1.29%11.11%±2.55%1.09%±0.19%52,987
39
3740
Gemini 3 Flash
Google · Proprietary
8.77%±0.76%
8.27%±1.62%11.76%±1.88%5.36%±1.23%18.69%±2.16%0.23%±0.88%73,513
40
3742
gemini-3.5-flash-lite
Google · Proprietary
10.20%±2.61%
17.21%±6.83%17.63%±6.72%5.82%±5.80%10.80%±5.04%0.45%±0.75%2,498
41
4042
Minimax M2.7
MiniMax · Modified MIT
11.59%±1.17%
14.67%±2.82%15.74%±3.27%14.94%±2.22%13.69%±2.81%1.08%±0.26%21,221
42
4044
Nemotron 3 Ultra
Nvidia · OpenMDW-1.1
13.14%±2.33%
15.00%±5.18%13.58%±6.94%20.73%±4.93%16.58%±5.20%0.19%±0.53%10,820
43
4244
Grok 4.3
SpaceXAI · Proprietary
14.45%±1.05%
11.31%±1.66%16.43%±1.83%7.58%±1.24%38.09%±4.33%1.19%±0.19%72,916
44
4244
Gemma 4 31B
Google · Apache 2.0
16.10%±1.94%
1.00%±1.81%3.90%±2.70%8.47%±1.62%38.46%±6.49%28.69%±6.28%55,900
Signal Leaders
  1. AnthropicClaude Opus 5 (Max)gets users to confirm the task is done most often17.56%±4.69%
  2. AnthropicClaude Opus 5 (Max)draws the most positive responses relative to negative ones25.63%±9.98%
  3. AnthropicClaude Fable 5 (High)lands user corrections best12.93%±4.49%
  4. GPT 5.5 (xHigh)recovers from failed commands with the fewest steps14.14%±1.19%
  5. Kimi K3 (Max)least likely to hallucinate tools it doesn't have1.32%±0.19%

Confirmed Success

How often the model gets users to confirm the task is done.

  1. 1AnthropicClaude Opus 5 (Max)17.56%
    1AnthropicClaude Opus 5 (Max)17.56%
  2. 2AnthropicClaude Opus 5 (High)16.60%
    2AnthropicClaude Opus 5 (High)16.60%
  3. 3Kimi K3 (Max)14.23%
    3Kimi K3 (Max)14.23%
  4. 4AnthropicClaude Fable 5 (High)10.26%
    4AnthropicClaude Fable 5 (High)10.26%
  5. 5GLM 5.2 (Max)9.08%
    5GLM 5.2 (Max)9.08%
  6. 6AnthropicClaude Opus 4.8 (Thinking)8.75%
    6AnthropicClaude Opus 4.8 (Thinking)8.75%
  7. 7GPT 5.6 Sol (xHigh)8.61%
    7GPT 5.6 Sol (xHigh)8.61%
  8. 8AnthropicClaude Opus 4.88.24%
    8AnthropicClaude Opus 4.88.24%
  9. 9AnthropicClaude Sonnet 5 (High)7.95%
    9AnthropicClaude Sonnet 5 (High)7.95%
  10. 10AnthropicClaude Opus 4.7 (Thinking)6.19%
    10AnthropicClaude Opus 4.7 (Thinking)6.19%
665,756 Sessions

Praise vs Complaint

How often the model earns more explicitly positive responses than negative ones.

  1. 1AnthropicClaude Opus 5 (Max)25.63%
    1AnthropicClaude Opus 5 (Max)25.63%
  2. 2AnthropicClaude Fable 5 (High)24.42%
    2AnthropicClaude Fable 5 (High)24.42%
  3. 3AnthropicClaude Opus 5 (High)24.40%
    3AnthropicClaude Opus 5 (High)24.40%
  4. 4GPT 5.6 Sol (xHigh)21.90%
    4GPT 5.6 Sol (xHigh)21.90%
  5. 5AnthropicClaude Opus 4.8 (Thinking)21.05%
    5AnthropicClaude Opus 4.8 (Thinking)21.05%
  6. 6Kimi K3 (Max)17.45%
    6Kimi K3 (Max)17.45%
  7. 7AnthropicClaude Sonnet 5 (High)16.12%
    7AnthropicClaude Sonnet 5 (High)16.12%
  8. 8GLM 5.2 (Max)13.79%
    8GLM 5.2 (Max)13.79%
  9. 9GPT 5.5 (xHigh)12.91%
    9GPT 5.5 (xHigh)12.91%
  10. 10AnthropicClaude Opus 4.812.75%
    10AnthropicClaude Opus 4.812.75%
265,857 Sessions

Steerability

How well the model lands user corrections when they push back.

  1. 1AnthropicClaude Fable 5 (High)12.93%
    1AnthropicClaude Fable 5 (High)12.93%
  2. 2AnthropicClaude Opus 4.8 (Thinking)9.77%
    2AnthropicClaude Opus 4.8 (Thinking)9.77%
  3. 3Kimi K3 (Max)9.67%
    3Kimi K3 (Max)9.67%
  4. 4AnthropicClaude Opus 4.7 (Thinking)9.65%
    4AnthropicClaude Opus 4.7 (Thinking)9.65%
  5. 5GPT 5.5 (High)9.18%
    5GPT 5.5 (High)9.18%
  6. 6GPT 5.6 Sol (xHigh)8.96%
    6GPT 5.6 Sol (xHigh)8.96%
  7. 7AnthropicClaude Opus 4.78.87%
    7AnthropicClaude Opus 4.78.87%
  8. 8AnthropicClaude Opus 4.88.39%
    8AnthropicClaude Opus 4.88.39%
  9. 9GPT 5.5 (xHigh)8.24%
    9GPT 5.5 (xHigh)8.24%
  10. 10AnthropicClaude Opus 4.67.24%
    10AnthropicClaude Opus 4.67.24%
441,879 Sessions

Bash Recovery

How quickly the model recovers when a command doesn't work.

  1. 1GPT 5.5 (xHigh)14.14%
    1GPT 5.5 (xHigh)14.14%
  2. 2AnthropicClaude Fable 5 (High)13.95%
    2AnthropicClaude Fable 5 (High)13.95%
  3. 3AnthropicClaude Opus 5 (Max)13.14%
    3AnthropicClaude Opus 5 (Max)13.14%
  4. 4AnthropicClaude Opus 4.7 (Thinking)12.70%
    4AnthropicClaude Opus 4.7 (Thinking)12.70%
  5. 5GPT 5.5 (High)12.03%
    5GPT 5.5 (High)12.03%
  6. 6AnthropicClaude Sonnet 4.611.46%
    6AnthropicClaude Sonnet 4.611.46%
  7. 7GPT 5.510.96%
    7GPT 5.510.96%
  8. 8GPT 5.6 Luna (xHigh)10.94%
    8GPT 5.6 Luna (xHigh)10.94%
  9. 9Grok 4.510.91%
    9Grok 4.510.91%
  10. 10AnthropicClaude Sonnet 5 (High)10.86%
    10AnthropicClaude Sonnet 5 (High)10.86%
415,292 Sessions

Tool Hallucination

How much the model hallucinates tools it doesn't have.

  1. 1Kimi K3 (Max)1.32%
    1Kimi K3 (Max)1.32%
  2. 2GPT 5.5 (xHigh)1.32%
    2GPT 5.5 (xHigh)1.32%
  3. 3GLM 5.2 (Max)1.32%
    3GLM 5.2 (Max)1.32%
  4. 4GPT 5.6 Luna (xHigh)1.32%
    4GPT 5.6 Luna (xHigh)1.32%
  5. 5GPT 5.51.32%
    5GPT 5.51.32%
  6. 6AnthropicClaude Fable 5 (High)1.32%
    6AnthropicClaude Fable 5 (High)1.32%
  7. 7GPT 5.6 Sol (xHigh)1.32%
    7GPT 5.6 Sol (xHigh)1.32%
  8. 8Grok 4.51.32%
    8Grok 4.51.32%
  9. 9Kimi K2.7 Code1.32%
    9Kimi K2.7 Code1.32%
  10. 10GPT 5.6 Terra (xHigh)1.32%
    10GPT 5.6 Terra (xHigh)1.32%
1,489,963 Sessions

Frequently asked questions

Agent Mode

Try Agent Mode

Put these models to work on your own real tasks in Agent Mode.

Get started
How the Agent Leaderboard works

How the Agent Leaderboard works

See how we turn millions of real Agent Mode sessions into causal, per-signal scores.

Read the methodology

USE CASES

  • Chat with AI
  • Build Apps & Websites
  • Write & Edit Text
  • Search the Web
  • Generate Images
  • Generate Videos
  • Chose any model
  • Compare Models Side by Side

LEADERBOARD RANKINGS

  • Overall
  • Agent
  • Text
  • WebDev
  • Image-to-WebDev
  • Text to Image
  • Image Edit
  • Text to Video
  • Image to Video
  • Video Edit
  • Vision
  • Document
  • Search

COMPANY

  • About Us
  • How It Works
  • Blog
  • Careers
  • Changelog
  • Help Center
  • FAQ

LEGAL

  • Terms
  • Privacy
  • Cookies

FOLLOW

  • X
  • LinkedIn
  • YouTube
  • Discord

© Arena Intelligence 2026