llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

LLMBoard ranking

Agent Model Leaderboard

Compare the current agent model ranking from the LLMBoard Agent Score, eligible agent benchmarks, official token pricing and supporting task evidence.

Data as of 2026-09-08

On this page

  • Ranking
  • Overview
  • Agent score
  • Composition
  • Top models
  • FAQ

Agent Model Ranking

This leaderboard ranking orders models by their domain-specific LLMBoard score aggregated from eligible benchmark evidence. Pricing, coverage and the source signal remain separate.

30 of 49 rows
Columns

Show columns

Sort by
Rank
Model
LLMBoard Agent Score
Official input / 1M
Official output / 1M
Score updated
Rank01ModelANClaude Opus 5AnthropicLLMBoard Agent Score86.35Official input / 1M$5Official output / 1M$25Score updatedSep 8, 2026
Rank02ModelANClaude Fable 5.1AnthropicLLMBoard Agent Score85.92Official input / 1M$10Official output / 1M$50Score updatedSep 8, 2026
Rank03ModelOPGPT-5.6-SolOpenAILLMBoard Agent Score82.84Official input / 1M$4Official output / 1M$20Score updatedSep 8, 2026
Rank04ModelANClaude Fable 5AnthropicLLMBoard Agent Score82.04Official input / 1M$10Official output / 1M$50Score updatedSep 8, 2026
Rank05ModelMAKimi K3Moonshot AILLMBoard Agent Score78.26Official input / 1M$3Official output / 1M$15Score updatedSep 8, 2026
Rank06ModelOPGPT-5.5OpenAILLMBoard Agent Score77.93Official input / 1M$5Official output / 1M$30Score updatedSep 8, 2026
Rank07ModelDEDeepSeek-V4-ProDeepSeekLLMBoard Agent Score75.09Official input / 1MN/AOfficial output / 1MN/AScore updatedSep 8, 2026
Rank08ModelZAGLM 5.2Zhipu AILLMBoard Agent Score73.87Official input / 1M$1.4Official output / 1M$4.4Score updatedSep 8, 2026
Rank09ModelXAGrok 4.5xAILLMBoard Agent Score72.64Official input / 1M$2Official output / 1M$6Score updatedSep 8, 2026
Rank10ModelXAGrok 4.6xAILLMBoard Agent Score72.21Official input / 1M$2Official output / 1M$6Score updatedSep 8, 2026
Rank11ModelANClaude Sonnet 5AnthropicLLMBoard Agent Score70.70Official input / 1M$2Official output / 1M$10Score updatedSep 8, 2026
Rank12ModelANClaude Opus 4.7AnthropicLLMBoard Agent Score70.42Official input / 1M$5Official output / 1M$25Score updatedSep 8, 2026
Rank13ModelANClaude Opus 4.8AnthropicLLMBoard Agent Score70.09Official input / 1M$5Official output / 1M$25Score updatedSep 8, 2026
Rank14ModelTEHy4TencentLLMBoard Agent Score69.85Official input / 1MN/AOfficial output / 1MN/AScore updatedSep 8, 2026
Rank15ModelZAGLM 5.3Zhipu AILLMBoard Agent Score66.97Official input / 1M$1.4Official output / 1M$4.4Score updatedSep 8, 2026
Rank16ModelANClaude Opus 4.6AnthropicLLMBoard Agent Score66.59Official input / 1M$5Official output / 1M$25Score updatedSep 8, 2026
Rank17ModelZAGLM 5.3 FlashZhipu AILLMBoard Agent Score65.64Official input / 1M$0.075Official output / 1M$0.25Score updatedSep 8, 2026
Rank18ModelOPGPT-5.6-TerraOpenAILLMBoard Agent Score64.79Official input / 1M$2Official output / 1M$12Score updatedSep 8, 2026
Rank19ModelGOGemini 3.8 FlashGoogleLLMBoard Agent Score62.62Official input / 1M$0.75Official output / 1M$3.75Score updatedSep 8, 2026
Rank20ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamLLMBoard Agent Score60.73Official input / 1M$2Official output / 1M$6Score updatedSep 8, 2026
Rank21ModelOPGPT-5.4OpenAILLMBoard Agent Score59.78Official input / 1M$2.5Official output / 1M$15Score updatedSep 8, 2026
Rank22ModelDEDeepSeek-V4-FlashDeepSeekLLMBoard Agent Score59.12Official input / 1M$0.14Official output / 1M$0.28Score updatedSep 8, 2026
Rank23ModelOPGPT-5.6-LunaOpenAILLMBoard Agent Score58.32Official input / 1M$0.20Official output / 1M$1.2Score updatedSep 8, 2026
Rank24ModelMAKimi K2.7 CodeMoonshot AILLMBoard Agent Score56.33Official input / 1M$0.95Official output / 1M$4Score updatedSep 8, 2026
Rank25ModelMAKimi K2.6Moonshot AILLMBoard Agent Score52.55Official input / 1M$0.95Official output / 1M$4Score updatedSep 8, 2026
Rank26ModelMEMuse Spark 1.2MetaLLMBoard Agent Score47.73Official input / 1M$1.25Official output / 1M$4.25Score updatedSep 8, 2026
Rank27ModelACQwen3.8 Flash NextAlibaba Cloud / Qwen TeamLLMBoard Agent Score47.68Official input / 1MN/AOfficial output / 1MN/AScore updatedSep 8, 2026
Rank28ModelANClaude Sonnet 4.6AnthropicLLMBoard Agent Score47.07Official input / 1M$3Official output / 1M$15Score updatedSep 8, 2026
Rank29ModelACQwen3.8 27BAlibaba Cloud / Qwen TeamLLMBoard Agent Score45.70Official input / 1MN/AOfficial output / 1MN/AScore updatedSep 8, 2026
Rank30ModelGOGemini 3.7 FlashGoogleLLMBoard Agent Score45.56Official input / 1M$0.75Official output / 1M$3.75Score updatedSep 8, 2026

Leaderboard ranking overview

This leaderboard overview connects current LLMBoard leaders with benchmark coverage and supporting source activity.

Ranked models49
SignalLLMBoard Agent Score

Agent score and confidence

Use confidence and observed task volume to judge how firmly each leaderboard position is supported; these source proportions remain separate from benchmark coverage.

Claude Fable 5.1Anthropic
15.87%
Claude Opus 5Anthropic
12.68%
Claude Fable 5Anthropic
10.23%
GPT-5.6-SolOpenAI
9.31%
Claude Opus 4.8Anthropic
9.17%
Kimi K3Moonshot AI
7.96%
Claude Sonnet 5Anthropic
7.41%
GPT-5.5OpenAI
7.21%
Hy4Tencent
6.92%
Claude Opus 4.7Anthropic
6.03%
GLM 5.2Zhipu AI
6.03%
Grok 4.5xAI
5.73%
4.54%Source score with confidence interval18.71%

Agent Model Leaderboard Vendor Mix

Vendor representation among the top-ranked models in this leaderboard.

Top-entry vendor mix

25 models
Anthropic7
OpenAI5
Moonshot AI3
Zhipu AI3
DeepSeek2
xAI2
Tencent1
Google1

The Top AI Models for Agent

A practical summary of the first five entries in this agent AI model leaderboard, with benchmark evidence and pricing kept in context.

About this leaderboard ranking

Common questions about the Agent Model Leaderboard.

What is the top model on Agent Model Leaderboard?

Claude Opus 5 is currently ranked first with 86.35 LLMBoard Agent Score.

What are the top three models on Agent Model Leaderboard?

The current leaders are Claude Fable 5.1 (rank #1), Claude Opus 5 (rank #2), and Claude Fable 5 (rank #3).

Which ranked model has the lowest official input price?

GLM 5.3 Flash has the lowest matched official input price at $0.075 per 1M tokens.

Which models are fastest on this leaderboard?

The fastest matched records are GLM 5.3 (472.69 tok/s via FriendliAI), Gemini 3.5 Flash (210.94 tok/s via Google), and GPT-5.5 (134.94 tok/s via OpenAI).

Does the top source result mean the model is best at every task?

No. Source Arena results describe a particular task or preference signal. Capability benchmarks, prices and runtime can produce different rankings.

0.00%4.75%9.50%14.25%19.00%123.5K381.1K1.2M3.6M11.2MSource Agent scoreObservations
49 models with comparable dataOpen a model by selecting its logo
How does this leaderboard use benchmarks?

The LLMBoard score normalizes eligible evidence by task dimension, applies domain weights and reports coverage separately. The ranking follows that score.

How many models are shown?

This page currently compares 49 models.

Why can a model be missing?

A model appears after LLMBoard calculates a score for this domain. A source Arena result is optional supporting evidence.

Top LLMBoard modelClaude Opus 586.35 score ยท 100.00% coverage
Closest challengerClaude Fable 5.10.43 LLMBoard points behind
Most observationsKimi K38205381 observations

Ranking basisThis agent AI model leaderboard uses the domain-specific LLMBoard score, evidence coverage and status. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    AN
    Claude Opus 5Anthropic
    LLMBoard Agent Score
    86.35
    Price
    $5 input / $25 output per 1M tokens
    Speed
    Up to 62.88 tok/s via Anthropic

    Strengths

    • Ranks #1 on the current agent LLMBoard ranking
    • 2,700,251 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  2. 02
    AN
    Claude Fable 5.1Anthropic
    LLMBoard Agent Score
    85.92
    Price
    $10 input / $50 output per 1M tokens
    Speed
    Up to 7.67 tok/s via Anthropic

    Strengths

    • Ranks #2 on the current agent LLMBoard ranking
    • 502,509 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  3. 03
    OP
    OpenAI
    LLMBoard Agent Score
    82.84
    Price
    $4 input / $20 output per 1M tokens
    Speed
    Up to 2.36 tok/s via OpenAI

    Strengths

    • Ranks #3 on the current agent LLMBoard ranking
    • 3,130,094 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  4. 04
    AN
    Anthropic
    LLMBoard Agent Score
    82.04
    Price
    $10 input / $50 output per 1M tokens
    Speed
    Up to 43.18 tok/s via Anthropic

    Strengths

    • Ranks #4 on the current agent LLMBoard ranking
    • 1,838,727 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  5. 05
    MA
    Moonshot AI
    LLMBoard Agent Score
    78.26
    Price
    $3 input / $15 output per 1M tokens

    Strengths

    • Ranks #5 on the current agent LLMBoard ranking
    • 8,205,381 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score

Selection summary

Best AI Models for Agent

Claude Opus 5 currently leads the agent ranking at 86.35. Compare evidence coverage, source activity and price separately before choosing a model for production.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

GPT-5.6-Sol
Claude Fable 5
Kimi K3
LLMBoard rank #1Claude Opus 586.35
LLMBoard rank #2Claude Fable 5.185.92
LLMBoard rank #3GPT-5.6-Sol82.84