llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Anthropic model product

Claude Opus 4

Claude Opus 4 is an Anthropic large language model in the Claude 4 family for coding, advanced reasoning, and extended tool-using tasks.

Updated Sep 8, 2026. Default version: Claude Opus 4

LLMBoard Score45.5Claude Opus 4
Coverage20%9 benchmark families
Context window200KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Claude Opus 4 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Claude Opus 4 LLMBoard score breakdown

Claude Opus 4 Benchmark Results

Benchmark scores for Claude Opus 4.

16 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkMMMU (validation)Score76.50%Rank03Participants4Percentile33.33%EvidenceCEvaluatedSep 8, 2026
BenchmarkTAU-bench RetailScore81.40%Rank03Participants25Percentile91.67%EvidenceCEvaluatedSep 8, 2026
BenchmarkTAU-bench AirlineScore59.60%Rank08Participants23Percentile68.18%EvidenceCEvaluatedSep 8, 2026
BenchmarkTerminal-BenchScore39.20%Rank10Participants25Percentile62.50%EvidenceCEvaluatedSep 8, 2026
BenchmarkARC-AGI v2Score8.60%Rank16Participants18Percentile11.76%EvidenceBEvaluatedSep 8, 2026
BenchmarkMMMLUScore88.80%Rank16Participants49Percentile68.75%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Search Style ControlScore1,150.60 ratingRank24Participants28Percentile14.81%EvidenceAEvaluatedAug 24, 2026
BenchmarkLM Arena Search FactualityScore1,137.74 ratingRank25Participants28Percentile11.11%EvidenceAEvaluatedAug 24, 2026
BenchmarkLM Arena SearchScore1,126.47 ratingRank28Participants28Percentile0.00%EvidenceAEvaluatedAug 24, 2026
BenchmarkSWE-Bench VerifiedScore72.50%Rank54Participants113Percentile52.68%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Vision Style ControlScore1,207.36 ratingRank58Participants107Percentile46.23%EvidenceAEvaluatedAug 27, 2026
BenchmarkLM Arena VisionScore1,191.50 ratingRank67Participants107Percentile37.74%EvidenceAEvaluatedAug 27, 2026
BenchmarkAIME 2025Score75.50%Rank82Participants119Percentile31.36%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Text Style ControlScore1,425.73 ratingRank82Participants210Percentile61.24%EvidenceAEvaluatedSep 2, 2026
BenchmarkGPQAScore79.60%Rank98Participants247Percentile60.57%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,376.75 ratingRank116Participants210Percentile44.98%EvidenceAEvaluatedSep 2, 2026

Claude Opus 4 Arena Results

Preference and agent-evaluation results for the default version.

30 of 73 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
Arenavision style controlCategorycreative writingRank08Rating / score1,258.38Votes192ObservationsN/AResult dateJan 9, 2026
ArenavisionCategorycreative writingRank09Rating / score1,253.43Votes192ObservationsN/AResult dateJan 9, 2026
Arenasearch style controlCategoryoverallRank24Rating / score1,150.60Votes31,037ObservationsN/AResult dateAug 24, 2026
Arenasearch factualityCategoryoverallRank25Rating / score1,137.74Votes13,998ObservationsN/AResult dateAug 24, 2026
ArenasearchCategoryoverallRank28Rating / score1,126.47Votes31,037ObservationsN/AResult dateAug 24, 2026
Arenatext style controlCategorycreative writingRank43Rating / score1,431.88Votes4,434ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorylonger queryRank45Rating / score1,465.57Votes6,942ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorygermanRank46Rating / score1,450.24Votes951ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryinstruction followingRank47Rating / score1,444.00Votes9,085ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycodingRank50Rating / score1,498.11Votes6,561ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry entertainment and sports and mediaRank52Rating / score1,420.90Votes6,263ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry writing and literature and languageRank53Rating / score1,430.62Votes7,772ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryhomeworkRank55Rating / score1,241.67Votes306ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryhomeworkRank58Rating / score1,227.88Votes306ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryoverallRank58Rating / score1,207.36Votes1,291ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryocrRank59Rating / score1,217.47Votes835ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorykoreanRank63Rating / score1,384.08Votes776ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryrussianRank63Rating / score1,441.00Votes2,551ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryspanishRank63Rating / score1,432.52Votes854ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank64Rating / score1,387.75Votes820ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryocrRank65Rating / score1,196.43Votes835ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryenglishRank66Rating / score1,195.60Votes956ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryoverallRank67Rating / score1,191.50Votes1,291ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategorydiagramRank67Rating / score1,208.48Votes278ObservationsN/AResult dateAug 27, 2026
ArenavisionCategorydiagramRank69Rating / score1,180.08Votes278ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategoryhard promptsRank70Rating / score1,456.15Votes16,134ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryenglishRank70Rating / score1,186.15Votes956ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorypolishRank71Rating / score1,419.91Votes3,566ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorymulti turnRank72Rating / score1,439.40Votes6,056ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard prompts englishRank73Rating / score1,459.70Votes7,151ObservationsN/AResult dateSep 2, 2026

Claude Opus 4 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
From $13 input, $64 output per 1M via Poe
Tracked offerings
12
12 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderPoeProvider model IDanthropic/claude-opus-4RegionglobalInput / 1M$13Output / 1M$64Context192.5KUpdatedSep 8, 2026
ProviderJiekou.AIProvider model IDclaude-opus-4-20250514RegionglobalInput / 1M$13.5Output / 1M$67.5Context200KUpdatedSep 8, 2026
ProviderNanoGPTProvider model IDclaude-opus-4-20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderVertexProvider model IDclaude-opus-4@20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderOpenRouterProvider model IDanthropic/claude-opus-4RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderVertex (Anthropic)Provider model IDclaude-opus-4@20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderMerge GatewayProvider model IDanthropic/claude-opus-4-20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderDigitalOceanProvider model IDanthropic-claude-opus-4RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderZenMuxProvider model IDanthropic/claude-opus-4RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
Provider302.AIProvider model IDclaude-opus-4-20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderVercel AI GatewayProvider model IDanthropic/claude-opus-4RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026
ProviderAbacusProvider model IDclaude-opus-4-20250514RegionglobalInput / 1M$15Output / 1M$75Context200KUpdatedSep 8, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Claude Opus 4 Runtime Performance

Provider-specific output speed and catalog latency for Claude Opus 4. Runtime does not affect the capability score.

2 rows
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderAnthropicOutput Speed100.00 tok/sCatalog Latency0.50 sMax Input200KMax Output32KUpdatedSep 8, 2026
ProviderGoogleOutput Speed42.00 tok/sCatalog Latency0.40 sMax Input200KMax Output128KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Claude Opus 4 Specifications

Technical details for the model's default version.

Version
Claude Opus 4
Released
May 22, 2025
Knowledge cutoff
Unknown
Parameters
N/A
Context window
200K
Max output
128K
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Claude Opus 4 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionClaude Opus 4ReleasedMay 22, 2025LLMBoard45.50ParametersN/AContext200KMax output128KOpen weightsNoLicenseProprietary

Models similar to Claude Opus 4

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#121-0.82
AN

Claude Haiku 4.5

Anthropic

44.68 LLMBoard

Details
#101+5.64
AN

Claude Opus 4.1

Anthropic

51.14 LLMBoard

Details
#150-6.79
AN

Claude Sonnet 4

Anthropic

38.71 LLMBoard

Details
#153-8.00
AN

Claude Sonnet 3.7

Anthropic

37.50 LLMBoard

Details
#85+10.04
AN

Claude Sonnet 4.5

Anthropic

55.54 LLMBoard

Details
#180-14.74
AN

Claude Sonnet 3.5

Anthropic

30.76 LLMBoard

Details

What is Claude Opus 4?

Key information about Claude Opus 4 and its available data.

Claude Opus 4 is a large language model from Anthropic and part of the Claude 4 family. It supports coding, advanced reasoning, web search and other tool use during extended tasks, parallel tool execution, and improved memory capabilities.

Data as of 2026-09-08.

FAQ

Common questions about Claude Opus 4.

When was Claude Opus 4 released?

Claude Opus 4's default version was released on May 22, 2025.

How much does Claude Opus 4 cost?

No official standard PAYG price is currently available for Claude Opus 4. The lowest tracked third-party offer starts at $13 input and $64 output via Poe.

Who created Claude Opus 4?

Claude Opus 4 was created by Anthropic.

What is the context window for Claude Opus 4?

The default version has a 200K token context window.

Is Claude Opus 4 open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Claude Opus 4?

12 provider offerings are linked to the default version.