llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Anthropic model product

Claude Opus 4.8

7, with reported improvements in software engineering, agentic tool use, reasoning, computer use, and knowledge work.

Updated Sep 8, 2026. Default version: Claude Opus 4.8

LLMBoard Score85.4Claude Opus 4.8
Coverage60%27 benchmark families
Context window1MTokens
Official input price$5Anthropic API

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Claude Opus 4.8 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Claude Opus 4.8 LLMBoard score breakdown

Claude Opus 4.8 Benchmark Results

Benchmark scores for Claude Opus 4.8.

30 of 44 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkIncludeScore87.60%Rank01Participants31Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkDeepSearchQAScore93.10%Rank02Participants10Percentile88.89%EvidenceCEvaluatedSep 8, 2026
BenchmarkGraphwalks parents >128kScore83.30%Rank02Participants7Percentile83.33%EvidenceCEvaluatedSep 8, 2026
BenchmarkScreenSpot ProScore87.90%Rank02Participants26Percentile96.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-bench MultilingualScore84.40%Rank02Participants43Percentile97.62%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Agent SteerabilityScore12.59%Rank02Participants49Percentile97.92%EvidenceAEvaluatedSep 5, 2026
BenchmarkOfficeQA ProScore66.20%Rank03Participants9Percentile75.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench MultimodalScore38.40%Rank03Participants4Percentile33.33%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench ProScore69.20%Rank03Participants55Percentile96.30%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench VerifiedScore88.60%Rank03Participants113Percentile98.21%EvidenceCEvaluatedSep 8, 2026
BenchmarkFinance Agent v2Score53.92%Rank04Participants26Percentile88.00%EvidenceBEvaluatedSep 8, 2026
BenchmarkFrontierSWEScore75.00%Rank04Participants16Percentile80.00%EvidenceBEvaluatedSep 8, 2026
BenchmarkOSWorld-VerifiedScore83.40%Rank04Participants24Percentile86.96%EvidenceCEvaluatedSep 8, 2026
BenchmarkFrontierCode 1.1Score46.50%Rank05Participants17Percentile75.00%EvidenceBEvaluatedSep 8, 2026
BenchmarkGraphwalks BFS >128kScore68.10%Rank05Participants11Percentile60.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Agent LeaderboardScore9.17%Rank05Participants49Percentile91.67%EvidenceAEvaluatedSep 5, 2026
BenchmarkLM Arena Agent Praise ComplaintScore18.04%Rank05Participants49Percentile91.67%EvidenceAEvaluatedSep 5, 2026
BenchmarkLM Arena Document Style ControlScore1,486.56 ratingRank05Participants33Percentile87.50%EvidenceAEvaluatedJul 30, 2026
BenchmarkGPQAScore93.60%Rank06Participants247Percentile97.97%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveBenchScore77.22%Rank06Participants38Percentile86.49%EvidenceBEvaluatedSep 8, 2026
BenchmarkMCP AtlasScore82.20%Rank06Participants34Percentile84.85%EvidenceCEvaluatedSep 8, 2026
BenchmarkCharXiv-RScore89.90%Rank07Participants55Percentile88.89%EvidenceCEvaluatedSep 8, 2026
BenchmarkCyberGymScore78.80%Rank07Participants15Percentile57.14%EvidenceCEvaluatedSep 8, 2026
BenchmarkFinance AgentScore53.90%Rank07Participants8Percentile14.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkHealthBench ProfessionalScore55.80%Rank07Participants10Percentile33.33%EvidenceCEvaluatedSep 8, 2026
BenchmarkTerminal-Bench 2.0Score74.60%Rank07Participants51Percentile88.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkTerminal-Bench 4.0Score23.60%Rank07Participants13Percentile50.00%EvidenceBEvaluatedSep 8, 2026
BenchmarkHumanity's Last ExamScore57.90%Rank09Participants103Percentile92.16%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Agent Bash Recovery StepsScore10.17%Rank09Participants49Percentile83.33%EvidenceAEvaluatedSep 5, 2026
BenchmarkLM Arena DocumentScore1,474.59 ratingRank09Participants33Percentile75.00%EvidenceAEvaluatedJul 30, 2026

Claude Opus 4.8 Arena Results

Preference and agent-evaluation results for the default version.

30 of 100 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
Arenaagent steerabilityCategoryoverallRank02Rating / score0.13VotesN/AObservations27.2KResult dateSep 5, 2026
ArenavisionCategoryhomeworkRank04Rating / score1,338.12Votes1,499ObservationsN/AResult dateAug 27, 2026
ArenaagentCategoryoverallRank05Rating / score0.09VotesN/AObservations1.9MResult dateSep 5, 2026
Arenaagent praise complaintCategoryoverallRank05Rating / score0.18VotesN/AObservations8.4KResult dateSep 5, 2026
Arenadocument style controlCategoryoverallRank05Rating / score1,486.56Votes8,388ObservationsN/AResult dateJul 30, 2026
Arenatext style controlCategoryinstruction followingRank05Rating / score1,492.26Votes16,979ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryhomeworkRank05Rating / score1,332.85Votes1,499ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorycodingRank06Rating / score1,534.26Votes13,439ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryindustry mathematicalRank07Rating / score1,486.69Votes2,144ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry business and management and financial operationsRank07Rating / score1,491.89Votes9,841ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategorydiagramRank07Rating / score1,321.10Votes3,684ObservationsN/AResult dateAug 27, 2026
Arenatext factualityCategoryspanishRank08Rating / score1,474.36Votes1,087ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryexpertRank08Rating / score1,525.94Votes5,208ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryfrenchRank08Rating / score1,508.82Votes1,751ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard promptsRank08Rating / score1,514.61Votes32,372ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard prompts englishRank08Rating / score1,515.06Votes14,694ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorylonger queryRank08Rating / score1,501.81Votes22,671ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorymulti turnRank08Rating / score1,497.24Votes8,708ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryenglishRank08Rating / score1,282.83Votes5,661ObservationsN/AResult dateAug 27, 2026
Arenaagent bash recovery stepsCategoryoverallRank09Rating / score0.10VotesN/AObservations35.6KResult dateSep 5, 2026
ArenadocumentCategoryoverallRank09Rating / score1,474.59Votes8,388ObservationsN/AResult dateJul 30, 2026
Arenatext factualityCategoryjapaneseRank09Rating / score1,449.69Votes427ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategorymathRank09Rating / score1,483.10Votes1,956ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategorypolishRank09Rating / score1,457.99Votes491ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry software and it servicesRank09Rating / score1,524.05Votes19,325ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryspanishRank09Rating / score1,486.38Votes1,434ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategorycreative writing visionRank09Rating / score1,293.93Votes907ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorygermanRank10Rating / score1,492.37Votes852ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryrussianRank10Rating / score1,494.63Votes5,083ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryocrRank10Rating / score1,300.61Votes9,631ObservationsN/AResult dateAug 27, 2026

Claude Opus 4.8 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
$5 input, $25 output per 1M
Official provider
Anthropic
Lowest third-party
From $0.425 input, $2.13 output per 1M via UnoRouter
Tracked offerings
44
30 of 41 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderKenariProvider model IDclaude-opus-4-8RegionglobalInput / 1MN/AOutput / 1MN/AContext1MUpdatedSep 8, 2026
ProviderUnoRouterProvider model IDclaude-opus-4-8RegionglobalInput / 1M$0.425Output / 1M$2.13Context1MUpdatedSep 8, 2026
ProviderXpersonaProvider model IDclaude-opus-4-8RegionglobalInput / 1M$1.5Output / 1M$9.25Context200KUpdatedSep 8, 2026
ProviderPoeProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$4.29Output / 1M$21.46Context1MUpdatedSep 8, 2026
ProviderOrcaRouterProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderNanoGPTProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderVertexProvider model IDclaude-opus-4-8@defaultRegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderImpossiblProvider model IDanthropic/claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderOpenRouterProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderVertex (Anthropic)Provider model IDclaude-opus-4-8@defaultRegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderDaoXEProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderCrossModelProvider model IDanthropic/claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderAnthropicProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
Providerrouting.runProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderAmazon BedrockProvider model IDanthropic.claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderOpperProvider model IDanthropic/claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderLLM GatewayProvider model IDanthropic/claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderMerge GatewayProvider model IDanthropic/claude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderCloudflare AI GatewayProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderFastRouterProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderZenMuxProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderGMI CloudProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderFreeModelProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderOfoxProvider model IDanthropic/claude-opus-4.8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderNeonProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderAzure Cognitive ServicesProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderOpenCode ZenProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderRequestyProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderAzureProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context1MUpdatedSep 8, 2026
ProviderAIHubMixProvider model IDclaude-opus-4-8RegionglobalInput / 1M$5Output / 1M$25Context200KUpdatedSep 8, 2026

Claude Opus 4.8 Runtime Performance

Provider-specific output speed and catalog latency for Claude Opus 4.8. Runtime does not affect the capability score.

2 rows
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderAnthropicOutput Speed42.00 tok/sCatalog Latency0.50 sMax Input1MMax Output128KUpdatedSep 8, 2026
ProviderVertex AIOutput Speed42.00 tok/sCatalog Latency0.50 sMax Input1MMax Output128KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Claude Opus 4.8 Specifications

Technical details for the model's default version.

Version
Claude Opus 4.8
Released
May 28, 2026
Knowledge cutoff
Unknown
Parameters
N/A
Context window
1M
Max output
128K
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Claude Opus 4.8 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionClaude Opus 4.8ReleasedMay 28, 2026LLMBoard85.39ParametersN/AContext1MMax output128KOpen weightsNoLicenseProprietary

Models similar to Claude Opus 4.8

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#19-3.13
AN

Claude Opus 4.7

Anthropic

82.26 LLMBoard

Details
#23-5.23
AN

Claude Sonnet 5

Anthropic

80.16 LLMBoard

Details
#5+5.91
AN

Claude Mythos

Anthropic

91.30 LLMBoard

Details
#4+6.24
AN

Claude Fable 5

Anthropic

91.63 LLMBoard

Details
#25-7.21
AN

Claude Opus 4.6

Anthropic

78.18 LLMBoard

Details
#3+7.85
AN

Claude Opus 5

Anthropic

93.24 LLMBoard

Details

What is Claude Opus 4.8?

Key information about Claude Opus 4.8 and its available data.

7. At release, Anthropic described it as its most capable general-access model, with high, extra (xhigh), and max effort levels for different task demands.

Data as of 2026-09-08.

FAQ

Common questions about Claude Opus 4.8.

When was Claude Opus 4.8 released?

Claude Opus 4.8's default version was released on May 28, 2026.

How much does Claude Opus 4.8 cost?

Claude Opus 4.8's official API price is $5 per million input tokens and $25 per million output tokens via Anthropic. The lowest tracked third-party offer starts at $0.425 input and $2.13 output via UnoRouter.

Who created Claude Opus 4.8?

Claude Opus 4.8 was created by Anthropic.

What is the context window for Claude Opus 4.8?

The default version has a 1M token context window.

Is Claude Opus 4.8 open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Claude Opus 4.8?

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

44 provider offerings are linked to the default version.