llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

OpenAI model product

o3

o3 is an OpenAI large language model for multi-step reasoning across text, code, and images.

Updated Sep 8, 2026. Default version: o3

LLMBoard Score51.7o3
Coverage40%22 benchmark families
Context window200KTokens
Official input price$2OpenAI API

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

o3 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

o3 LLMBoard score breakdown

o3 Benchmark Results

Benchmark scores for o3.

30 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkAider-PolyglotScore81.30%Rank03Participants22Percentile90.48%EvidenceCEvaluatedSep 8, 2026
BenchmarkCOLLIEScore98.40%Rank03Participants10Percentile77.78%EvidenceCEvaluatedSep 8, 2026
BenchmarkMathVistaScore86.80%Rank03Participants39Percentile94.74%EvidenceCEvaluatedSep 8, 2026
BenchmarkARC-AGIScore88.00%Rank05Participants9Percentile50.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkAIME 2024Score91.60%Rank06Participants53Percentile90.38%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau-benchScore63.00%Rank06Participants6Percentile0.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMMUScore82.90%Rank07Participants64Percentile90.48%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 RetailScore80.20%Rank08Participants27Percentile73.08%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 AirlineScore64.80%Rank09Participants24Percentile65.22%EvidenceCEvaluatedSep 8, 2026
BenchmarkMulti-ChallengeScore60.40%Rank10Participants29Percentile67.86%EvidenceCEvaluatedSep 8, 2026
BenchmarkERQAScore64.00%Rank12Participants26Percentile56.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkFrontierMathScore15.80%Rank13Participants17Percentile25.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkVideoMMMUScore83.30%Rank13Participants26Percentile52.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Search Style ControlScore1,191.88 ratingRank15Participants28Percentile48.15%EvidenceAEvaluatedAug 24, 2026
BenchmarkARC-AGI v2Score6.50%Rank17Participants18Percentile5.88%EvidenceBEvaluatedSep 8, 2026
BenchmarkLM Arena SearchScore1,144.26 ratingRank22Participants28Percentile22.22%EvidenceAEvaluatedAug 24, 2026
BenchmarkMMMU-ProScore76.40%Rank28Participants69Percentile60.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Search FactualityScore1,129.33 ratingRank28Participants28Percentile0.00%EvidenceAEvaluatedAug 24, 2026
BenchmarkCharXiv-RScore78.60%Rank29Participants55Percentile48.15%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 TelecomScore58.20%Rank31Participants36Percentile14.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkBrowseCompScore49.70%Rank47Participants63Percentile25.81%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Vision Style ControlScore1,215.58 ratingRank54Participants107Percentile50.00%EvidenceAEvaluatedAug 27, 2026
BenchmarkLM Arena VisionScore1,214.21 ratingRank57Participants107Percentile47.17%EvidenceAEvaluatedAug 27, 2026
BenchmarkAIME 2025Score86.40%Rank61Participants119Percentile49.15%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench VerifiedScore69.10%Rank70Participants113Percentile38.39%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore83.30%Rank75Participants247Percentile69.92%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Text Style ControlScore1,431.36 ratingRank75Participants210Percentile64.59%EvidenceAEvaluatedSep 2, 2026
BenchmarkHumanity's Last ExamScore14.70%Rank81Participants103Percentile21.57%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,409.75 ratingRank94Participants210Percentile55.50%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena Text FactualityScore1,404.15 ratingRank102Participants121Percentile15.83%EvidenceAEvaluatedSep 2, 2026

o3 Arena Results

Preference and agent-evaluation results for the default version.

30 of 98 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
Arenavision style controlCategoryentity recognitionRank11Rating / score1,232.33Votes496ObservationsN/AResult dateAug 27, 2026
Arenasearch style controlCategoryoverallRank15Rating / score1,191.88Votes20,644ObservationsN/AResult dateAug 24, 2026
ArenavisionCategorycaptioningRank17Rating / score1,210.59Votes542ObservationsN/AResult dateAug 27, 2026
ArenavisionCategorycreative writingRank17Rating / score1,197.19Votes1,880ObservationsN/AResult dateJan 9, 2026
Arenavision style controlCategorycreative writingRank18Rating / score1,191.82Votes1,880ObservationsN/AResult dateJan 9, 2026
ArenavisionCategoryentity recognitionRank19Rating / score1,234.78Votes496ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategorycaptioningRank19Rating / score1,182.22Votes542ObservationsN/AResult dateAug 27, 2026
ArenasearchCategoryoverallRank22Rating / score1,144.26Votes20,644ObservationsN/AResult dateAug 24, 2026
Arenasearch factualityCategoryoverallRank28Rating / score1,129.33Votes3,689ObservationsN/AResult dateAug 24, 2026
Arenatext style controlCategorypolishRank38Rating / score1,457.32Votes3,826ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank39Rating / score1,424.31Votes1,405ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryhumorRank45Rating / score1,211.25Votes1,634ObservationsN/AResult dateAug 27, 2026
ArenatextCategoryjapaneseRank46Rating / score1,403.44Votes1,405ObservationsN/AResult dateSep 2, 2026
ArenatextCategorypolishRank46Rating / score1,437.72Votes3,826ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry medicine and healthcareRank46Rating / score1,474.86Votes3,319ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryhumorRank48Rating / score1,222.39Votes1,634ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryhomeworkRank52Rating / score1,247.12Votes1,935ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorymathRank53Rating / score1,447.57Votes3,679ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategorycreative writing visionRank53Rating / score1,194.26Votes1,586ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryenglishRank54Rating / score1,218.41Votes21,630ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryoverallRank54Rating / score1,215.58Votes45,594ObservationsN/AResult dateAug 27, 2026
ArenavisionCategorychineseRank55Rating / score1,234.96Votes2,261ObservationsN/AResult dateAug 27, 2026
ArenavisionCategorycreative writing visionRank56Rating / score1,202.81Votes1,586ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryhomeworkRank56Rating / score1,234.11Votes1,935ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategorychineseRank56Rating / score1,228.25Votes2,261ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorykoreanRank57Rating / score1,394.11Votes1,187ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryenglishRank57Rating / score1,221.77Votes21,630ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryocrRank57Rating / score1,217.89Votes14,351ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryoverallRank57Rating / score1,214.21Votes45,594ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategorydiagramRank57Rating / score1,231.87Votes3,556ObservationsN/AResult dateAug 27, 2026

o3 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
$2 input, $8 output per 1M
Official provider
OpenAI
Lowest third-party
From $1.8 input, $7.2 output per 1M via Poe
Tracked offerings
13
12 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderPoeProvider model IDopenai/o3RegionglobalInput / 1M$1.8Output / 1M$7.2Context200KUpdatedSep 8, 2026
ProviderNanoGPTProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderImpossiblProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderOpenRouterProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderOpenAIProvider model IDo3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderLLM GatewayProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderNEAR AI CloudProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderMerge GatewayProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderCloudflare AI GatewayProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderVercel AI GatewayProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderEden AIProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026
ProviderKilo GatewayProvider model IDopenai/o3RegionglobalInput / 1M$2Output / 1M$8Context200KUpdatedSep 8, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

o3 Runtime Performance

Provider-specific output speed and catalog latency for o3. Runtime does not affect the capability score.

1 row
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderOpenAIOutput Speed50.00 tok/sCatalog Latency20.00 sMax Input200KMax Output100KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

o3 Specifications

Technical details for the model's default version.

Version
o3
Released
Apr 16, 2025
Knowledge cutoff
May 31, 2024
Parameters
N/A
Context window
200K
Max output
100K
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

o3 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
Versiono3ReleasedApr 16, 2025LLMBoard51.69ParametersN/AContext200KMax output100KOpen weightsNoLicenseProprietary

Models similar to o3

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#98+0.41
OP

GPT-5.5-Instant

OpenAI

52.10 LLMBoard

Details
#108-2.86
OP

GPT-5

OpenAI

48.83 LLMBoard

Details
#109-2.88
OP

GPT-5.1-Codex

OpenAI

48.81 LLMBoard

Details
#114-3.80
OP

GPT-5.4-nano

OpenAI

47.89 LLMBoard

Details
#81+5.62
OP

GPT-5.2-Codex

OpenAI

57.31 LLMBoard

Details
#118-5.71
OP

GPT-OSS-20B

OpenAI

45.98 LLMBoard

Details

What is o3?

Key information about o3 and its available data.

o3 is an OpenAI large language model designed for reasoning tasks in math, science, coding, and visual analysis. It also supports technical writing and instruction-following.

Data as of 2026-09-08.

FAQ

Common questions about o3.

When was o3 released?

o3's default version was released on Apr 16, 2025.

How much does o3 cost?

o3's official API price is $2 per million input tokens and $8 per million output tokens via OpenAI. The lowest tracked third-party offer starts at $1.8 input and $7.2 output via Poe.

Who created o3?

o3 was created by OpenAI.

What is the context window for o3?

The default version has a 200K token context window.

Is o3 open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer o3?

13 provider offerings are linked to the default version.