llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

xAI model product

Grok 3

Grok 3 is an llm from xAI launched on February 17, 2025.

Updated Sep 8, 2026. Default version: Grok-3

LLMBoard Score53.3Grok-3
Coverage0%5 benchmark families
Context window128KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Grok 3 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Grok-3 LLMBoard score breakdown

Grok 3 Benchmark Results

Benchmark scores for Grok-3.

7 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkAIME 2024Score93.30%Rank03Participants53Percentile96.15%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveCodeBenchScore79.40%Rank12Participants75Percentile85.14%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMMUScore78.00%Rank18Participants64Percentile73.02%EvidenceCEvaluatedSep 8, 2026
BenchmarkAIME 2025Score93.30%Rank29Participants119Percentile76.27%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore84.60%Rank64Participants247Percentile74.39%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,425.32 ratingRank72Participants210Percentile66.03%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena Text Style ControlScore1,411.47 ratingRank98Participants210Percentile53.59%EvidenceAEvaluatedSep 2, 2026

Grok 3 Arena Results

Preference and agent-evaluation results for the default version.

30 of 58 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenatextCategoryindustry medicine and healthcareRank28Rating / score1,466.98Votes1,606ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryfrenchRank42Rating / score1,461.08Votes392ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry entertainment and sports and mediaRank42Rating / score1,415.88Votes5,660ObservationsN/AResult dateSep 2, 2026
ArenatextCategorycreative writingRank48Rating / score1,413.96Votes4,697ObservationsN/AResult dateSep 2, 2026
ArenatextCategorygermanRank52Rating / score1,431.06Votes784ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry legal and governmentRank53Rating / score1,445.00Votes1,869ObservationsN/AResult dateSep 2, 2026
ArenatextCategorylonger queryRank54Rating / score1,438.61Votes4,862ObservationsN/AResult dateSep 2, 2026
ArenatextCategorypolishRank54Rating / score1,433.11Votes1,618ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryjapaneseRank60Rating / score1,388.72Votes823ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry writing and literature and languageRank63Rating / score1,410.69Votes7,725ObservationsN/AResult dateSep 2, 2026
ArenatextCategorykoreanRank69Rating / score1,374.09Votes520ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryexclude tiesRank70Rating / score1,418.43Votes22,978ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorygermanRank70Rating / score1,418.65Votes784ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryenglishRank71Rating / score1,437.67Votes17,053ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard prompts englishRank71Rating / score1,441.38Votes5,818ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryoverallRank72Rating / score1,425.32Votes32,415ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryfrenchRank72Rating / score1,448.67Votes392ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryinstruction followingRank73Rating / score1,408.60Votes9,626ObservationsN/AResult dateSep 2, 2026
ArenatextCategorynon englishRank73Rating / score1,410.26Votes15,353ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard promptsRank74Rating / score1,433.17Votes10,576ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry life and physical and social scienceRank76Rating / score1,436.88Votes5,714ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryrussianRank76Rating / score1,415.12Votes2,766ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank76Rating / score1,370.91Votes823ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycreative writingRank77Rating / score1,398.23Votes4,697ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry entertainment and sports and mediaRank77Rating / score1,397.15Votes5,660ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry writing and literature and languageRank77Rating / score1,405.79Votes7,725ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorypolishRank77Rating / score1,414.22Votes1,618ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymulti turnRank78Rating / score1,424.58Votes4,648ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry medicine and healthcareRank80Rating / score1,446.00Votes1,606ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry software and it servicesRank83Rating / score1,437.86Votes9,098ObservationsN/AResult dateSep 2, 2026

Grok 3 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
From $3 input, $15 output per 1M via Helicone
Tracked offerings
2
2 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderHeliconeProvider model IDgrok-3RegionglobalInput / 1M$3Output / 1M$15Context131.1KUpdatedSep 8, 2026
ProviderPoeProvider model IDxai/grok-3RegionglobalInput / 1M$3Output / 1M$15Context131.1KUpdatedSep 8, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Grok 3 Runtime Performance

Provider-specific output speed and catalog latency for Grok-3. Runtime does not affect the capability score.

1 row
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderxAIOutput Speed100.00 tok/sCatalog Latency0.70 sMax Input128KMax Output8KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Grok 3 Specifications

Technical details for the model's default version.

Version
Grok-3
Released
Feb 17, 2025
Knowledge cutoff
Nov 17, 2024
Parameters
N/A
Context window
128K
Max output
8K
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Grok 3 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionGrok-3ReleasedFeb 17, 2025LLMBoard53.30ParametersN/AContext128KMax output8KOpen weightsNoLicenseProprietary

Models similar to Grok 3

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#97-0.98
XA

Grok 4 Fast

xAI

52.32 LLMBoard

Details
#87+1.91
XA

Grok 4

xAI

55.21 LLMBoard

Details
#86+2.11
XA

Grok 4.20 Beta Reasoning

xAI

55.41 LLMBoard

Details
#112-5.29
XA

Grok 4.1 Thinking

xAI

48.01 LLMBoard

Details
#113-5.33
XA

Grok 3 Mini

xAI

47.97 LLMBoard

Details
#130-9.86
XA

Grok 4.1

xAI

43.44 LLMBoard

Details

What is Grok 3?

Key information about Grok 3 and its available data.

Grok 3 is an llm from xAI that was trained with ten times more compute than Grok 2. The Grok 3 family includes Grok 3 Reasoning and Grok 3 Mini Reasoning for complex problem-solving.

Data as of 2026-09-08.

FAQ

Common questions about Grok 3.

When was Grok 3 released?

Grok 3's default version was released on Feb 17, 2025.

How much does Grok 3 cost?

No official standard PAYG price is currently available for Grok 3. The lowest tracked third-party offer starts at $3 input and $15 output via Helicone.

Who created Grok 3?

Grok 3 was created by xAI.

What is the context window for Grok 3?

The default version has a 128K token context window.

Is Grok 3 open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Grok 3?

2 provider offerings are linked to the default version.