llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

xAI model product

Grok 4 Heavy

Grok 4 Heavy is xAI’s multi-agent version of Grok 4, using multiple Grok 4 agents that work in parallel and collaborate on solutions.

Updated Sep 8, 2026. Default version: Grok-4 Heavy

LLMBoard Score67.2Grok-4 Heavy
Coverage20%6 benchmark families
Context windowN/ATokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Grok 4 Heavy Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Grok-4 Heavy LLMBoard score breakdown

Grok 4 Heavy Benchmark Results

Benchmark scores for Grok-4 Heavy.

6 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkHMMT25Score96.70%Rank01Participants28Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkUSAMO25Score61.90%Rank02Participants3Percentile50.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkAIME 2025Score100.00%Rank04Participants119Percentile97.46%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveCodeBenchScore79.40%Rank13Participants75Percentile83.78%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanity's Last ExamScore50.70%Rank23Participants103Percentile78.43%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore88.40%Rank35Participants247Percentile86.18%EvidenceCEvaluatedSep 8, 2026

Grok 4 Heavy Arena Results

Preference and agent-evaluation results for the default version.

No Arena results

The default version does not have a matching Arena result yet.

Grok 4 Heavy Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No provider prices

The default version has no current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Grok 4 Heavy Runtime Performance

Provider-specific output speed and catalog latency for Grok-4 Heavy. Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Grok 4 Heavy Specifications

Technical details for the model's default version.

Version
Grok-4 Heavy
Released
Unknown
Knowledge cutoff
Dec 31, 2024
Parameters
N/A
Context window
N/A
Max output
N/A
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Grok 4 Heavy Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionGrok-4 HeavyReleasedN/ALLMBoard67.18ParametersN/AContextN/AMax outputN/AOpen weightsNoLicenseProprietary

Models similar to Grok 4 Heavy

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#32+9.37
XA

Grok 4.5

xAI

76.55 LLMBoard

Details
#28+9.63
XA

Grok 4.6

xAI

76.81 LLMBoard

Details
#86-11.77
XA

Grok 4.20 Beta Reasoning

xAI

55.41 LLMBoard

Details
#87-11.97
XA

Grok 4

xAI

55.21 LLMBoard

Details
#93-13.88
XA

Grok 3

xAI

53.30 LLMBoard

Details
#97-14.86
XA

Grok 4 Fast

xAI

52.32 LLMBoard

Details

What is Grok 4 Heavy?

Key information about Grok 4 Heavy and its available data.

Grok 4 Heavy is an AI model from xAI that runs multiple Grok 4 agents in parallel, compares their solutions, and combines their insights. It uses approximately 10x more test-time compute than regular Grok 4 and scored over 50% on text-only problems from the Humanities Last Exam and a perfect result on AIME 2025.

Data as of 2026-09-08.

FAQ

Common questions about Grok 4 Heavy.

When was Grok 4 Heavy released?

A release date is not available for the default version.

How much does Grok 4 Heavy cost?

No official standard PAYG price is currently available for Grok 4 Heavy.

Who created Grok 4 Heavy?

Grok 4 Heavy was created by xAI.

What is the context window for Grok 4 Heavy?

A context window is not available for the default version.

Is Grok 4 Heavy open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Grok 4 Heavy?

No provider offering is currently linked to the default version.

Browse runtime rankings