xAI model product
Grok 2 is an LLM from xAI for chat, coding, reasoning, visual math reasoning, and document-based question answering.
Updated Sep 8, 2026. Default version: Grok-2
This profile uses the model's current scored version. Arena ratings and prices are shown separately.
Benchmark scores for Grok-2.
Benchmark | Score | Rank | Participants | Percentile | Evidence | Evaluated |
|---|
| BenchmarkDocVQA | Score93.60% | Rank09 | Participants28 | Percentile70.37% | EvidenceC | Evaluated |
| BenchmarkMathVista | Score69.00% | Rank16 | Participants39 | Percentile60.53% | EvidenceC | Evaluated |
| BenchmarkHumanEval | Score88.40% | Rank21 | Participants66 | Percentile69.23% | EvidenceC | Evaluated |
| BenchmarkMMLU | Score87.50% | Rank24 | Participants101 | Percentile77.00% | EvidenceC | Evaluated |
| BenchmarkMATH | Score76.10% | Rank26 | Participants71 | Percentile64.29% | EvidenceC | Evaluated |
| BenchmarkMMMU | Score66.10% | Rank38 | Participants64 | Percentile41.27% | EvidenceC | Evaluated |
| BenchmarkMMLU-Pro | Score75.50% | Rank77 | Participants138 | Percentile44.53% | EvidenceC | Evaluated |
| BenchmarkLM Arena Text Style Control | Score1,335.44 rating | Rank159 | Participants210 | Percentile24.40% | EvidenceA | Evaluated |
| BenchmarkLM Arena Text | Score1,304.35 rating | Rank160 | Participants210 | Percentile23.92% | EvidenceA | Evaluated |
| BenchmarkGPQA | Score56.00% | Rank178 | Participants247 | Percentile28.05% | EvidenceC | Evaluated |
Preference and agent-evaluation results for the default version.
Arena | Category | Rank | Rating / score | Votes | Observations | Result date |
|---|
| Arenatext style control | Categoryjapanese | Rank130 | Rating / score1,270.15 | Votes1,478 | ObservationsN/A | Result date |
| Arenatext | Categoryjapanese | Rank131 | Rating / score1,244.19 | Votes1,478 | ObservationsN/A | Result date |
| Arenatext style control | Categorykorean | Rank131 | Rating / score1,278.82 | Votes952 | ObservationsN/A | Result date |
| Arenatext | Categorykorean | Rank134 | Rating / score1,236.53 | Votes952 | ObservationsN/A | Result date |
| Arenatext style control | Categorygerman | Rank137 | Rating / score1,313.85 | Votes1,609 | ObservationsN/A | Result date |
| Arenatext | Categoryfrench | Rank138 | Rating / score1,317.32 | Votes604 | ObservationsN/A | Result date |
| Arenatext | Categorygerman | Rank138 | Rating / score1,286.59 | Votes1,609 | ObservationsN/A | Result date |
| Arenatext style control | Categoryfrench | Rank138 | Rating / score1,351.78 | Votes604 | ObservationsN/A | Result date |
| Arenatext style control | Categorycreative writing | Rank142 | Rating / score1,316.84 | Votes9,339 | ObservationsN/A | Result date |
| Arenatext style control | Categoryindustry entertainment and sports and media | Rank145 | Rating / score1,314.28 | Votes10,293 | ObservationsN/A | Result date |
| Arenatext style control | Categoryindustry legal and government | Rank145 | Rating / score1,365.48 | Votes3,813 | ObservationsN/A | Result date |
| Arenatext style control | Categoryindustry writing and literature and language | Rank145 | Rating / score1,326.71 | Votes17,406 | ObservationsN/A | Result date |
| Arenatext | Categoryspanish | Rank147 | Rating / score1,280.39 | Votes671 | ObservationsN/A | Result date |
| Arenatext style control | Categoryindustry medicine and healthcare | Rank149 | Rating / score1,355.97 | Votes3,206 | ObservationsN/A | Result date |
| Arenatext style control | Categoryspanish | Rank150 | Rating / score1,311.82 | Votes671 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry entertainment and sports and media | Rank152 | Rating / score1,281.95 | Votes10,293 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry legal and government | Rank153 | Rating / score1,325.76 | Votes3,813 | ObservationsN/A | Result date |
| Arenatext style control | Categorynon english | Rank153 | Rating / score1,318.83 | Votes29,136 | ObservationsN/A | Result date |
| Arenatext style control | Categoryrussian | Rank154 | Rating / score1,326.26 | Votes8,745 | ObservationsN/A | Result date |
| Arenatext style control | Categorychinese | Rank155 | Rating / score1,342.17 | Votes5,317 | ObservationsN/A | Result date |
| Arenatext | Categorychinese | Rank156 | Rating / score1,289.24 | Votes5,317 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry medicine and healthcare | Rank156 | Rating / score1,305.94 | Votes3,206 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry writing and literature and language | Rank158 | Rating / score1,291.53 | Votes17,406 | ObservationsN/A | Result date |
| Arenatext style control | Categoryindustry life and physical and social science | Rank158 | Rating / score1,351.38 | Votes10,796 | ObservationsN/A | Result date |
| Arenatext style control | Categorylonger query | Rank158 | Rating / score1,335.29 | Votes8,902 | ObservationsN/A | Result date |
| Arenatext | Categorycreative writing | Rank159 | Rating / score1,284.10 | Votes9,339 | ObservationsN/A | Result date |
| Arenatext style control | Categoryexclude ties | Rank159 | Rating / score1,287.76 | Votes40,763 | ObservationsN/A | Result date |
| Arenatext style control | Categoryoverall | Rank159 | Rating / score1,335.44 | Votes63,498 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry life and physical and social science | Rank160 | Rating / score1,314.41 | Votes10,796 | ObservationsN/A | Result date |
| Arenatext | Categoryindustry mathematical | Rank160 | Rating / score1,283.70 | Votes7,604 | ObservationsN/A | Result date |
Official vendor API pricing appears first, followed by individual provider offers.
The default version has no current input or output token prices.
Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.
Provider-specific output speed and catalog latency for Grok-2. Runtime does not affect the capability score.
Provider | Output Speed | Catalog Latency | Max Input | Max Output | Updated |
|---|
| ProviderxAI | Output Speed85.00 tok/s | Catalog Latency0.70 s | Max Input128K | Max Output8K | Updated |
Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.
Technical details for the model's default version.
Available versions of this model. The score column identifies the version used in the overall ranking.
Version | Released | LLMBoard | Parameters | Context | Max output | Open weights | License |
|---|
| VersionGrok-2 | Released | LLMBoard18.43 | ParametersN/A | Context128K | Max output8K | Open weightsNo | LicenseProprietary |
Recommendations prioritize the same model type and family, then the closest LLMBoard score.
Key information about Grok 2 and its available data.
Grok 2 is a language model developed by xAI. It is described as supporting chat, coding, reasoning, visual math reasoning, document-based question answering, and academic benchmarks in reasoning, reading comprehension, math, and science.
Data as of 2026-09-08.
Common questions about Grok 2.
Grok 2's default version was released on Aug 13, 2024.
No official standard PAYG price is currently available for Grok 2.
Grok 2 was created by xAI.
The default version has a 128K token context window.
No. The default version is not marked as having publicly available weights.
No provider offering is currently linked to the default version.