Capability ranking
Use this reasoning model leaderboard to compare the current ranking from the LLMBoard Reasoning Score after applying benchmark evidence requirements.
Updated 2026-09-08
This leaderboard ranking orders models by their LLMBoard Reasoning Score aggregated from eligible benchmark evidence.
Rank | Model | LLMBoard Reasoning Score | Context | Official input / 1M | Official output / 1M |
|---|
| Rank01 | ModelOP | LLMBoard Reasoning Score100.00 | Context1.1M | Official input / 1M$10 | Official output / 1M$50 |
| Rank02 | ModelAN | LLMBoard Reasoning Score98.88 | Context1M | Official input / 1M$5 | Official output / 1M$25 |
| Rank03 | ModelAC | LLMBoard Reasoning Score97.56 | Context1M | Official input / 1M$2 | Official output / 1M$6 |
| Rank04 | ModelOP | LLMBoard Reasoning Score95.94 | Context1.1M | Official input / 1M$5 | Official output / 1M$30 |
| Rank05 | ModelGO | LLMBoard Reasoning Score93.51 | Context1M | Official input / 1M$2 | Official output / 1M$12 |
| Rank06 | ModelAC | LLMBoard Reasoning Score93.24 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A |
| Rank07 | ModelAC | LLMBoard Reasoning Score93.24 | Context1M | Official input / 1M$0.15 | Official output / 1M$0.47 |
| Rank08 | ModelAC | LLMBoard Reasoning Score92.90 | Context1M | Official input / 1M$2.5 | Official output / 1M$7.5 |
| Rank09 | ModelOP | LLMBoard Reasoning Score88.94 | Context1M | Official input / 1M$2.5 | Official output / 1M$15 |
| Rank10 | ModelGO | LLMBoard Reasoning Score87.05 | Context1M | Official input / 1M$1.25 | Official output / 1M$10 |
| Rank11 | ModelGO | LLMBoard Reasoning Score86.59 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A |
| Rank12 | ModelXI | LLMBoard Reasoning Score83.88 | Context1M | Official input / 1M$0.435 | Official output / 1M$0.87 |
| Rank13 | ModelAC | LLMBoard Reasoning Score83.52 | Context1M | Official input / 1M$0.50 | Official output / 1M$3 |
| Rank14 | ModelAC | LLMBoard Reasoning Score83.36 | Context262.1K | Official input / 1M$0.60 | Official output / 1M$3.6 |
| Rank15 | ModelAN | LLMBoard Reasoning Score81.23 | Context1M | Official input / 1M$5 | Official output / 1M$25 |
| Rank16 | ModelGO | LLMBoard Reasoning Score78.44 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A |
| Rank17 | ModelOP | LLMBoard Reasoning Score77.79 | Context400K | Official input / 1M$21 | Official output / 1M$168 |
| Rank18 | ModelAC | LLMBoard Reasoning Score77.70 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A |
| Rank19 | ModelAN | LLMBoard Reasoning Score77.52 | Context200K | Official input / 1MN/A | Official output / 1MN/A |
| Rank20 | ModelAN | LLMBoard Reasoning Score76.66 | Context200K | Official input / 1M$3 | Official output / 1M$15 |
| Rank21 | ModelOP | LLMBoard Reasoning Score73.72 | Context400K | Official input / 1M$1.75 | Official output / 1M$14 |
| Rank22 | ModelAC | LLMBoard Reasoning Score73.62 | Context1M | Official input / 1M$0.50 | Official output / 1M$3 |
| Rank23 | ModelAN | LLMBoard Reasoning Score73.56 | Context200K | Official input / 1MN/A | Official output / 1MN/A |
| Rank24 | ModelGO | LLMBoard Reasoning Score72.21 | Context1M | Official input / 1MN/A | Official output / 1MN/A |
| Rank25 | ModelCO | LLMBoard Reasoning Score71.22 | Context128K | Official input / 1M$2.5 | Official output / 1M$10 |
| Rank26 | ModelAC | LLMBoard Reasoning Score70.65 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A |
| Rank27 | ModelGO | LLMBoard Reasoning Score70.44 | Context1M | Official input / 1M$0.50 | Official output / 1M$3 |
| Rank28 | ModelDE | LLMBoard Reasoning Score70.01 | Context131.1K | Official input / 1MN/A | Official output / 1MN/A |
| Rank29 | ModelGO | LLMBoard Reasoning Score69.47 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A |
| Rank30 | ModelGO | LLMBoard Reasoning Score68.58 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A |
Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.
Key findings from the current reasoning leaderboard ranking and its supporting benchmark coverage.
GPT-6-Astra leads this page at 100.00, 1.12 points ahead of Claude Opus 4.8.
Qwen3.8 Max is the highest-ranked open-weight option at #3. Nova Micro has the lowest official input price among ranked models.
The leading reasoning scores in this benchmark-backed ranking, with the axis focused on the competitive leaderboard range.
Compare reasoning benchmark strength with each model's broader capability score and leaderboard position.
Official vendor API prices plotted against the reasoning metric used on this page. Price remains a separate decision signal.
Vendor concentration and model access among the first 25 products in the ranking.
A data-backed look at the first five models in this reasoning leaderboard ranking, including benchmark context, price and output speed where available.
Common questions about the Best AI for Reasoning leaderboard, benchmark evidence and ranking method.
GPT-6-Astra is currently ranked first with a reasoning score of 100.00.
The current leaders are GPT-6-Astra (100.00 LLMBoard Reasoning Score), Claude Opus 4.8 (98.88 LLMBoard Reasoning Score), and Qwen3.8 Max (97.56 LLMBoard Reasoning Score).
Nova Micro has the lowest current official input price among ranked models at $0.04 per 1M tokens.
The fastest matched records are Llama 3.1 8B (2,047.00 tok/s via Cerebras), Llama 3.1 70B (1,204.00 tok/s via Cerebras), and GPT-4.1-mini (467.39 tok/s via OpenAI).
Qwen3.8 Max is the highest-ranked open-weight model at rank #3.
Gemini 1.5 Pro has the largest listed context window at 2.1M tokens.
Each model appears once using its current scored version. The leaderboard ranking follows the capability named in the title, while the overall page uses the LLMBoard score aggregated from eligible benchmark evidence.
This page currently ranks 117 unique model products.
Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.
No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.
Ranking basisThis reasoning AI model leaderboard uses the reasoning capability score shown on this page. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.
Selection summary
GPT-6-Astra is currently the best-ranked LLM for reasoning with a 100.00 reasoning score. The score is a relative ranking signal, so price, speed, context and evidence coverage should still be checked separately.
Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.