llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Capability ranking

Best AI for Reasoning

Use this reasoning model leaderboard to compare the current ranking from the LLMBoard Reasoning Score after applying benchmark evidence requirements.

Updated 2026-09-08

On this page

  • Ranking
  • Overview
  • Leaders
  • Capability
  • Value
  • Composition
  • Top models
  • FAQ

Reasoning Model Leaderboard Ranking

This leaderboard ranking orders models by their LLMBoard Reasoning Score aggregated from eligible benchmark evidence.

30 of 117 rows
Columns

Show columns

Sort by
Rank
Model
LLMBoard Reasoning Score
Context
Official input / 1M
Official output / 1M
Rank01ModelOPGPT-6-AstraOpenAILLMBoard Reasoning Score100.00Context1.1MOfficial input / 1M$10Official output / 1M$50
Rank02ModelANClaude Opus 4.8AnthropicLLMBoard Reasoning Score98.88Context1MOfficial input / 1M$5Official output / 1M$25
Rank03ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score97.56Context1MOfficial input / 1M$2Official output / 1M$6
Rank04ModelOPGPT-5.5OpenAILLMBoard Reasoning Score95.94Context1.1MOfficial input / 1M$5Official output / 1M$30
Rank05ModelGOGemini 3.1 ProGoogleLLMBoard Reasoning Score93.51Context1MOfficial input / 1M$2Official output / 1M$12
Rank06ModelACQwen3.8 Flash NextAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score93.24ContextN/AOfficial input / 1MN/AOfficial output / 1MN/A
Rank07ModelACQwen3.8 FlashAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score93.24Context1MOfficial input / 1M$0.15Official output / 1M$0.47
Rank08ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score92.90Context1MOfficial input / 1M$2.5Official output / 1M$7.5
Rank09ModelOPGPT-5.4OpenAILLMBoard Reasoning Score88.94Context1MOfficial input / 1M$2.5Official output / 1M$15
Rank10ModelGOGemini 2.5 ProGoogleLLMBoard Reasoning Score87.05Context1MOfficial input / 1M$1.25Official output / 1M$10
Rank11ModelGOGemma 4 31BGoogleLLMBoard Reasoning Score86.59Context262.1KOfficial input / 1MN/AOfficial output / 1MN/A
Rank12ModelXIMiMo V2.5 ProXiaomiLLMBoard Reasoning Score83.88Context1MOfficial input / 1M$0.435Official output / 1M$0.87
Rank13ModelACQwen3.7 PlusAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score83.52Context1MOfficial input / 1M$0.50Official output / 1M$3
Rank14ModelACQwen3.5 397B A17BAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score83.36Context262.1KOfficial input / 1M$0.60Official output / 1M$3.6
Rank15ModelANClaude Opus 4.6AnthropicLLMBoard Reasoning Score81.23Context1MOfficial input / 1M$5Official output / 1M$25
Rank16ModelGOGemma 4 26B A4BGoogleLLMBoard Reasoning Score78.44Context262.1KOfficial input / 1MN/AOfficial output / 1MN/A
Rank17ModelOPGPT-5.2-ProOpenAILLMBoard Reasoning Score77.79Context400KOfficial input / 1M$21Official output / 1M$168
Rank18ModelACQwen3.8 27BAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score77.70Context262.1KOfficial input / 1MN/AOfficial output / 1MN/A
Rank19ModelANClaude Sonnet 3.5AnthropicLLMBoard Reasoning Score77.52Context200KOfficial input / 1MN/AOfficial output / 1MN/A
Rank20ModelANClaude Sonnet 4.6AnthropicLLMBoard Reasoning Score76.66Context200KOfficial input / 1M$3Official output / 1M$15
Rank21ModelOPGPT-5.2OpenAILLMBoard Reasoning Score73.72Context400KOfficial input / 1M$1.75Official output / 1M$14
Rank22ModelACQwen3.6 PlusAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score73.62Context1MOfficial input / 1M$0.50Official output / 1M$3
Rank23ModelANClaude Opus 3AnthropicLLMBoard Reasoning Score73.56Context200KOfficial input / 1MN/AOfficial output / 1MN/A
Rank24ModelGOGemini 3 ProGoogleLLMBoard Reasoning Score72.21Context1MOfficial input / 1MN/AOfficial output / 1MN/A
Rank25ModelCOCommand R+CohereLLMBoard Reasoning Score71.22Context128KOfficial input / 1M$2.5Official output / 1M$10
Rank26ModelACQwen3 235B A22B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Reasoning Score70.65Context262.1KOfficial input / 1MN/AOfficial output / 1MN/A
Rank27ModelGOGemini 3 FlashGoogleLLMBoard Reasoning Score70.44Context1MOfficial input / 1M$0.50Official output / 1M$3
Rank28ModelDEDeepSeek-R1DeepSeekLLMBoard Reasoning Score70.01Context131.1KOfficial input / 1MN/AOfficial output / 1MN/A
Rank29ModelGOGemma 4 12BGoogleLLMBoard Reasoning Score69.47ContextN/AOfficial input / 1MN/AOfficial output / 1MN/A
Rank30ModelGOGemma 2 27BGoogleLLMBoard Reasoning Score68.58ContextN/AOfficial input / 1MN/AOfficial output / 1MN/A

Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.

What this leaderboard says

Key findings from the current reasoning leaderboard ranking and its supporting benchmark coverage.

Ranked models117
Vendors20
Open-weight models70
Official prices49

GPT-6-Astra leads this page at 100.00, 1.12 points ahead of Claude Opus 4.8.

Qwen3.8 Max is the highest-ranked open-weight option at #3. Nova Micro has the lowest official input price among ranked models.

Reasoning Leaderboard Leaders

The leading reasoning scores in this benchmark-backed ranking, with the axis focused on the competitive leaderboard range.

Top reasoning models

Reasoning Benchmark Profile

Compare reasoning benchmark strength with each model's broader capability score and leaderboard position.

ModelReasoningKnowledgeMathCodingInstruction
GPT-6-AstraOpenAI100.00N/AN/A94.72N/A
Claude Opus 4.8Anthropic98.8862.75N/A83.59N/A
Qwen3.8 MaxAlibaba Cloud / Qwen Team97.56N/AN/A71.34N/A
GPT-5.5OpenAI95.94N/AN/A71.24N/A
Gemini 3.1 ProGoogle93.5189.16N/A54.52N/A
31.4450.1468.8487.53106.23-7.9120.7749.4478.12106.80LLMBoard Reasoning ScoreLLMBoard Overall Score
80 models with comparable dataOpen a model by selecting its logo

Capability at each price point

Official vendor API prices plotted against the reasoning metric used on this page. Price remains a separate decision signal.

020406080100$0.05$0.1$0.2$0.5$1.0$2.0$5.0$10LLMBoard Reasoning ScoreOfficial API price blend (8:1 input/output), USD / 1M
Efficient frontier49 models with official PAYG prices

Reasoning Leaderboard Composition

Vendor concentration and model access among the first 25 products in the ranking.

Top-model vendor mix

25 models
Alibaba Cloud / Qwen Team8
OpenAI5
Anthropic5
Google5
Xiaomi1
Cohere1

Access model

Open weights
8
Closed weights
17
Official price available
18

The Top AI Models for Reasoning

A data-backed look at the first five models in this reasoning leaderboard ranking, including benchmark context, price and output speed where available.

About this leaderboard

Common questions about the Best AI for Reasoning leaderboard, benchmark evidence and ranking method.

What is the best model on Best AI for Reasoning?

GPT-6-Astra is currently ranked first with a reasoning score of 100.00.

Which three AI models lead this capability ranking?

The current leaders are GPT-6-Astra (100.00 LLMBoard Reasoning Score), Claude Opus 4.8 (98.88 LLMBoard Reasoning Score), and Qwen3.8 Max (97.56 LLMBoard Reasoning Score).

Which ranked model has the lowest official input price?

Nova Micro has the lowest current official input price among ranked models at $0.04 per 1M tokens.

Which ranked models have the highest measured output speed?

The fastest matched records are Llama 3.1 8B (2,047.00 tok/s via Cerebras), Llama 3.1 70B (1,204.00 tok/s via Cerebras), and GPT-4.1-mini (467.39 tok/s via OpenAI).

Which open-weight model ranks highest?

Qwen3.8 Max is the highest-ranked open-weight model at rank #3.

Which model offers the longest context in this ranking?

Gemini 1.5 Pro has the largest listed context window at 2.1M tokens.

How is this leaderboard ranked?

Each model appears once using its current scored version. The leaderboard ranking follows the capability named in the title, while the overall page uses the LLMBoard score aggregated from eligible benchmark evidence.

How many models are included?

This page currently ranks 117 unique model products.

Where does pricing come from?

Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.

Does Arena affect the score?

No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.

Top reasoning modelGPT-6-Astra100.00 reasoning score
Best open-weight modelQwen3.8 MaxRank #3
Lowest official inputNova Micro$0.04 / 1M
Largest contextGemini 1.5 Pro2.1M tokens

Ranking basisThis reasoning AI model leaderboard uses the reasoning capability score shown on this page. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    OP
    GPT-6-AstraOpenAI
    LLMBoard Reasoning Score
    100.00
    Price
    $10 input / $50 output per 1M tokens
    Speed
    30.66 tok/s via OpenAI

    Strengths

    • Ranks #1 with a 100.00 LLMBoard Reasoning Score
    • 1.1M maximum context tokens

    Considerations

    • 40% benchmark coverage
    • Weights are not marked as open
  2. 02
    AN
    Claude Opus 4.8Anthropic
    LLMBoard Reasoning Score
    98.88
    Price
    $5 input / $25 output per 1M tokens
    Speed
    42.00 tok/s via Vertex AI

    Strengths

    • Ranks #2 with a 98.88 LLMBoard Reasoning Score
    • 1M maximum context tokens

    Considerations

    • 60% benchmark coverage
    • Weights are not marked as open
  3. 03
    AC
    Alibaba Cloud / Qwen Team
    LLMBoard Reasoning Score
    97.56
    Price
    $2 input / $6 output per 1M tokens
    Speed
    60.87 tok/s via DeepInfra

    Strengths

    • Ranks #3 with a 97.56 LLMBoard Reasoning Score
    • Open-weight model version
    • Lowest official input price among the top models

    Considerations

    • 40% benchmark coverage
  4. 04
    OP
    OpenAI
    LLMBoard Reasoning Score
    95.94
    Price
    $5 input / $30 output per 1M tokens
    Speed
    134.94 tok/s via OpenAI

    Strengths

    • Ranks #4 with a 95.94 LLMBoard Reasoning Score
    • 1.1M maximum context tokens

    Considerations

    • 40% benchmark coverage
    • Weights are not marked as open
  5. 05
    GO
    Google
    LLMBoard Reasoning Score
    93.51
    Price
    $2 input / $12 output per 1M tokens
    Speed
    43.79 tok/s via Google

    Strengths

    • Ranks #5 with a 93.51 LLMBoard Reasoning Score
    • Lowest official input price among the top models
    • 1M maximum context tokens

    Considerations

    • 60% benchmark coverage
    • Weights are not marked as open

Selection summary

Best AI Models for Reasoning

GPT-6-Astra is currently the best-ranked LLM for reasoning with a 100.00 reasoning score. The score is a relative ranking signal, so price, speed, context and evidence coverage should still be checked separately.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Qwen3.8 Max
GPT-5.5
Gemini 3.1 Pro
Capability rank #1GPT-6-Astra30.66 tok/s via OpenAI
Capability rank #2Claude Opus 4.842.00 tok/s via Vertex AI
Capability rank #3Qwen3.8 Max60.87 tok/s via DeepInfra