llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Capability ranking

Best AI for Instruction Following

Use this instruction-following model leaderboard to compare the current ranking from the LLMBoard Instruction Score after applying benchmark evidence requirements.

Updated 2026-09-08

On this page

  • Ranking
  • Overview
  • Leaders
  • Capability
  • Value
  • Composition
  • Top models
  • FAQ

Instruction Model Leaderboard Ranking

This leaderboard ranking orders models by their LLMBoard Instruction Score aggregated from eligible benchmark evidence.

30 of 44 rows
Columns

Show columns

Sort by
Rank
Model
LLMBoard Instruction Score
Context
Official input / 1M
Official output / 1M
Output speed
Rank01ModelOPo3 miniOpenAILLMBoard Instruction Score92.47Context200KOfficial input / 1M$1.1Official output / 1M$4.4Output speed115.00 tok/s
Rank02ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamLLMBoard Instruction Score87.39Context1MOfficial input / 1M$2.5Official output / 1M$7.5Output speedN/A
Rank03ModelACQwen3.7 PlusAlibaba Cloud / Qwen TeamLLMBoard Instruction Score86.62Context1MOfficial input / 1M$0.50Official output / 1M$3Output speedN/A
Rank04ModelNVNemotron 3 UltraNVIDIALLMBoard Instruction Score86.01ContextN/AOfficial input / 1M$0.50Official output / 1M$2.5Output speedN/A
Rank05ModelACQwen3 235B A22B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score85.01Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank06ModelACQwen3 235B A22BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score78.10Context262.1KOfficial input / 1M$0.70Official output / 1M$2.8Output speed21.74 tok/s
Rank07ModelACQwen3 VL 235B A22BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score77.96Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank08ModelACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score77.11Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank09ModelACQwen3.6 PlusAlibaba Cloud / Qwen TeamLLMBoard Instruction Score76.91Context1MOfficial input / 1M$0.50Official output / 1M$3Output speedN/A
Rank10ModelACQwen3.5 27BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score74.52Context262.1KOfficial input / 1M$0.30Official output / 1M$2.4Output speedN/A
Rank11ModelACQwen3.5 397B A17BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score74.07Context262.1KOfficial input / 1M$0.60Official output / 1M$3.6Output speedN/A
Rank12ModelNVNemotron 3 SuperNVIDIALLMBoard Instruction Score73.69Context262.1KOfficial input / 1M$0.20Official output / 1M$0.80Output speedN/A
Rank13ModelMAKimi K2Moonshot AILLMBoard Instruction Score73.47ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank14ModelACQwen3.5 122B A10BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score72.24Context262.1KOfficial input / 1M$0.40Official output / 1M$3.2Output speedN/A
Rank15ModelACQwen3 Next 80B A3BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score68.71Context65.5KOfficial input / 1M$0.50Official output / 1M$2Output speedN/A
Rank16ModelACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score62.97ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank17ModelACQwen3 Next 80B A3B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score62.18Context65.5KOfficial input / 1M$0.50Official output / 1M$6Output speedN/A
Rank18ModelGOGemma 3 27BGoogleLLMBoard Instruction Score61.50Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33.00 tok/s
Rank19ModelCOCommand A+CohereLLMBoard Instruction Score61.37ContextN/AOfficial input / 1M$2.5Official output / 1M$10Output speedN/A
Rank20ModelACQwen3.5 35B A3BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score58.82Context262.1KOfficial input / 1M$0.25Official output / 1M$2Output speedN/A
Rank21ModelLALFM2.5 2.6BLiquid AILLMBoard Instruction Score54.31ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank22ModelGOGemma 3 12BGoogleLLMBoard Instruction Score51.39Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33.00 tok/s
Rank23ModelGOGemma 3 4BGoogleLLMBoard Instruction Score50.92Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33.00 tok/s
Rank24ModelACQwen3.5 9BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score50.67ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank25ModelACQwen3.5 4BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score45.79ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank26ModelOPGPT-4.5OpenAILLMBoard Instruction Score45.00Context128KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed50.00 tok/s
Rank27ModelACQwen2.5 72BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score44.03Context131.1KOfficial input / 1M$1.4Official output / 1M$5.6Output speed10.00 tok/s
Rank28ModelACQwen3 VL 32BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score42.19ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank29ModelACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score41.96Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank30ModelACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score40.64Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A

Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.

What this leaderboard says

Key findings from the current instruction leaderboard ranking and its supporting benchmark coverage.

Ranked models44
Vendors7
Open-weight models35
Official prices20

o3 mini leads this page at 92.47, 5.08 points ahead of Qwen3.7 Max.

Nemotron 3 Ultra is the highest-ranked open-weight option at #4. GPT-4.1-nano has the lowest official input price among ranked models.

Instruction Leaderboard Leaders

The leading instruction scores in this benchmark-backed ranking, with the axis focused on the competitive leaderboard range.

Top instruction models

Instruction Benchmark Profile

Compare instruction benchmark strength with each model's broader capability score and leaderboard position.

ModelReasoningKnowledgeMathCodingInstruction
o3 miniOpenAI65.4646.7963.78N/A92.47
Qwen3.7 MaxAlibaba Cloud / Qwen Team92.9088.6593.4280.8687.39
Qwen3.7 PlusAlibaba Cloud / Qwen Team83.5282.3469.2659.6086.62
Nemotron 3 UltraNVIDIAN/A76.62N/A50.2086.01
Qwen3 235B A22B ThinkingAlibaba Cloud / Qwen Team70.6560.7864.16N/A85.01
-9.2518.4946.2473.98101.72-6.1316.0938.3160.5482.76LLMBoard Instruction ScoreLLMBoard Overall Score
44 models with comparable dataOpen a model by selecting its logo

Capability at each price point

Official vendor API prices plotted against the instruction metric used on this page. Price remains a separate decision signal.

020406080100$0.2$0.5$1.0$2.0LLMBoard Instruction ScoreOfficial API price blend (8:1 input/output), USD / 1M
Efficient frontier20 models with official PAYG prices

Instruction Leaderboard Composition

Vendor concentration and model access among the first 25 products in the ranking.

Top-model vendor mix

25 models
Alibaba Cloud / Qwen Team16
Google3
NVIDIA2
OpenAI1
Moonshot AI1
Cohere1
Liquid AI1

Access model

Open weights
21
Closed weights
4
Official price available
14

The Top AI Models for Instruction Following

A data-backed look at the first five models in this instruction following leaderboard ranking, including benchmark context, price and output speed where available.

About this leaderboard

Common questions about the Best AI for Instruction Following leaderboard, benchmark evidence and ranking method.

What is the best model on Best AI for Instruction Following?

o3 mini is currently ranked first with a instruction score of 92.47.

Which three AI models lead this capability ranking?

The current leaders are o3 mini (92.47 LLMBoard Instruction Score), Qwen3.7 Max (87.39 LLMBoard Instruction Score), and Qwen3.7 Plus (86.62 LLMBoard Instruction Score).

Which ranked model has the lowest official input price?

GPT-4.1-nano has the lowest current official input price among ranked models at $0.1 per 1M tokens.

Which ranked models have the highest measured output speed?

The fastest matched records are GPT-4.1-mini (467.39 tok/s via OpenAI), GPT-4.1-nano (138.14 tok/s via OpenAI), and GPT-4o (132.00 tok/s via OpenAI).

Which open-weight model ranks highest?

Nemotron 3 Ultra is the highest-ranked open-weight model at rank #4.

Which model offers the longest context in this ranking?

GPT-4.1 has the largest listed context window at 1M tokens.

How is this leaderboard ranked?

Each model appears once using its current scored version. The leaderboard ranking follows the capability named in the title, while the overall page uses the LLMBoard score aggregated from eligible benchmark evidence.

How many models are included?

This page currently ranks 44 unique model products.

Where does pricing come from?

Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.

Does Arena affect the score?

No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.

Top instruction modelo3 mini92.47 instruction score
Best open-weight modelNemotron 3 UltraRank #4
Lowest official inputGPT-4.1-nano$0.1 / 1M
Largest contextGPT-4.11M tokens

Ranking basisThis instruction following AI model leaderboard uses the instruction capability score shown on this page. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    OP
    o3 miniOpenAI
    LLMBoard Instruction Score
    92.47
    Price
    $1.1 input / $4.4 output per 1M tokens
    Speed
    115.00 tok/s via Azure

    Strengths

    • Ranks #1 with a 92.47 LLMBoard Instruction Score
    • 200K maximum context tokens

    Considerations

    • Weights are not marked as open
  2. 02
    AC
    Qwen3.7 MaxAlibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    87.39
    Price
    $2.5 input / $7.5 output per 1M tokens

    Strengths

    • Ranks #2 with a 87.39 LLMBoard Instruction Score
    • 1M maximum context tokens

    Considerations

    • Weights are not marked as open
  3. 03
    AC
    Alibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    86.62
    Price
    $0.5 input / $3 output per 1M tokens

    Strengths

    • Ranks #3 with a 86.62 LLMBoard Instruction Score
    • Lowest official input price among the top models
    • 1M maximum context tokens

    Considerations

    • Weights are not marked as open
  4. 04
    NV
    NVIDIA
    LLMBoard Instruction Score
    86.01
    Price
    $0.5 input / $2.5 output per 1M tokens

    Strengths

    • Ranks #4 with a 86.01 LLMBoard Instruction Score
    • Open-weight model version
    • Lowest official input price among the top models

    Considerations

    • 60% benchmark coverage
  5. 05
    AC
    Alibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    85.01

    Strengths

    • Ranks #5 with a 85.01 LLMBoard Instruction Score
    • Open-weight model version
    • 262K maximum context tokens

Selection summary

Best AI Models for Instruction Following

o3 mini is currently the best-ranked LLM for instruction following with a 92.47 instruction score. The score is a relative ranking signal, so price, speed, context and evidence coverage should still be checked separately.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Qwen3.7 Plus
Nemotron 3 Ultra
Qwen3 235B A22B Thinking
Capability rank #1o3 mini115.00 tok/s via Azure
Capability rank #2Qwen3.7 Max$2.5 input / $7.5 output per 1M tokens
Capability rank #3Qwen3.7 Plus$0.5 input / $3 output per 1M tokens