Capability ranking
Use this instruction-following model leaderboard to compare the current ranking from the LLMBoard Instruction Score after applying benchmark evidence requirements.
Updated 2026-09-08
This leaderboard ranking orders models by their LLMBoard Instruction Score aggregated from eligible benchmark evidence.
Rank | Model | LLMBoard Instruction Score | Context | Official input / 1M | Official output / 1M | Output speed |
|---|
| Rank01 | ModelOP | LLMBoard Instruction Score92.47 | Context200K | Official input / 1M$1.1 | Official output / 1M$4.4 | Output speed115.00 tok/s |
| Rank02 | ModelAC | LLMBoard Instruction Score87.39 | Context1M | Official input / 1M$2.5 | Official output / 1M$7.5 | Output speedN/A |
| Rank03 | ModelAC | LLMBoard Instruction Score86.62 | Context1M | Official input / 1M$0.50 | Official output / 1M$3 | Output speedN/A |
| Rank04 | ModelNV | LLMBoard Instruction Score86.01 | ContextN/A | Official input / 1M$0.50 | Official output / 1M$2.5 | Output speedN/A |
| Rank05 | ModelAC | LLMBoard Instruction Score85.01 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank06 | ModelAC | LLMBoard Instruction Score78.10 | Context262.1K | Official input / 1M$0.70 | Official output / 1M$2.8 | Output speed21.74 tok/s |
| Rank07 | ModelAC | LLMBoard Instruction Score77.96 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank08 | ModelAC | LLMBoard Instruction Score77.11 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank09 | ModelAC | LLMBoard Instruction Score76.91 | Context1M | Official input / 1M$0.50 | Official output / 1M$3 | Output speedN/A |
| Rank10 | ModelAC | LLMBoard Instruction Score74.52 | Context262.1K | Official input / 1M$0.30 | Official output / 1M$2.4 | Output speedN/A |
| Rank11 | ModelAC | LLMBoard Instruction Score74.07 | Context262.1K | Official input / 1M$0.60 | Official output / 1M$3.6 | Output speedN/A |
| Rank12 | ModelNV | LLMBoard Instruction Score73.69 | Context262.1K | Official input / 1M$0.20 | Official output / 1M$0.80 | Output speedN/A |
| Rank13 | ModelMA | LLMBoard Instruction Score73.47 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank14 | ModelAC | LLMBoard Instruction Score72.24 | Context262.1K | Official input / 1M$0.40 | Official output / 1M$3.2 | Output speedN/A |
| Rank15 | ModelAC | LLMBoard Instruction Score68.71 | Context65.5K | Official input / 1M$0.50 | Official output / 1M$2 | Output speedN/A |
| Rank16 | ModelAC | LLMBoard Instruction Score62.97 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank17 | ModelAC | LLMBoard Instruction Score62.18 | Context65.5K | Official input / 1M$0.50 | Official output / 1M$6 | Output speedN/A |
| Rank18 | ModelGO | LLMBoard Instruction Score61.50 | Context131.1K | Official input / 1MN/A | Official output / 1MN/A | Output speed33.00 tok/s |
| Rank19 | ModelCO | LLMBoard Instruction Score61.37 | ContextN/A | Official input / 1M$2.5 | Official output / 1M$10 | Output speedN/A |
| Rank20 | ModelAC | LLMBoard Instruction Score58.82 | Context262.1K | Official input / 1M$0.25 | Official output / 1M$2 | Output speedN/A |
| Rank21 | ModelLA | LLMBoard Instruction Score54.31 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank22 | ModelGO | LLMBoard Instruction Score51.39 | Context131.1K | Official input / 1MN/A | Official output / 1MN/A | Output speed33.00 tok/s |
| Rank23 | ModelGO | LLMBoard Instruction Score50.92 | Context131.1K | Official input / 1MN/A | Official output / 1MN/A | Output speed33.00 tok/s |
| Rank24 | ModelAC | LLMBoard Instruction Score50.67 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank25 | ModelAC | LLMBoard Instruction Score45.79 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank26 | ModelOP | LLMBoard Instruction Score45.00 | Context128K | Official input / 1MN/A | Official output / 1MN/A | Output speed50.00 tok/s |
| Rank27 | ModelAC | LLMBoard Instruction Score44.03 | Context131.1K | Official input / 1M$1.4 | Official output / 1M$5.6 | Output speed10.00 tok/s |
| Rank28 | ModelAC | LLMBoard Instruction Score42.19 | ContextN/A | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank29 | ModelAC | LLMBoard Instruction Score41.96 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
| Rank30 | ModelAC | LLMBoard Instruction Score40.64 | Context262.1K | Official input / 1MN/A | Official output / 1MN/A | Output speedN/A |
Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.
Key findings from the current instruction leaderboard ranking and its supporting benchmark coverage.
o3 mini leads this page at 92.47, 5.08 points ahead of Qwen3.7 Max.
Nemotron 3 Ultra is the highest-ranked open-weight option at #4. GPT-4.1-nano has the lowest official input price among ranked models.
The leading instruction scores in this benchmark-backed ranking, with the axis focused on the competitive leaderboard range.
Compare instruction benchmark strength with each model's broader capability score and leaderboard position.
Official vendor API prices plotted against the instruction metric used on this page. Price remains a separate decision signal.
Vendor concentration and model access among the first 25 products in the ranking.
A data-backed look at the first five models in this instruction following leaderboard ranking, including benchmark context, price and output speed where available.
Common questions about the Best AI for Instruction Following leaderboard, benchmark evidence and ranking method.
o3 mini is currently ranked first with a instruction score of 92.47.
The current leaders are o3 mini (92.47 LLMBoard Instruction Score), Qwen3.7 Max (87.39 LLMBoard Instruction Score), and Qwen3.7 Plus (86.62 LLMBoard Instruction Score).
GPT-4.1-nano has the lowest current official input price among ranked models at $0.1 per 1M tokens.
The fastest matched records are GPT-4.1-mini (467.39 tok/s via OpenAI), GPT-4.1-nano (138.14 tok/s via OpenAI), and GPT-4o (132.00 tok/s via OpenAI).
Nemotron 3 Ultra is the highest-ranked open-weight model at rank #4.
GPT-4.1 has the largest listed context window at 1M tokens.
Each model appears once using its current scored version. The leaderboard ranking follows the capability named in the title, while the overall page uses the LLMBoard score aggregated from eligible benchmark evidence.
This page currently ranks 44 unique model products.
Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.
No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.
Ranking basisThis instruction following AI model leaderboard uses the instruction capability score shown on this page. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.
Selection summary
o3 mini is currently the best-ranked LLM for instruction following with a 92.47 instruction score. The score is a relative ranking signal, so price, speed, context and evidence coverage should still be checked separately.
Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.