llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Scoring & Data

Use rankings to narrow the field, not end the decision

Start with the aggregate view, then review capability rankings, original sources, price, and your own use case.

Scoring Principles

A ranking summarizes evidence

Different sources are made comparable while their original meaning remains visible.

How to read the overall score

The score shows a model relative position across the current eligible evidence set. It is useful for screening, but it is not IQ and a score of 97 is not seven percent more capable than a score of 90.

How missing results are handled

Missing evaluations are not treated as zero. Coverage is shown alongside the score, and limited evidence reduces ranking confidence so models cannot lead on just a few results.

How related benchmarks are handled

Closely related evaluations are organized into benchmark families so one capability does not gain extra influence simply because it has many similar tests.

How model names are matched

A model can use different names across sources. Confirmed aliases map to a specific model and version; records that cannot be matched reliably remain separate.

What evidence coverage means

Coverage shows how much relevant capability evidence is available. Broad coverage usually makes a rank more stable; limited coverage is better read as an early signal.

Why capability rankings differ

Coding, reasoning, math, knowledge, and instruction following each sort by their relevant capability measures. They are not renamed copies of the overall ranking.

Ranking status

Status describes evidence confidence, not whether a model is useful.

RankedCore capability coverage meets the current threshold for the main ranking.
ProvisionalUseful results exist, but coverage or benchmark-family diversity is still limited.

Signals that stay separate

Model selection does not have one universal number. Non-capability signals deserve their own comparisons.

CapabilityCan the model perform this type of task?
PriceDoes the operating cost fit the use case?
SpeedAre response and generation times fast enough?
AdoptionHow widely is it used? Adoption does not replace capability.

Output speed, catalog latency, observed TTFT, throughput, and reliability are published in a standalone runtime view. None of these fields are folded into capability scores.

Review the original evidence

Open a core benchmark or LM Arena to see its original metrics, ranking, and update date.

Browse core benchmarks