llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Runtime performance

AI Model Speed and Latency Leaderboard

Use this runtime leaderboard to compare the current output-speed ranking, catalog latency, provider time to first token and reliability. Runtime remains independent from capability benchmark scores and price.

Data as of 2026-09-08

Model-provider records218175 canonical versions
Complete speed + latency213Comparable catalog records
Provider aggregates16Observed 7-day metrics
Fastest catalog speed2,220.00 tok/sLlama 3.3 70B Instruct via Cerebras

On this page

  • Model ranking
  • Speed vs. latency
  • Provider view
  • Top models
  • Definitions
  • FAQ

Model-provider Runtime Leaderboard

This leaderboard ranking uses generated output tokens received per second. Sort by speed or catalog latency to match your workload; capability benchmark results do not change the order.

30 of 213 rows
Columns

Show columns

Sort by
Model
Provider
Output Speed
Catalog Latency
LLMBoard score
Official input / 1M
Max Input
Updated
ModelMELlama 3.3 70B InstructMetaProviderCerebrasOutput Speed2,220.00 tok/sCatalog Latency0.65 sLLMBoard score20.20Official input / 1MN/AMax Input128KUpdatedSep 8, 2026
ModelMELlama 3.1 8B InstructMetaProviderCerebrasOutput Speed2,047.00 tok/sCatalog Latency0.20 sLLMBoard scoreN/AOfficial input / 1MN/AMax Input131.1KUpdatedSep 8, 2026
ModelMAMin istral 3 (3B Reasoning 2512)Mistral AIProviderMistral AIOutput Speed1,222.00 tok/sCatalog Latency0.16 sLLMBoard score19.35Official input / 1M$0.04Max Input131.1KUpdatedSep 8, 2026
ModelMELlama 3.1 70B InstructMetaProviderCerebrasOutput Speed1,204.00 tok/sCatalog Latency0.20 sLLMBoard score11.47Official input / 1MN/AMax Input128KUpdatedSep 8, 2026
ModelOPGPT OSS 20BOpenAIProviderGroqOutput Speed1,000.00 tok/sCatalog Latency0.38 sLLMBoard score45.98Official input / 1MN/AMax Input131KUpdatedSep 8, 2026
ModelMELlama 4 ScoutMetaProviderGroqOutput Speed776.10 tok/sCatalog Latency1.08 sLLMBoard score12.39Official input / 1MN/AMax Input10MUpdatedSep 8, 2026
ModelMELlama 3.1 8B InstructMetaProviderGroqOutput Speed750.00 tok/sCatalog Latency0.50 sLLMBoard scoreN/AOfficial input / 1MN/AMax Input131.1KUpdatedSep 8, 2026
ModelOPGPT-5 nanoOpenAIProviderOpenAIOutput Speed500.00 tok/sCatalog Latency0.30 sLLMBoard score29.51Official input / 1M$0.05Max Input400KUpdatedSep 8, 2026
ModelOPGPT OSS 120BOpenAIProviderGroqOutput Speed500.00 tok/sCatalog Latency0.50 sLLMBoard score44.05Official input / 1MN/AMax Input131KUpdatedSep 8, 2026
ModelZAGLM-5.3Zhipu AIProviderFriendliAIOutput Speed472.69 tok/sCatalog Latency1.50 sLLMBoard score88.97Official input / 1M$1.4Max Input1MUpdatedSep 8, 2026
ModelOPGPT-4.1 miniOpenAIProviderOpenAIOutput Speed467.39 tok/sCatalog Latency0.82 sLLMBoard score23.57Official input / 1M$0.40Max Input1MUpdatedSep 8, 2026
ModelOPGPT-5.1 MediumOpenAIProviderOpenAIOutput Speed369.03 tok/sCatalog Latency2.65 sLLMBoard score60.79Official input / 1M$1.25Max Input400KUpdatedSep 8, 2026
ModelINMercury 2InceptionProviderInceptionOutput Speed346.33 tok/sCatalog Latency1.64 sLLMBoard score37.76Official input / 1M$0.25Max Input128KUpdatedSep 8, 2026
ModelMELlama 4 MaverickMetaProviderGroqOutput Speed307.30 tok/sCatalog Latency0.27 sLLMBoard score23.44Official input / 1MN/AMax Input1MUpdatedSep 8, 2026
ModelMELlama 3.3 70B InstructMetaProviderGroqOutput Speed268.00 tok/sCatalog Latency0.65 sLLMBoard score20.20Official input / 1MN/AMax Input128KUpdatedSep 8, 2026
ModelMELlama 3.1 70B InstructMetaProviderGroqOutput Speed250.00 tok/sCatalog Latency0.50 sLLMBoard score11.47Official input / 1MN/AMax Input128KUpdatedSep 8, 2026
ModelMAMinistral 3 (8B Reasoning 2512)Mistral AIProviderMistral AIOutput Speed237.50 tok/sCatalog Latency0.34 sLLMBoard score27.87Official input / 1M$0.10Max Input262.1KUpdatedSep 8, 2026
ModelGOGemini 3.5 FlashGoogleProviderGoogleOutput Speed210.94 tok/sCatalog Latency5.84 sLLMBoard score76.67Official input / 1M$1.5Max Input1MUpdatedSep 8, 2026
ModelOPGPT-5.5 InstantOpenAIProviderOpenAIOutput Speed206.89 tok/sCatalog Latency1.89 sLLMBoard score52.10Official input / 1MN/AMax Input400KUpdatedSep 8, 2026
ModelOPGPT-5.1 Codex MiniOpenAIProviderOpenAIOutput Speed200.00 tok/sCatalog Latency1.00 sLLMBoard scoreN/AOfficial input / 1MN/AMax Input400KUpdatedSep 8, 2026
ModelGOGemini 2.0 FlashGoogleProviderGoogleOutput Speed183.00 tok/sCatalog Latency0.40 sLLMBoard score27.37Official input / 1MN/AMax Input1MUpdatedSep 8, 2026
ModelMELlama 3.2 3B InstructMetaProviderDeepInfraOutput Speed171.50 tok/sCatalog Latency0.24 sLLMBoard scoreN/AOfficial input / 1MN/AMax Input128KUpdatedSep 8, 2026
ModelOPGPT-5 miniOpenAIProviderOpenAIOutput Speed167.61 tok/sCatalog Latency1.51 sLLMBoard score43.73Official input / 1M$0.25Max Input400KUpdatedSep 8, 2026
ModelMAMistral Small 4Mistral AIProviderMistral AIOutput Speed163.71 tok/sCatalog Latency1.89 sLLMBoard score32.09Official input / 1M$0.15Max Input256KUpdatedSep 8, 2026
ModelGOGemini 1.5 Flash 8BGoogleProviderGoogleOutput Speed150.00 tok/sCatalog Latency0.30 sLLMBoard scoreN/AOfficial input / 1MN/AMax Input1MUpdatedSep 8, 2026
ModelGOGemini 1.5 FlashGoogleProviderGoogleOutput Speed150.00 tok/sCatalog Latency0.30 sLLMBoard score10.18Official input / 1MN/AMax Input1MUpdatedSep 8, 2026
ModelOPGPT-4.1 nanoOpenAIProviderOpenAIOutput Speed138.14 tok/sCatalog Latency1.35 sLLMBoard score2.48Official input / 1M$0.10Max Input1MUpdatedSep 8, 2026
ModelMADevstral MediumMistral AIProviderMistral AIOutput Speed137.10 tok/sCatalog Latency0.23 sLLMBoard score23.28Official input / 1M$0.40Max Input128KUpdatedSep 8, 2026
ModelMADevstral Small 1.1Mistral AIProviderMistral AIOutput Speed137.10 tok/sCatalog Latency0.23 sLLMBoard score16.11Official input / 1M$0.10Max Input128KUpdatedSep 8, 2026
ModelMAMistral Small 3.1 24B BaseMistral AIProviderMistral AIOutput Speed137.10 tok/sCatalog Latency0.23 sLLMBoard score6.55Official input / 1MN/AMax Input128KUpdatedSep 8, 2026

Lowest catalog latency in the current records: Min istral 3 (3B Reasoning 2512) via Mistral AI at 0.16s.

Output speed vs. catalog latency

Each point is one canonical model version offered by one provider in the runtime leaderboard. Upper-left is generally preferable: higher token generation speed with lower catalog latency.

0.10 tok/s0.50 tok/s2.00 tok/s10.00 tok/s20.00 tok/s100.00 tok/s500.00 tok/s2,000.00 tok/s0.20 s0.50 s1.00 s2.00 s5.00 s10.00 s20.00 sOutput speedCatalog latency
Each mark is one model-provider record213 records

Provider runtime: observed over 7 days

This provider ranking uses observed character throughput and time to first token. Character throughput is not converted into model token throughput or benchmark capability.

16 rows
Columns

Show columns

Sort by
Provider
P50 TTFT
P95 TTFT
P50 Throughput
P5 Throughput
7d Calls
Success Rate
Error Rate
Updated
ProviderxAIP50 TTFT1,096.00 msP95 TTFT1,704.80 msP50 Throughput81.65 char/sP5 Throughput18.04 char/s7d Calls56Success Rate96.40%Error Rate3.60%UpdatedSep 8, 2026
ProviderFriendliAIP50 TTFT1,163.50 msP95 TTFT14,865.50 msP50 Throughput263.98 char/sP5 Throughput40.68 char/s7d Calls36Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderMistral AIP50 TTFT1,350.00 msP95 TTFT1,968.60 msP50 Throughput131.43 char/sP5 Throughput59.14 char/s7d Calls13Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderDeepSeekP50 TTFT1,364.00 msP95 TTFT1,998.50 msP50 Throughput183.71 char/sP5 Throughput45.48 char/s7d Calls14Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderStepFunP50 TTFT1,418.00 msP95 TTFT1,895.60 msP50 Throughput220.72 char/sP5 Throughput21.16 char/s7d Calls139Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderInceptionP50 TTFT1,598.00 msP95 TTFT1,644.80 msP50 Throughput353.66 char/sP5 Throughput346.33 char/s7d Calls3Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderGoogleP50 TTFT1,734.00 msP95 TTFT8,806.50 msP50 Throughput192.02 char/sP5 Throughput28.71 char/s7d Calls4,034Success Rate98.80%Error Rate1.20%UpdatedSep 8, 2026
ProviderDeepInfraP50 TTFT1,754.00 msP95 TTFT12,446.25 msP50 Throughput65.69 char/sP5 Throughput7.94 char/s7d Calls311Success Rate99.00%Error Rate1.00%UpdatedSep 8, 2026
ProviderMeta Model APIP50 TTFT1,942.50 msP95 TTFT7,620.10 msP50 Throughput33.35 char/sP5 Throughput8.49 char/s7d Calls548Success Rate99.60%Error Rate0.40%UpdatedSep 8, 2026
ProviderOpenAIP50 TTFT1,978.00 msP95 TTFT10,244.20 msP50 Throughput162.32 char/sP5 Throughput14.13 char/s7d Calls534Success Rate95.10%Error Rate4.90%UpdatedSep 8, 2026
ProviderAnthropicP50 TTFT3,427.00 msP95 TTFT13,175.20 msP50 Throughput24.41 char/sP5 Throughput8.04 char/s7d Calls558Success Rate97.70%Error Rate2.30%UpdatedSep 8, 2026
ProviderXiaomiP50 TTFT3,765.00 msP95 TTFT15,612.35 msP50 Throughput25.00 char/sP5 Throughput4.22 char/s7d Calls4Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026
ProviderZAIP50 TTFT5,335.00 msP95 TTFT9,690.30 msP50 Throughput37.95 char/sP5 Throughput4.22 char/s7d Calls63Success Rate95.20%Error Rate4.80%UpdatedSep 8, 2026
ProviderMiniMaxP50 TTFT5,837.00 msP95 TTFT6,046.60 msP50 Throughput8.92 char/sP5 Throughput6.61 char/s7d Calls9Success Rate55.60%Error Rate44.40%UpdatedSep 8, 2026
ProviderAzureP50 TTFTN/AP95 TTFTN/AP50 ThroughputN/AP5 ThroughputN/A7d Calls0Success RateN/AError RateN/AUpdatedSep 8, 2026
ProviderMoonshot AIP50 TTFTN/AP95 TTFTN/AP50 Throughput2.93 char/sP5 Throughput1.24 char/s7d Calls6Success Rate100.00%Error Rate0.00%UpdatedSep 8, 2026

The Top AI Models for Speed and Latency

The five fastest records in the current runtime leaderboard, with catalog latency, official token price and independent benchmark context shown where available.

How to read the runtime leaderboard

Runtime leaderboard metrics answer deployment questions and should not be read as model capability scores or a quality benchmark ranking.

Output SpeedGenerated output tokens received per second. Higher is faster.
Catalog LatencyProvider catalog latency in seconds. Lower is faster; it is not the 7-day TTFT measure.
TTFTTime to first token from observed provider calls. P50 is typical; P95 shows slower-tail behavior.
Observed ThroughputCharacters per second across a 7-day provider window. It remains separate from output tokens per second.
ReliabilityCalls, success rate, and error rate describe the observed provider window, not model capability.
Independent SignalRuntime does not change the LLMBoard capability leaderboard, ranking bands, or benchmark evidence coverage.

Runtime source: LLM Stats / ZeroEval snapshot. Speed follows the public output-tokens-per-second convention. Read the source definition.

FAQ

Common questions about model speed, latency and provider runtime.

Which model has the fastest output speed?

Llama 3.3 70B Instruct currently records the fastest output speed at 2,220.00 tok/s via Cerebras.

Which model has the lowest catalog latency?

Min istral 3 (3B Reasoning 2512) has the lowest catalog latency at 0.16s via Mistral AI.

What are the three fastest model records?

The current fastest records are Llama 3.3 70B Instruct (2,220.00 tok/s via Cerebras, 0.65s catalog latency), Llama 3.1 8B Instruct (2,047.00 tok/s via Cerebras, 0.20s catalog latency), and Min istral 3 (3B Reasoning 2512) (1,222.00 tok/s via Mistral AI, 0.16s catalog latency).

Does faster output mean a better model?

No. Output speed and latency describe runtime behavior for a provider record. They do not change the LLMBoard capability score or benchmark ranking.

Which provider has the best observed runtime?

Inception leads the available 7-day observed throughput records. Provider observations use characters per second and should not be read as model tokens per second.

10.00 char/s20.00 char/s50.00 char/s100.00 char/s200.00 char/s1,000.00 ms2,000.00 ms5,000.00 msMedian observed throughputMedian TTFT
Each mark is one provider aggregate14 records
Are runtime prices included in the speed ranking?

No. Official input and output prices are shown as separate context fields; they do not affect the runtime leaderboard order.

Ranking basisThis speed and latency AI model leaderboard uses descending provider catalog output speed in generated tokens per second. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    ME
    Llama 3.3 70B InstructMeta
    Output speed
    2,220.00 tok/s
    Speed
    2,220.00 tok/s via Cerebras

    Strengths

    • Ranks #1 by output tokens per second
    • 0.65s catalog latency in the same provider record
    • 128K maximum input tokens

    Considerations

    • This speed is provider-specific and does not measure model capability
    • Catalog latency is not the same as observed 7-day TTFT
  2. 02
    ME
    Llama 3.1 8B InstructMeta
    Output speed
    2,047.00 tok/s
    Speed
    2,047.00 tok/s via Cerebras

    Strengths

    • Ranks #2 by output tokens per second
    • 0.20s catalog latency in the same provider record
    • 131K maximum input tokens

    Considerations

    • This speed is provider-specific and does not measure model capability
    • Catalog latency is not the same as observed 7-day TTFT
  3. 03
    MA
    Mistral AI
    Output speed
    1,222.00 tok/s
    Price
    $0.0 input / $0.0 output per 1M tokens
    Speed
    1,222.00 tok/s via Mistral AI

    Strengths

    • Ranks #3 by output tokens per second
    • 0.16s catalog latency in the same provider record
    • 131K maximum input tokens

    Considerations

    • This speed is provider-specific and does not measure model capability
    • Catalog latency is not the same as observed 7-day TTFT
  4. 04
    ME
    Meta
    Output speed
    1,204.00 tok/s
    Speed
    1,204.00 tok/s via Cerebras

    Strengths

    • Ranks #4 by output tokens per second
    • 0.20s catalog latency in the same provider record
    • 128K maximum input tokens

    Considerations

    • This speed is provider-specific and does not measure model capability
    • Catalog latency is not the same as observed 7-day TTFT
  5. 05
    OP
    OpenAI
    Output speed
    1,000.00 tok/s
    Speed
    1,000.00 tok/s via Groq

    Strengths

    • Ranks #5 by output tokens per second
    • 0.38s catalog latency in the same provider record
    • 131K maximum input tokens

    Considerations

    • This speed is provider-specific and does not measure model capability
    • Catalog latency is not the same as observed 7-day TTFT

Selection summary

Best AI Models for Speed and Latency

Llama 3.3 70B Instruct has the fastest current catalog record at 2,220.00 tok/s. The best AI model for runtime-sensitive work also depends on latency, provider reliability, price and workload-specific quality.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Min istral 3 (3B Reasoning 2512)
Llama 3.1 70B Instruct
GPT OSS 20B
Speed rank #1Llama 3.3 70B Instruct2,220.00 tok/s via Cerebras
Speed rank #2Llama 3.1 8B Instruct2,047.00 tok/s via Cerebras
Speed rank #3Min istral 3 (3B Reasoning 2512)1,222.00 tok/s via Mistral AI ยท $0.0 input / $0.0 output per 1M tokens