Runtime performance
Use this runtime leaderboard to compare the current output-speed ranking, catalog latency, provider time to first token and reliability. Runtime remains independent from capability benchmark scores and price.
Data as of 2026-09-08
This leaderboard ranking uses generated output tokens received per second. Sort by speed or catalog latency to match your workload; capability benchmark results do not change the order.
Model | Provider | Output Speed | Catalog Latency | LLMBoard score | Official input / 1M | Max Input | Updated |
|---|
| ModelME | ProviderCerebras | Output Speed2,220.00 tok/s | Catalog Latency0.65 s | LLMBoard score20.20 | Official input / 1MN/A | Max Input128K | Updated |
| ModelME | ProviderCerebras | Output Speed2,047.00 tok/s | Catalog Latency0.20 s | LLMBoard scoreN/A | Official input / 1MN/A | Max Input131.1K | Updated |
| ModelMA | ProviderMistral AI | Output Speed1,222.00 tok/s | Catalog Latency0.16 s | LLMBoard score19.35 | Official input / 1M$0.04 | Max Input131.1K | Updated |
| ModelME | ProviderCerebras | Output Speed1,204.00 tok/s | Catalog Latency0.20 s | LLMBoard score11.47 | Official input / 1MN/A | Max Input128K | Updated |
| ModelOP | ProviderGroq | Output Speed1,000.00 tok/s | Catalog Latency0.38 s | LLMBoard score45.98 | Official input / 1MN/A | Max Input131K | Updated |
| ModelME | ProviderGroq | Output Speed776.10 tok/s | Catalog Latency1.08 s | LLMBoard score12.39 | Official input / 1MN/A | Max Input10M | Updated |
| ModelME | ProviderGroq | Output Speed750.00 tok/s | Catalog Latency0.50 s | LLMBoard scoreN/A | Official input / 1MN/A | Max Input131.1K | Updated |
| ModelOP | ProviderOpenAI | Output Speed500.00 tok/s | Catalog Latency0.30 s | LLMBoard score29.51 | Official input / 1M$0.05 | Max Input400K | Updated |
| ModelOP | ProviderGroq | Output Speed500.00 tok/s | Catalog Latency0.50 s | LLMBoard score44.05 | Official input / 1MN/A | Max Input131K | Updated |
| ModelZA | ProviderFriendliAI | Output Speed472.69 tok/s | Catalog Latency1.50 s | LLMBoard score88.97 | Official input / 1M$1.4 | Max Input1M | Updated |
| ModelOP | ProviderOpenAI | Output Speed467.39 tok/s | Catalog Latency0.82 s | LLMBoard score23.57 | Official input / 1M$0.40 | Max Input1M | Updated |
| ModelOP | ProviderOpenAI | Output Speed369.03 tok/s | Catalog Latency2.65 s | LLMBoard score60.79 | Official input / 1M$1.25 | Max Input400K | Updated |
| ModelIN | ProviderInception | Output Speed346.33 tok/s | Catalog Latency1.64 s | LLMBoard score37.76 | Official input / 1M$0.25 | Max Input128K | Updated |
| ModelME | ProviderGroq | Output Speed307.30 tok/s | Catalog Latency0.27 s | LLMBoard score23.44 | Official input / 1MN/A | Max Input1M | Updated |
| ModelME | ProviderGroq | Output Speed268.00 tok/s | Catalog Latency0.65 s | LLMBoard score20.20 | Official input / 1MN/A | Max Input128K | Updated |
| ModelME | ProviderGroq | Output Speed250.00 tok/s | Catalog Latency0.50 s | LLMBoard score11.47 | Official input / 1MN/A | Max Input128K | Updated |
| ModelMA | ProviderMistral AI | Output Speed237.50 tok/s | Catalog Latency0.34 s | LLMBoard score27.87 | Official input / 1M$0.10 | Max Input262.1K | Updated |
| ModelGO | ProviderGoogle | Output Speed210.94 tok/s | Catalog Latency5.84 s | LLMBoard score76.67 | Official input / 1M$1.5 | Max Input1M | Updated |
| ModelOP | ProviderOpenAI | Output Speed206.89 tok/s | Catalog Latency1.89 s | LLMBoard score52.10 | Official input / 1MN/A | Max Input400K | Updated |
| ModelOP | ProviderOpenAI | Output Speed200.00 tok/s | Catalog Latency1.00 s | LLMBoard scoreN/A | Official input / 1MN/A | Max Input400K | Updated |
| ModelGO | ProviderGoogle | Output Speed183.00 tok/s | Catalog Latency0.40 s | LLMBoard score27.37 | Official input / 1MN/A | Max Input1M | Updated |
| ModelME | ProviderDeepInfra | Output Speed171.50 tok/s | Catalog Latency0.24 s | LLMBoard scoreN/A | Official input / 1MN/A | Max Input128K | Updated |
| ModelOP | ProviderOpenAI | Output Speed167.61 tok/s | Catalog Latency1.51 s | LLMBoard score43.73 | Official input / 1M$0.25 | Max Input400K | Updated |
| ModelMA | ProviderMistral AI | Output Speed163.71 tok/s | Catalog Latency1.89 s | LLMBoard score32.09 | Official input / 1M$0.15 | Max Input256K | Updated |
| ModelGO | ProviderGoogle | Output Speed150.00 tok/s | Catalog Latency0.30 s | LLMBoard scoreN/A | Official input / 1MN/A | Max Input1M | Updated |
| ModelGO | ProviderGoogle | Output Speed150.00 tok/s | Catalog Latency0.30 s | LLMBoard score10.18 | Official input / 1MN/A | Max Input1M | Updated |
| ModelOP | ProviderOpenAI | Output Speed138.14 tok/s | Catalog Latency1.35 s | LLMBoard score2.48 | Official input / 1M$0.10 | Max Input1M | Updated |
| ModelMA | ProviderMistral AI | Output Speed137.10 tok/s | Catalog Latency0.23 s | LLMBoard score23.28 | Official input / 1M$0.40 | Max Input128K | Updated |
| ModelMA | ProviderMistral AI | Output Speed137.10 tok/s | Catalog Latency0.23 s | LLMBoard score16.11 | Official input / 1M$0.10 | Max Input128K | Updated |
| ModelMA | ProviderMistral AI | Output Speed137.10 tok/s | Catalog Latency0.23 s | LLMBoard score6.55 | Official input / 1MN/A | Max Input128K | Updated |
Lowest catalog latency in the current records: Min istral 3 (3B Reasoning 2512) via Mistral AI at 0.16s.
Each point is one canonical model version offered by one provider in the runtime leaderboard. Upper-left is generally preferable: higher token generation speed with lower catalog latency.
This provider ranking uses observed character throughput and time to first token. Character throughput is not converted into model token throughput or benchmark capability.
Provider | P50 TTFT | P95 TTFT | P50 Throughput | P5 Throughput | 7d Calls | Success Rate | Error Rate | Updated |
|---|
| ProviderxAI | P50 TTFT1,096.00 ms | P95 TTFT1,704.80 ms | P50 Throughput81.65 char/s | P5 Throughput18.04 char/s | 7d Calls56 | Success Rate96.40% | Error Rate3.60% | Updated |
| ProviderFriendliAI | P50 TTFT1,163.50 ms | P95 TTFT14,865.50 ms | P50 Throughput263.98 char/s | P5 Throughput40.68 char/s | 7d Calls36 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderMistral AI | P50 TTFT1,350.00 ms | P95 TTFT1,968.60 ms | P50 Throughput131.43 char/s | P5 Throughput59.14 char/s | 7d Calls13 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderDeepSeek | P50 TTFT1,364.00 ms | P95 TTFT1,998.50 ms | P50 Throughput183.71 char/s | P5 Throughput45.48 char/s | 7d Calls14 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderStepFun | P50 TTFT1,418.00 ms | P95 TTFT1,895.60 ms | P50 Throughput220.72 char/s | P5 Throughput21.16 char/s | 7d Calls139 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderInception | P50 TTFT1,598.00 ms | P95 TTFT1,644.80 ms | P50 Throughput353.66 char/s | P5 Throughput346.33 char/s | 7d Calls3 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderGoogle | P50 TTFT1,734.00 ms | P95 TTFT8,806.50 ms | P50 Throughput192.02 char/s | P5 Throughput28.71 char/s | 7d Calls4,034 | Success Rate98.80% | Error Rate1.20% | Updated |
| ProviderDeepInfra | P50 TTFT1,754.00 ms | P95 TTFT12,446.25 ms | P50 Throughput65.69 char/s | P5 Throughput7.94 char/s | 7d Calls311 | Success Rate99.00% | Error Rate1.00% | Updated |
| ProviderMeta Model API | P50 TTFT1,942.50 ms | P95 TTFT7,620.10 ms | P50 Throughput33.35 char/s | P5 Throughput8.49 char/s | 7d Calls548 | Success Rate99.60% | Error Rate0.40% | Updated |
| ProviderOpenAI | P50 TTFT1,978.00 ms | P95 TTFT10,244.20 ms | P50 Throughput162.32 char/s | P5 Throughput14.13 char/s | 7d Calls534 | Success Rate95.10% | Error Rate4.90% | Updated |
| ProviderAnthropic | P50 TTFT3,427.00 ms | P95 TTFT13,175.20 ms | P50 Throughput24.41 char/s | P5 Throughput8.04 char/s | 7d Calls558 | Success Rate97.70% | Error Rate2.30% | Updated |
| ProviderXiaomi | P50 TTFT3,765.00 ms | P95 TTFT15,612.35 ms | P50 Throughput25.00 char/s | P5 Throughput4.22 char/s | 7d Calls4 | Success Rate100.00% | Error Rate0.00% | Updated |
| ProviderZAI | P50 TTFT5,335.00 ms | P95 TTFT9,690.30 ms | P50 Throughput37.95 char/s | P5 Throughput4.22 char/s | 7d Calls63 | Success Rate95.20% | Error Rate4.80% | Updated |
| ProviderMiniMax | P50 TTFT5,837.00 ms | P95 TTFT6,046.60 ms | P50 Throughput8.92 char/s | P5 Throughput6.61 char/s | 7d Calls9 | Success Rate55.60% | Error Rate44.40% | Updated |
| ProviderAzure | P50 TTFTN/A | P95 TTFTN/A | P50 ThroughputN/A | P5 ThroughputN/A | 7d Calls0 | Success RateN/A | Error RateN/A | Updated |
| ProviderMoonshot AI | P50 TTFTN/A | P95 TTFTN/A | P50 Throughput2.93 char/s | P5 Throughput1.24 char/s | 7d Calls6 | Success Rate100.00% | Error Rate0.00% | Updated |
The five fastest records in the current runtime leaderboard, with catalog latency, official token price and independent benchmark context shown where available.
Runtime leaderboard metrics answer deployment questions and should not be read as model capability scores or a quality benchmark ranking.
Runtime source: LLM Stats / ZeroEval snapshot. Speed follows the public output-tokens-per-second convention. Read the source definition.
Common questions about model speed, latency and provider runtime.
Llama 3.3 70B Instruct currently records the fastest output speed at 2,220.00 tok/s via Cerebras.
Min istral 3 (3B Reasoning 2512) has the lowest catalog latency at 0.16s via Mistral AI.
The current fastest records are Llama 3.3 70B Instruct (2,220.00 tok/s via Cerebras, 0.65s catalog latency), Llama 3.1 8B Instruct (2,047.00 tok/s via Cerebras, 0.20s catalog latency), and Min istral 3 (3B Reasoning 2512) (1,222.00 tok/s via Mistral AI, 0.16s catalog latency).
No. Output speed and latency describe runtime behavior for a provider record. They do not change the LLMBoard capability score or benchmark ranking.
Inception leads the available 7-day observed throughput records. Provider observations use characters per second and should not be read as model tokens per second.
No. Official input and output prices are shown as separate context fields; they do not affect the runtime leaderboard order.
Ranking basisThis speed and latency AI model leaderboard uses descending provider catalog output speed in generated tokens per second. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.
Selection summary
Llama 3.3 70B Instruct has the fastest current catalog record at 2,220.00 tok/s. The best AI model for runtime-sensitive work also depends on latency, provider reliability, price and workload-specific quality.
Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.