llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Model catalog

Speech-to-Text Model Leaderboard

This speech-to-text leaderboard compares the available benchmark ranking, score status, per-minute pricing and streaming capabilities without combining unlike evaluations.

Data as of 2026-09-08

On this page

  • Catalog
  • Overview
  • Pricing
  • Modalities
  • FAQ

Speech-to-Text Leaderboard Models

Browse the speech-to-text leaderboard ranking alongside the full catalog. Models without enough comparable benchmark evidence remain unscored.

30 of 125 rows
Columns

Show columns

Sort by
Model
LLMBoard score
Input -> output
Native price
Provider
Context
Released
ModelAUARK-ASR-3BAudio8LLMBoard score70.77Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelOTMOSS-Transcribe-preview-2BOpenmoss TeamLLMBoard score70.07Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelOTMOSS-Transcribe-DiarizeOpenmoss TeamLLMBoard score69.36Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelCOcohere-transcribe-03-2026CoherelabsLLMBoard score68.66Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelAUARK-ASR-0.6BAudio8LLMBoard score67.96Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelSOZipformer-cr-ctc-transducer-XL-290MSoundsgoodaiLLMBoard score67.25Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelACQwen3-ASR-1.7BAlibaba Cloud / Qwen TeamLLMBoard score66.55Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelMIPhi-4-multimodal-instructMicrosoftLLMBoard score65.84Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVcanary-1b-flashNVIDIALLMBoard score64.43Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelKYstt-2.6b-enKyutaiLLMBoard score63.73Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelACQwen3-ASR-0.6BAlibaba Cloud / Qwen TeamLLMBoard score63.02Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVcanary-1bNVIDIALLMBoard score62.32Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelMAmoonshine-streaming-mediumMoonshine AiLLMBoard score61.61Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelSOZipformer-transducer-XL-290MSoundsgoodaiLLMBoard score60.91Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVparakeet-tdt-1.1bNVIDIALLMBoard score60.21Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelAUAudio8-ASR-0.1BAudio8LLMBoard score59.15Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelZAGLM-ASR-Nano-2512Zhipu AILLMBoard score59.15Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelMIVoxtral-Mini-3B-2507MistralaiLLMBoard score58.10Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVcanary-180m-flashNVIDIALLMBoard score57.39Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVcanary-1b-v2NVIDIALLMBoard score55.99Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelDWdistil-large-v3.5Distil WhisperLLMBoard score55.28Input -> outputunspecified -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelAMAmazon TranscribeAmazonLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelNVCanary Qwen 2.5BNVIDIALLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelGOChirp 3GoogleLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelCOCohere TranscribeCohereLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelALFun-Realtime-ASR-previewAlibabaLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelGOGemini 2.0 Flash LiteGoogleLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelGOGemini 2.5 ProGoogleLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelGOGemini 3 Flash (High)GoogleLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A
ModelGOGemini 3.1 Flash-Lite Preview (Minimal)GoogleLLMBoard score55.09Input -> outputaudio -> textNative priceN/AProviderN/AContextN/AReleasedN/A

This leaderboard uses evaluation and benchmark evidence only for scores. Price, context and model specifications remain separate from the ranking.

Leaderboard overview

Models125
Vendors39
Scored models114
Ranking status114 provisional
Models12539 vendors
Multimodal inputs0Products accepting more than one input modality

Vendor catalog mix

125 models
NVIDIA15
Google14
Facebook12
OpenAI12
AssemblyAI6
Microsoft6
Moonshine Ai5
Alibaba4

Speech-to-Text Leaderboard Pricing

Use price as context for the leaderboard, not as part of its benchmark ranking. Each listing keeps its original unit.

per minute

Up to $0.0018 / minute
2
$0.0018 / minute-$0.009 / minute
2
Above $0.009 / minute
1

Speech-to-Text Input and Output Modalities

Compare input and output support alongside the leaderboard; modality support does not change the benchmark ranking.

ModelIn: textIn: imageIn: audioIn: videoOut: textOut: imageOut: audioOut: video
Amazon TranscribeAmazon--Yes-Yes---
AARK-ASR-0.6BAudio8----Yes---
AARK-ASR-3BAudio8----Yes---
Sasr-conformer-loquaciousSpeechbrain----Yes---
Sasr-wav2vec2-librispeechSpeechbrain----Yes---
AAudio8-ASR-0.1BAudio8----Yes---
BestAssemblyAI--Yes-Yes---
Canary Qwen 2.5BNVIDIA--Yes-Yes---
canary-180m-flashNVIDIA----Yes---
canary-1bNVIDIA----Yes---
canary-1b-flashNVIDIA----Yes---
canary-1b-v2NVIDIA----Yes---

About Speech-to-Text Model Leaderboard

How does this speech-to-text leaderboard work?

114 models currently have a domain-specific LLMBoard score. The leaderboard ranking uses matched benchmark evidence, while coverage and provisional status identify incomplete evidence.

Which speech-to-text model has the lowest listed price?

Whisper Large V3 Turbo has the lowest current native price at $0.0007 / minute via Groq.

Which speech-to-text model has the highest LLMBoard score?

ARK-ASR-3B leads the available LLMBoard scores at 70.77.

How many speech-to-text models support multiple input modalities?

0 of 125 catalog models have more than one listed input modality.

How many speech-to-text models have no native price?

120 catalog models currently have no positive native-unit price listing.

How are prices displayed?

Price availability

With native price
5
Without native price
120

Prices retain the unit returned by the catalog, such as per image, per second, per minute or per million characters. They are never shown as token prices.

Why can a model have no price?

Some models do not have a current price listing in their native billing unit.

Lowest listed native priceWhisper Large V3 Turbo$0.0007 / minute via Groq