llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

DeepSeek model product

DeepSeek-V2.5

5 is an upgraded large language model that combines general-purpose and coding capabilities.

Updated Sep 8, 2026. Default version: DeepSeek-V2.5

LLMBoard Score13.7DeepSeek-V2.5
Coverage20%11 benchmark families
Context window8.2KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

DeepSeek-V2.5 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

DeepSeek-V2.5 LLMBoard score breakdown

DeepSeek-V2.5 Benchmark Results

Benchmark scores for DeepSeek-V2.5.

17 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkAiderScore72.20%Rank01Participants4Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkDS-Arena-CodeScore63.10%Rank01Participants1Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkDS-FIM-EvalScore78.30%Rank01Participants1Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveCodeBench(01-09)Score41.80%Rank01Participants1Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkAlignBenchScore80.40%Rank02Participants4Percentile66.67%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanEval-MulScore73.80%Rank02Participants2Percentile0.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkAlpacaEval 2.0Score50.50%Rank03Participants4Percentile33.33%EvidenceCEvaluatedSep 8, 2026
BenchmarkMT-BenchScore90.20%Rank03Participants12Percentile81.82%EvidenceCEvaluatedSep 8, 2026
BenchmarkBBHScore84.30%Rank05Participants12Percentile63.64%EvidenceCEvaluatedSep 8, 2026
BenchmarkArena HardScore76.20%Rank08Participants26Percentile72.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkGSM8kScore95.10%Rank11Participants48Percentile78.72%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanEvalScore89.00%Rank16Participants66Percentile76.92%EvidenceCEvaluatedSep 8, 2026
BenchmarkMATHScore74.70%Rank29Participants71Percentile60.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMLUScore80.40%Rank60Participants101Percentile41.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench VerifiedScore16.80%Rank112Participants113Percentile0.89%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,293.70 ratingRank164Participants210Percentile22.01%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena Text Style ControlScore1,323.21 ratingRank166Participants210Percentile21.05%EvidenceAEvaluatedSep 2, 2026

DeepSeek-V2.5 Arena Results

Preference and agent-evaluation results for the default version.

30 of 56 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenatextCategoryjapaneseRank135Rating / score1,227.86Votes161ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank139Rating / score1,252.10Votes161ObservationsN/AResult dateSep 2, 2026
ArenatextCategorykoreanRank140Rating / score1,209.14Votes344ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorykoreanRank144Rating / score1,254.58Votes344ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryfrenchRank146Rating / score1,288.77Votes222ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycreative writingRank146Rating / score1,309.31Votes1,090ObservationsN/AResult dateSep 2, 2026
ArenatextCategorychineseRank150Rating / score1,317.47Votes446ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorychineseRank150Rating / score1,355.44Votes446ObservationsN/AResult dateSep 2, 2026
ArenatextCategorygermanRank151Rating / score1,257.88Votes140ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry writing and literature and languageRank153Rating / score1,314.30Votes1,930ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryfrenchRank154Rating / score1,325.39Votes222ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry entertainment and sports and mediaRank155Rating / score1,301.84Votes1,216ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry software and it servicesRank155Rating / score1,364.68Votes1,688ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymathRank156Rating / score1,288.48Votes1,031ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorylonger queryRank156Rating / score1,337.19Votes1,037ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryrussianRank156Rating / score1,322.55Votes1,034ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryrussianRank157Rating / score1,288.52Votes1,034ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycodingRank157Rating / score1,374.30Votes1,079ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorygermanRank157Rating / score1,289.50Votes140ObservationsN/AResult dateSep 2, 2026
ArenatextCategorycreative writingRank158Rating / score1,284.91Votes1,090ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryexpertRank158Rating / score1,266.24Votes441ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry entertainment and sports and mediaRank158Rating / score1,274.96Votes1,216ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry software and it servicesRank158Rating / score1,313.53Votes1,688ObservationsN/AResult dateSep 2, 2026
ArenatextCategorylonger queryRank158Rating / score1,300.59Votes1,037ObservationsN/AResult dateSep 2, 2026
ArenatextCategorycodingRank159Rating / score1,309.73Votes1,079ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard prompts englishRank160Rating / score1,299.04Votes1,132ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryinstruction followingRank160Rating / score1,279.73Votes2,970ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard promptsRank161Rating / score1,289.13Votes1,790ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry medicine and healthcareRank161Rating / score1,291.29Votes291ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymulti turnRank161Rating / score1,296.82Votes1,178ObservationsN/AResult dateSep 2, 2026

DeepSeek-V2.5 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No provider prices

The default version has no current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

DeepSeek-V2.5 Runtime Performance

Provider-specific output speed and catalog latency for DeepSeek-V2.5. Runtime does not affect the capability score.

2 rows
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderDeepSeekOutput Speed100.00 tok/sCatalog Latency0.50 sMax Input8.2KMax Output8.2KUpdatedSep 8, 2026
ProviderDeepInfraOutput Speed63.00 tok/sCatalog Latency0.50 sMax Input8.2KMax Output8.2KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

DeepSeek-V2.5 Specifications

Technical details for the model's default version.

Version
DeepSeek-V2.5
Released
May 8, 2024
Knowledge cutoff
Unknown
Parameters
236B
Context window
8.2K
Max output
8.2K
Inputs
text
Outputs
text
Open weights
Yes
License
deepseek

DeepSeek-V2.5 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionDeepSeek-V2.5ReleasedMay 8, 2024LLMBoard13.68Parameters236BContext8.2KMax output8.2KOpen weightsYesLicensedeepseek

Models similar to DeepSeek-V2.5

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#229+0.70
DE

DeepSeek-R1-Distill-Qwen

DeepSeek

14.38 LLMBoard

Details
#240-2.61
DE

DeepSeek-R1-Distill-Llama

DeepSeek

11.07 LLMBoard

Details
#194+11.30
DE

DeepSeek-V3

DeepSeek

24.98 LLMBoard

Details
#294-13.68
DE

DeepSeek-VL2

DeepSeek

0.00 LLMBoard

Details
#149+25.37
DE

DeepSeek-V3.1

DeepSeek

39.05 LLMBoard

Details
#134+28.97
DE

DeepSeek-R1

DeepSeek

42.65 LLMBoard

Details

What is DeepSeek-V2.5?

Key information about DeepSeek-V2.5 and its available data.

5 combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It is optimized for writing, instruction following, and alignment with human preferences.

Data as of 2026-09-08.

FAQ

Common questions about DeepSeek-V2.5.

When was DeepSeek-V2.5 released?

DeepSeek-V2.5's default version was released on May 8, 2024.

How much does DeepSeek-V2.5 cost?

No official standard PAYG price is currently available for DeepSeek-V2.5.

Who created DeepSeek-V2.5?

DeepSeek-V2.5 was created by DeepSeek.

What is the context window for DeepSeek-V2.5?

The default version has a 8.2K token context window.

Is DeepSeek-V2.5 open weight?

Yes. The default version is marked as open weight under deepseek.

How many API providers offer DeepSeek-V2.5?

No provider offering is currently linked to the default version.