llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Meituan model product

LongCat Flash Chat

LongCat Flash Chat is Meituan’s open-source 560B-parameter Mixture-of-Experts (MoE) model for conversational and agentic tasks.

Updated Sep 8, 2026. Default version: LongCat-Flash-Chat

LLMBoard Score35.9LongCat-Flash-Chat
Coverage60%16 benchmark families
Context window128KTokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

LongCat Flash Chat Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

LongCat-Flash-Chat LLMBoard score breakdown

LongCat Flash Chat Benchmark Results

Benchmark scores for LongCat-Flash-Chat.

19 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkCMMLUScore84.34%Rank03Participants6Percentile60.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkZebraLogicScore89.30%Rank04Participants8Percentile57.14%EvidenceCEvaluatedSep 8, 2026
BenchmarkTerminal-BenchScore39.51%Rank09Participants25Percentile66.67%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMLUScore89.71%Rank12Participants101Percentile89.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkMATH-500Score96.40%Rank13Participants32Percentile61.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 AirlineScore58.00%Rank13Participants24Percentile47.83%EvidenceCEvaluatedSep 8, 2026
BenchmarkDROPScore79.06%Rank16Participants30Percentile48.28%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanEvalScore88.41%Rank19Participants66Percentile72.31%EvidenceCEvaluatedSep 8, 2026
BenchmarkIFEvalScore89.65%Rank20Participants68Percentile71.64%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 RetailScore71.27%Rank20Participants27Percentile26.92%EvidenceCEvaluatedSep 8, 2026
BenchmarkTau2 TelecomScore73.68%Rank24Participants36Percentile34.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMLU-ProScore82.68%Rank37Participants138Percentile73.72%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveCodeBenchScore48.02%Rank50Participants75Percentile33.78%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Text Style ControlScore1,435.76 ratingRank69Participants210Percentile67.46%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena TextScore1,426.03 ratingRank71Participants210Percentile66.51%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena Text FactualityScore1,427.99 ratingRank82Participants121Percentile32.50%EvidenceAEvaluatedSep 2, 2026
BenchmarkSWE-Bench VerifiedScore60.40%Rank84Participants113Percentile25.89%EvidenceCEvaluatedSep 8, 2026
BenchmarkAIME 2025Score61.25%Rank101Participants119Percentile15.25%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore73.23%Rank127Participants247Percentile48.78%EvidenceCEvaluatedSep 8, 2026

LongCat Flash Chat Arena Results

Preference and agent-evaluation results for the default version.

30 of 81 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenatextCategoryindustry medicine and healthcareRank36Rating / score1,459.93Votes666ObservationsN/AResult dateSep 2, 2026
ArenatextCategorycodingRank39Rating / score1,473.04Votes7,755ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategorycodingRank39Rating / score1,501.81Votes6,683ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard prompts englishRank42Rating / score1,459.86Votes8,709ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycodingRank43Rating / score1,502.69Votes7,755ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry software and it servicesRank44Rating / score1,465.74Votes10,911ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryspanishRank44Rating / score1,445.19Votes401ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryenglishRank46Rating / score1,451.28Votes13,159ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategorychineseRank46Rating / score1,480.57Votes1,381ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryfrenchRank47Rating / score1,458.33Votes182ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryhard prompts englishRank49Rating / score1,477.69Votes7,661ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry software and it servicesRank50Rating / score1,488.59Votes10,911ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymathRank53Rating / score1,440.81Votes686ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryexpertRank53Rating / score1,465.52Votes1,978ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryindustry mathematicalRank53Rating / score1,435.48Votes1,219ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryindustry software and it servicesRank53Rating / score1,481.30Votes10,236ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryexpertRank54Rating / score1,451.78Votes2,605ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry mathematicalRank55Rating / score1,443.20Votes542ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorychineseRank55Rating / score1,482.79Votes1,719ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryenglishRank55Rating / score1,458.76Votes13,159ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard prompts englishRank55Rating / score1,475.53Votes8,709ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategorymathRank57Rating / score1,427.87Votes1,391ObservationsN/AResult dateSep 2, 2026
Arenatext factualityCategoryenglishRank58Rating / score1,452.90Votes13,063ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry business and management and financial operationsRank59Rating / score1,428.25Votes2,149ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorykoreanRank62Rating / score1,384.62Votes470ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard promptsRank63Rating / score1,460.10Votes17,806ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank63Rating / score1,388.15Votes252ObservationsN/AResult dateSep 2, 2026
ArenatextCategorychineseRank64Rating / score1,466.11Votes1,719ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard promptsRank66Rating / score1,439.25Votes17,806ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryexpertRank66Rating / score1,465.17Votes2,605ObservationsN/AResult dateSep 2, 2026

LongCat Flash Chat Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No provider prices

The default version has no current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

LongCat Flash Chat Runtime Performance

Provider-specific output speed and catalog latency for LongCat-Flash-Chat. Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

LongCat Flash Chat Specifications

Technical details for the model's default version.

Version
LongCat-Flash-Chat
Released
Aug 29, 2025
Knowledge cutoff
Unknown
Parameters
560B
Context window
128K
Max output
128K
Inputs
text
Outputs
text
Open weights
Yes
License
MIT

LongCat Flash Chat Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionLongCat-Flash-ChatReleasedAug 29, 2025LLMBoard35.85Parameters560BContext128KMax output128KOpen weightsYesLicenseMIT

Models similar to LongCat Flash Chat

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#175-3.98
ME

LongCat Flash Lite

Meituan

31.87 LLMBoard

Details
#72+23.64
ME

LongCat Flash Thinking

Meituan

59.49 LLMBoard

Details
#1600.00
MI

MiniMax M1 80K

MiniMax

35.85 LLMBoard

Details
#158+0.40
MA

Kimi K2

Moonshot AI

36.25 LLMBoard

Details
#161-0.45
XA

Grok Code Fast 1

xAI

35.40 LLMBoard

Details
#162-0.59
OP

o1

OpenAI

35.26 LLMBoard

Details

What is LongCat Flash Chat?

Key information about LongCat Flash Chat and its available data.

3B parameters based on contextual demands. It supports a 128K context and is optimized for reasoning, coding, instruction following, tool use, and complex multi-step interactions.

Data as of 2026-09-08.

FAQ

Common questions about LongCat Flash Chat.

When was LongCat Flash Chat released?

LongCat Flash Chat's default version was released on Aug 29, 2025.

How much does LongCat Flash Chat cost?

No official standard PAYG price is currently available for LongCat Flash Chat.

Who created LongCat Flash Chat?

LongCat Flash Chat was created by Meituan.

What is the context window for LongCat Flash Chat?

The default version has a 128K token context window.

Is LongCat Flash Chat open weight?

Yes. The default version is marked as open weight under MIT.

How many API providers offer LongCat Flash Chat?

No provider offering is currently linked to the default version.

Browse runtime rankings