llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Meta model product

Muse Spark

Muse Spark is a natively multimodal reasoning model developed by Meta Superintelligence Labs.

Updated Sep 8, 2026. Default version: Muse Spark

LLMBoard Score72.1Muse Spark
Coverage60%18 benchmark families
Context windowN/ATokens
Official input priceN/AOfficial price unavailable

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Muse Spark Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Muse Spark LLMBoard score breakdown

Muse Spark Benchmark Results

Benchmark scores for Muse Spark.

26 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkFrontierScience ResearchScore38.30%Rank01Participants4Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkHealthBench HardScore42.80%Rank01Participants9Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkIPhO 2025Score82.60%Rank01Participants2Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkMedXpertQAScore78.40%Rank01Participants12Percentile100.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkZEROBenchScore0.33 pointsRank03Participants10Percentile77.78%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveCodeBench ProScore0.80 pointsRank04Participants4Percentile0.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkSimpleVQAScore0.71 pointsRank04Participants14Percentile76.92%EvidenceCEvaluatedSep 8, 2026
BenchmarkScreenSpot ProScore84.10%Rank05Participants26Percentile84.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Vision Style ControlScore1,293.73 ratingRank05Participants107Percentile96.23%EvidenceAEvaluatedAug 27, 2026
BenchmarkLM Arena VisionScore1,306.23 ratingRank07Participants107Percentile94.34%EvidenceAEvaluatedAug 27, 2026
BenchmarkHumanity's Last ExamScore58.40%Rank08Participants103Percentile93.14%EvidenceCEvaluatedSep 8, 2026
BenchmarkDeepSearchQAScore74.80%Rank09Participants10Percentile11.11%EvidenceCEvaluatedSep 8, 2026
BenchmarkARC-AGI v2Score42.50%Rank10Participants18Percentile47.06%EvidenceCEvaluatedSep 8, 2026
BenchmarkERQAScore64.70%Rank11Participants26Percentile60.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Text Style ControlScore1,488.17 ratingRank11Participants210Percentile95.22%EvidenceAEvaluatedSep 2, 2026
BenchmarkMMMU-ProScore80.40%Rank13Participants69Percentile82.35%EvidenceCEvaluatedSep 8, 2026
BenchmarkCharXiv-RScore86.40%Rank14Participants55Percentile75.93%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Document Style ControlScore1,466.33 ratingRank14Participants33Percentile59.38%EvidenceAEvaluatedJul 30, 2026
BenchmarkTau2 TelecomScore91.50%Rank16Participants36Percentile57.14%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,473.42 ratingRank18Participants210Percentile91.87%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena DocumentScore1,443.25 ratingRank20Participants33Percentile40.63%EvidenceAEvaluatedJul 30, 2026
BenchmarkTerminal-Bench 2.0Score59.00%Rank24Participants51Percentile54.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena Text FactualityScore1,466.50 ratingRank24Participants121Percentile80.83%EvidenceAEvaluatedSep 2, 2026
BenchmarkSWE-Bench VerifiedScore77.40%Rank27Participants113Percentile76.79%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore89.50%Rank31Participants247Percentile87.80%EvidenceCEvaluatedSep 8, 2026
BenchmarkSWE-Bench ProScore52.40%Rank46Participants55Percentile16.67%EvidenceCEvaluatedSep 8, 2026

Muse Spark Arena Results

Preference and agent-evaluation results for the default version.

30 of 90 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
Arenatext style controlCategoryfrenchRank02Rating / score1,524.77Votes468ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryhumorRank02Rating / score1,346.00Votes227ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryhumorRank02Rating / score1,325.94Votes227ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategorygermanRank03Rating / score1,508.70Votes235ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorykoreanRank04Rating / score1,473.97Votes248ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategoryenglishRank05Rating / score1,288.53Votes2,329ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryoverallRank05Rating / score1,293.73Votes5,585ObservationsN/AResult dateAug 27, 2026
ArenatextCategoryfrenchRank06Rating / score1,501.62Votes468ObservationsN/AResult dateSep 2, 2026
ArenatextCategorygermanRank06Rating / score1,495.01Votes235ObservationsN/AResult dateSep 2, 2026
ArenavisionCategorycreative writing visionRank06Rating / score1,327.48Votes334ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategorycreative writing visionRank06Rating / score1,303.49Votes334ObservationsN/AResult dateAug 27, 2026
Arenatext style controlCategoryindustry legal and governmentRank07Rating / score1,503.74Votes974ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryenglishRank07Rating / score1,302.94Votes2,329ObservationsN/AResult dateAug 27, 2026
ArenavisionCategoryoverallRank07Rating / score1,306.23Votes5,585ObservationsN/AResult dateAug 27, 2026
Arenavision style controlCategoryocrRank08Rating / score1,301.78Votes4,021ObservationsN/AResult dateAug 27, 2026
ArenatextCategorykoreanRank09Rating / score1,459.84Votes248ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryenglishRank09Rating / score1,495.50Votes6,471ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry business and management and financial operationsRank09Rating / score1,490.07Votes2,747ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry medicine and healthcareRank10Rating / score1,503.00Votes955ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryspanishRank10Rating / score1,481.55Votes501ObservationsN/AResult dateSep 2, 2026
Arenavision style controlCategorychineseRank10Rating / score1,342.05Votes450ObservationsN/AResult dateAug 27, 2026
ArenatextCategoryindustry legal and governmentRank11Rating / score1,489.12Votes974ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryoverallRank11Rating / score1,488.17Votes13,572ObservationsN/AResult dateSep 2, 2026
ArenavisionCategoryocrRank11Rating / score1,306.97Votes4,021ObservationsN/AResult dateAug 27, 2026
Arenatext factualityCategorycreative writingRank12Rating / score1,463.87Votes1,571ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryexclude tiesRank12Rating / score1,500.93Votes10,291ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard prompts englishRank12Rating / score1,507.61Votes4,353ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry entertainment and sports and mediaRank12Rating / score1,462.19Votes2,510ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry software and it servicesRank12Rating / score1,519.97Votes5,318ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorynon englishRank12Rating / score1,476.97Votes7,101ObservationsN/AResult dateSep 2, 2026

Muse Spark Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
N/A
Official provider
N/A
Lowest third-party
N/A
Tracked offerings
0
No provider prices

The default version has no current input or output token prices.

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Muse Spark Runtime Performance

Provider-specific output speed and catalog latency for Muse Spark. Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Muse Spark Specifications

Technical details for the model's default version.

Version
Muse Spark
Released
Apr 8, 2026
Knowledge cutoff
Unknown
Parameters
N/A
Context window
N/A
Max output
N/A
Inputs
image, text
Outputs
text
Open weights
No
License
Proprietary

Muse Spark Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionMuse SparkReleasedApr 8, 2026LLMBoard72.09ParametersN/AContextN/AMax outputN/AOpen weightsNoLicenseProprietary

Models similar to Muse Spark

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#24+6.89
ME

Muse Spark 1.2

Meta

78.98 LLMBoard

Details
#14+13.08
ME

Muse Spark 1.1

Meta

85.17 LLMBoard

Details
#79-14.70
ME

Muse Glimmer 30B

Meta

57.39 LLMBoard

Details
#6+18.73
ME

Muse Spark 1.3

Meta

90.82 LLMBoard

Details
#200-48.65
ME

Llama 4 Maverick

Meta

23.44 LLMBoard

Details
#203-49.21
ME

Llama 3.1 405B

Meta

22.88 LLMBoard

Details

What is Muse Spark?

Key information about Muse Spark and its available data.

Muse Spark is the first model in the Muse family, with support for tool-use, visual chain of thought, and multi-agent orchestration. Its Contemplating mode orchestrates multiple agents reasoning in parallel and achieved 58% on Humanity's Last Exam and 38% on FrontierScience Research.

Data as of 2026-09-08.

FAQ

Common questions about Muse Spark.

When was Muse Spark released?

Muse Spark's default version was released on Apr 8, 2026.

How much does Muse Spark cost?

No official standard PAYG price is currently available for Muse Spark.

Who created Muse Spark?

Muse Spark was created by Meta.

What is the context window for Muse Spark?

A context window is not available for the default version.

Is Muse Spark open weight?

No. The default version is not marked as having publicly available weights.

How many API providers offer Muse Spark?

No provider offering is currently linked to the default version.

Browse runtime rankings