llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model Directory

Scoring & Data

Scoring & Data
1224 models729 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll Benchmarks
llmboard.aiCopyright 2026 llmboard.ai

Microsoft model product

Phi 4

Phi 4 is an open large language model from Microsoft for reasoning, coding, and knowledge tasks.

Updated Sep 8, 2026. Default version: Phi 4

LLMBoard Score7.5Phi 4
Coverage40%13 benchmark families
Context window16KTokens
Official input price$0.125Azure API

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Similar models
  • About
  • FAQ

Phi 4 Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Phi 4 LLMBoard score breakdown

Phi 4 Benchmark Results

Benchmark scores for Phi 4.

15 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkPhiBenchScore56.20%Rank03Participants3Percentile0.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanEval+Score82.80%Rank05Participants10Percentile55.56%EvidenceCEvaluatedSep 8, 2026
BenchmarkArena HardScore75.40%Rank09Participants26Percentile68.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkDROPScore75.50%Rank19Participants30Percentile37.93%EvidenceCEvaluatedSep 8, 2026
BenchmarkMATHScore80.40%Rank19Participants71Percentile74.29%EvidenceCEvaluatedSep 8, 2026
BenchmarkMGSMScore80.60%Rank19Participants31Percentile40.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkLiveBenchScore47.60%Rank36Participants38Percentile5.41%EvidenceCEvaluatedSep 8, 2026
BenchmarkHumanEvalScore82.60%Rank41Participants66Percentile38.46%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMLUScore84.80%Rank45Participants101Percentile56.00%EvidenceCEvaluatedSep 8, 2026
BenchmarkSimpleQAScore3.00%Rank45Participants47Percentile4.35%EvidenceCEvaluatedSep 8, 2026
BenchmarkIFEvalScore63.00%Rank66Participants68Percentile2.99%EvidenceCEvaluatedSep 8, 2026
BenchmarkMMLU-ProScore70.40%Rank89Participants138Percentile35.77%EvidenceCEvaluatedSep 8, 2026
BenchmarkGPQAScore56.10%Rank177Participants247Percentile28.46%EvidenceCEvaluatedSep 8, 2026
BenchmarkLM Arena TextScore1,216.47 ratingRank196Participants210Percentile6.70%EvidenceAEvaluatedSep 2, 2026
BenchmarkLM Arena Text Style ControlScore1,255.98 ratingRank199Participants210Percentile5.26%EvidenceAEvaluatedSep 2, 2026

Phi 4 Arena Results

Preference and agent-evaluation results for the default version.

30 of 56 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
ArenatextCategoryjapaneseRank153Rating / score1,158.40Votes500ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryjapaneseRank158Rating / score1,194.11Votes500ObservationsN/AResult dateSep 2, 2026
ArenatextCategorykoreanRank160Rating / score1,150.74Votes319ObservationsN/AResult dateSep 2, 2026
ArenatextCategorygermanRank162Rating / score1,222.17Votes523ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorykoreanRank164Rating / score1,195.86Votes319ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryspanishRank165Rating / score1,233.53Votes127ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorygermanRank167Rating / score1,255.68Votes523ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryspanishRank167Rating / score1,269.28Votes127ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryfrenchRank168Rating / score1,223.96Votes205ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryfrenchRank170Rating / score1,265.74Votes205ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry mathematicalRank176Rating / score1,253.84Votes2,350ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry mathematicalRank177Rating / score1,278.30Votes2,350ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymathRank178Rating / score1,245.70Votes2,764ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorymathRank179Rating / score1,264.71Votes2,764ObservationsN/AResult dateSep 2, 2026
ArenatextCategorychineseRank186Rating / score1,211.42Votes1,503ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryexpertRank187Rating / score1,202.80Votes1,124ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryexpertRank187Rating / score1,267.87Votes1,124ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorychineseRank188Rating / score1,267.21Votes1,503ObservationsN/AResult dateSep 2, 2026
ArenatextCategorycodingRank189Rating / score1,231.63Votes3,305ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard prompts englishRank189Rating / score1,289.12Votes3,804ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard promptsRank190Rating / score1,219.45Votes5,747ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryhard prompts englishRank190Rating / score1,230.01Votes3,804ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry legal and governmentRank190Rating / score1,241.04Votes1,341ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategorycodingRank190Rating / score1,306.18Votes3,305ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryindustry legal and governmentRank191Rating / score1,298.77Votes1,341ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry software and it servicesRank192Rating / score1,228.13Votes5,580ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryinstruction followingRank192Rating / score1,201.32Votes9,162ObservationsN/AResult dateSep 2, 2026
Arenatext style controlCategoryhard promptsRank192Rating / score1,277.49Votes5,747ObservationsN/AResult dateSep 2, 2026
ArenatextCategoryindustry medicine and healthcareRank193Rating / score1,201.02Votes1,037ObservationsN/AResult dateSep 2, 2026
ArenatextCategorymulti turnRank193Rating / score1,206.07Votes3,517ObservationsN/AResult dateSep 2, 2026

Phi 4 Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
$0.125 input, $0.50 output per 1M
Official provider
Azure
Lowest third-party
From $0.07 input, $0.14 output per 1M via OpenRouter
Tracked offerings
3
3 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderOpenRouterProvider model IDmicrosoft/phi-4RegionglobalInput / 1M$0.07Output / 1M$0.14Context16.4KUpdatedSep 8, 2026
ProviderAzure Cognitive ServicesProvider model IDphi-4RegionglobalInput / 1M$0.125Output / 1M$0.50Context128KUpdatedSep 8, 2026
ProviderAzureProvider model IDphi-4RegionglobalInput / 1M$0.125Output / 1M$0.50Context128KUpdatedSep 8, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Phi 4 Runtime Performance

Provider-specific output speed and catalog latency for Phi 4. Runtime does not affect the capability score.

1 row
Columns

Show columns

Sort by
Provider
Output Speed
Catalog Latency
Max Input
Max Output
Updated
ProviderDeepInfraOutput Speed33.00 tok/sCatalog Latency0.20 sMax Input16KMax Output16KUpdatedSep 8, 2026

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Phi 4 Specifications

Technical details for the model's default version.

Version
Phi 4
Released
Dec 12, 2024
Knowledge cutoff
Jun 1, 2024
Parameters
14.7B
Context window
16K
Max output
16K
Inputs
text
Outputs
text
Open weights
Yes
License
MIT

Phi 4 Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 row
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionPhi 4ReleasedDec 12, 2024LLMBoard7.48Parameters14.7BContext16KMax output16KOpen weightsYesLicenseMIT

Models similar to Phi 4

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#257-2.04
MI

Phi 4 multimodal

Microsoft

5.44 LLMBoard

Details
#260-3.89
MI

Phi 3.5 MoE

Microsoft

3.59 LLMBoard

Details
#236+4.72
MI

Phi 4 Mini Reasoning

Microsoft

12.20 LLMBoard

Details
#286-7.48
MI

Phi 3.5 mini

Microsoft

0.00 LLMBoard

Details
#287-7.48
MI

Phi 3.5 vision

Microsoft

0.00 LLMBoard

Details
#288-7.48
MI

Phi 4 Mini

Microsoft

0.00 LLMBoard

Details

What is Phi 4?

Key information about Phi 4 and its available data.

Phi 4 is an open model from Microsoft built for advanced reasoning, coding, and knowledge tasks. It was developed using synthetic data, filtered web data, academic texts, and supervised fine-tuning.

Data as of 2026-09-08.

FAQ

Common questions about Phi 4.

When was Phi 4 released?

Phi 4's default version was released on Dec 12, 2024.

How much does Phi 4 cost?

Phi 4's official API price is $0.125 per million input tokens and $0.50 per million output tokens via Azure. The lowest tracked third-party offer starts at $0.07 input and $0.14 output via OpenRouter.

Who created Phi 4?

Phi 4 was created by Microsoft.

What is the context window for Phi 4?

The default version has a 16K token context window.

Is Phi 4 open weight?

Yes. The default version is marked as open weight under MIT.

How many API providers offer Phi 4?

3 provider offerings are linked to the default version.