Model comparison

Grok 4.3 vs Mistral Large 4

Grok 4.3 and Mistral Large 4 score almost the same on the Noometry Index (43.8 vs 43.1), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Mistral Large 4 Mistral AI

43.1

Rank #99 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and Mistral Large 4 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 36.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 55.2% for Grok 4.3 and 27.4% for Mistral Large 4.
  • Mistral Large 4 is cheaper at $0.68 / $2.09 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Mistral Large 4 accepts more context: 1.05M tokens versus 1M.

Side by side

Grok 4.3 and Mistral Large 4 specifications
Grok 4.3Mistral Large 4
ProviderxAIMistral AI
Noometry Index43.843.1
Released2026-04-172026-10-06
WeightsProprietaryProprietary
Context window1M1.05M
Max output30K262K
Input $ / M tokens$1.25$0.68
Output $ / M tokens$2.50$2.09
Results tracked4015

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large 4 leads

Grok 4.3: 41.6 (#121), Mistral Large 4: 48.6 (#57)

Coding benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena WebDev13571541
LMArena Coding14151475
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Mistral Large 4: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Mistral Large 4
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Mistral Large 4: 22.5 (#192)

Reasoning benchmarks
BenchmarkGrok 4.3Mistral Large 4
NYT Connections (extended)55.2%27.4%
LMArena Hard Prompts13961444
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Mistral Large 4: 40.4 (#91)

Math benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Math13881488
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Mistral Large 4: 36.6 (#166)

Knowledge benchmarks
BenchmarkGrok 4.3Mistral Large 4
SimpleQA Verified33.2%20%
LMArena Expert13851447
GPQA Diamond88.8%—

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Mistral Large 4: —

Multimodal benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Mistral Large 4 leads

Grok 4.3: 50.5 (#120), Mistral Large 4: 52.6 (#82)

Multilingual benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Non-English13851415
LMArena Chinese14221491
LMArena Russian13991414
LMArena French1412—
LMArena German1395—
LMArena Japanese1379—
LMArena Korean1356—
LMArena Spanish1398—

Instruction Following Mistral Large 4 leads

Grok 4.3: 72.1 (#140), Mistral Large 4: 75.0 (#76)

Instruction Following benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Instruction Following13661424

Long Context Mistral Large 4 leads

Grok 4.3: 42.5 (#123), Mistral Large 4: 43.6 (#89)

Long Context benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Longer Query13931429

Writing & Preference Mistral Large 4 leads

Grok 4.3: 58.5 (#118), Mistral Large 4: 60.4 (#97)

Writing & Preference benchmarks
BenchmarkGrok 4.3Mistral Large 4
LMArena Text13971427
LMArena Creative Writing13801361
LMArena Multi-Turn14061424
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Mistral Large 4?

Grok 4.3 and Mistral Large 4 score almost the same on the Noometry Index (43.8 vs 43.1), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or Mistral Large 4?

Mistral Large 4 is cheaper. It lists at $0.68 per million input tokens and $2.09 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Grok 4.3 or Mistral Large 4 better for coding?

Mistral Large 4 scores higher on coding benchmarks: 48.6 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Mistral Large 4 does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.3 and Mistral Large 4 share?

15 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Mistral Large 4 has 15.

Related comparisons

Go deeper