Model comparison

GPT-4.1 vs Grok Build 0.1

GPT-4.1 and Grok Build 0.1 score almost the same on the Noometry Index (35.9 vs 36.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 11.7.
  • Grok Build 0.1 is cheaper at $1 / $2 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 256K.

Side by side

GPT-4.1 and Grok Build 0.1 specifications
GPT-4.1Grok Build 0.1
ProviderOpenAIxAI
Noometry Index35.936.4
Released2025-04-142026-04-16
WeightsProprietaryProprietary
Context window1.05M256K
Max output33K256K
Input $ / M tokens$2$1
Output $ / M tokens$8$2
Results tracked523

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

GPT-4.1: 34.4 (#238), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkGPT-4.1Grok Build 0.1
SWE-bench Verified48.5%—
SWE-bench Verified (bash only)39.6%—
Aider Polyglot52.4%—
SciCode—50.2%
WeirdML39%—
LMArena Coding1391—
CadEval42%—
ALE-Bench558.1—

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1Grok Build 0.1
Berkeley Function Calling Leaderboard54%—
GBAEval—2.4%

Reasoning Grok Build 0.1 leads

GPT-4.1: 11.7 (#339), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkGPT-4.1Grok Build 0.1
ARC-AGI-20.4%—
SimpleBench27%—
Kagi LLM Benchmark52.3%—
ARC-AGI-15.5%—
CritPt—9.1%
Chess Puzzles6%—
EnigmaEval2.2%—
LMArena Hard Prompts1384—
DTBench68.3%—
LMCA25.6%—
Epoch Capabilities Index136.78—
ForecastBench61.5—

Math Not comparable

GPT-4.1: 22.3 (#280), Grok Build 0.1: —

Math benchmarks
BenchmarkGPT-4.1Grok Build 0.1
FrontierMath (Tiers 1-3)6%—
OTIS Mock AIME 2024-202538.3%—
Omni-MATH47.1%—
LMArena Math1370—
MATH Level 583%—
FrontierMath (Feb 2025 set)5.5%—
FrontierMath Tier 4 (v1)0%—

Knowledge Not comparable

GPT-4.1: 37.1 (#160), Grok Build 0.1: —

Knowledge benchmarks
BenchmarkGPT-4.1Grok Build 0.1
GPQA Diamond66.9%—
Humanity's Last Exam5.4%—
SimpleQA Verified31.1%—
MMLU-Pro81.1%—
Vectara Hallucination Rate5.6%—
GPQA (HELM)65.9%—
LMArena Expert1364—

Multimodal Not comparable

GPT-4.1: 38.2 (#67), Grok Build 0.1: —

Multimodal benchmarks
BenchmarkGPT-4.1Grok Build 0.1
LMArena Vision1211—
GeoBench72%—

Multilingual Not comparable

GPT-4.1: 49.4 (#133), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkGPT-4.1Grok Build 0.1
LMArena Non-English1370—
LMArena Chinese1382—
LMArena French1382—
LMArena German1381—
LMArena Japanese1319—
LMArena Korean1339—
LMArena Russian1377—
LMArena Spanish1376—

Instruction Following Not comparable

GPT-4.1: 71.3 (#153), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkGPT-4.1Grok Build 0.1
IFEval83.8%—
LMArena Instruction Following1367—

Long Context Not comparable

GPT-4.1: 40.0 (#163), Grok Build 0.1: —

Long Context benchmarks
BenchmarkGPT-4.1Grok Build 0.1
Fiction.LiveBench63.9%—
LMArena Longer Query1385—

Writing & Preference Not comparable

GPT-4.1: 57.6 (#125), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkGPT-4.1Grok Build 0.1
LMArena Text1383—
LMArena Creative Writing1363—
EQ-Bench Creative Writing1420—
WildBench85.4%—
LMArena Multi-Turn1398—

Frequently asked questions

Is GPT-4.1 better than Grok Build 0.1?

GPT-4.1 and Grok Build 0.1 score almost the same on the Noometry Index (35.9 vs 36.4), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-4.1 or Grok Build 0.1?

Grok Build 0.1 is cheaper. It lists at $1 per million input tokens and $2 per million output tokens; GPT-4.1 lists at $2 and $8.

Is GPT-4.1 or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 256K.

How many benchmarks do GPT-4.1 and Grok Build 0.1 share?

0 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper