Model comparison

C4ai Aya Expanse 32b vs GPT-4.1

C4ai Aya Expanse 32b and GPT-4.1 score almost the same on the Noometry Index (35.9 vs 35.9), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

C4ai Aya Expanse 32b Cohere

35.9

Rank #221 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 18 benchmarks with published results for both. C4ai Aya Expanse 32b scores higher in 3 categories and GPT-4.1 in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GPT-4.1 leads 57.6 to 42.2.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 10.9% for C4ai Aya Expanse 32b and 5.6% for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 128K.
  • C4ai Aya Expanse 32b has downloadable open weights; the other is API-only.

Side by side

C4ai Aya Expanse 32b and GPT-4.1 specifications
C4ai Aya Expanse 32bGPT-4.1
ProviderCohereOpenAI
Noometry Index35.935.9
Released2024-10-242025-04-14
WeightsOpenProprietary
Context window128K1.05M
Max output4K33K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

C4ai Aya Expanse 32b: 34.8 (#231), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Coding11971391
SWE-bench Verified—48.5%
SWE-bench Verified (bash only)—39.6%
Aider Polyglot—52.4%
WeirdML—39%
CadEval—42%
ALE-Bench—558.1

Agentic & Tool Use Not comparable

C4ai Aya Expanse 32b: —, GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
Berkeley Function Calling Leaderboard—54%

Reasoning C4ai Aya Expanse 32b leads

C4ai Aya Expanse 32b: 23.3 (#180), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Hard Prompts11931384
ARC-AGI-2—0.4%
SimpleBench—27%
Kagi LLM Benchmark—52.3%
ARC-AGI-1—5.5%
Chess Puzzles—6%
EnigmaEval—2.2%
DTBench—68.3%
LMCA—25.6%
Epoch Capabilities Index—136.78
ForecastBench—61.5

Math C4ai Aya Expanse 32b leads

C4ai Aya Expanse 32b: 34.0 (#197), GPT-4.1: 22.3 (#280)

Math benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Math12001370
FrontierMath (Tiers 1-3)—6%
OTIS Mock AIME 2024-2025—38.3%
Omni-MATH—47.1%
MATH Level 5—83%
FrontierMath (Feb 2025 set)—5.5%
FrontierMath Tier 4 (v1)—0%

Knowledge GPT-4.1 leads

C4ai Aya Expanse 32b: 33.2 (#206), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
Vectara Hallucination Rate10.9%5.6%
LMArena Expert11821364
GPQA Diamond—66.9%
Humanity's Last Exam—5.4%
SimpleQA Verified—31.1%
MMLU-Pro—81.1%
GPQA (HELM)—65.9%

Multimodal Not comparable

C4ai Aya Expanse 32b: —, GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Vision—1211
GeoBench—72%

Multilingual GPT-4.1 leads

C4ai Aya Expanse 32b: 38.4 (#230), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Non-English12131370
LMArena Chinese12111382
LMArena French12481382
LMArena German11991381
LMArena Japanese11631319
LMArena Korean11581339
LMArena Russian12271377
LMArena Spanish11931376

Instruction Following GPT-4.1 leads

C4ai Aya Expanse 32b: 62.6 (#237), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Instruction Following11961367
IFEval—83.8%

Long Context GPT-4.1 leads

C4ai Aya Expanse 32b: 37.2 (#220), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Longer Query12281385
Fiction.LiveBench—63.9%

Writing & Preference GPT-4.1 leads

C4ai Aya Expanse 32b: 42.2 (#235), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkC4ai Aya Expanse 32bGPT-4.1
LMArena Text12241383
LMArena Creative Writing12001363
LMArena Multi-Turn11901398
EQ-Bench Creative Writing—1420
WildBench—85.4%

Frequently asked questions

Is C4ai Aya Expanse 32b better than GPT-4.1?

C4ai Aya Expanse 32b and GPT-4.1 score almost the same on the Noometry Index (35.9 vs 35.9), so choose on price, context window or the category you care about most.

Is C4ai Aya Expanse 32b or GPT-4.1 better for coding?

They score almost the same on coding (34.8 vs 34.4); test both on your own repository before choosing.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 128K.

How many benchmarks do C4ai Aya Expanse 32b and GPT-4.1 share?

18 benchmarks have published results for both models. C4ai Aya Expanse 32b has 18 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper