Model comparison

C4ai Aya Expanse 32b vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 35.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

C4ai Aya Expanse 32b Cohere

35.9

Rank #221 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 17 benchmarks with published results for both. C4ai Aya Expanse 32b scores higher in 0 categories and Grok 4.6 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 23.3.
  • Grok 4.6 accepts more context: 500K tokens versus 128K.
  • C4ai Aya Expanse 32b has downloadable open weights; the other is API-only.

Side by side

C4ai Aya Expanse 32b and Grok 4.6 specifications
C4ai Aya Expanse 32bGrok 4.6
ProviderCoherexAI
Noometry Index35.956.9
Released2024-10-242026-08-12
WeightsOpenProprietary
Context window128K500K
Max output4K500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1849

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

C4ai Aya Expanse 32b: 34.8 (#231), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Coding11971465
DeepSWE—67.5%
FrontierCode—48%
CursorBench—41.4%
LMArena WebDev—1617
FrontierSWE—25.3%
SciCode—56.5%
WeirdML—67.3%
ALE-Bench—1,508

Agentic & Tool Use Not comparable

C4ai Aya Expanse 32b: —, Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
APEX-Agents—65.3%
GDP.pdf—17.2%
Vending-Bench 2—9,047

Reasoning Grok 4.6 leads

C4ai Aya Expanse 32b: 23.3 (#180), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Hard Prompts11931447
ARC-AGI-2—67.1%
SimpleBench—75.9%
NYT Connections (extended)—80%
ARC-AGI-1—87.5%
CritPt—19.7%
Chess Puzzles—40%
EBR-Bench—30.5%
Mystery Game Puzzles—34%
DTBench—97.3%
LMCA—48.5%
Epoch Capabilities Index—156.44

Math Grok 4.6 leads

C4ai Aya Expanse 32b: 34.0 (#197), Grok 4.6: 67.0 (#24)

Math benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Math12001423
FrontierMath (Tiers 1-3)—66%
FrontierMath Tier 4—31.7%
OTIS Mock AIME 2024-2025—99.2%
ProofBench—51%

Knowledge Grok 4.6 leads

C4ai Aya Expanse 32b: 33.2 (#206), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Expert11821467
GPQA Diamond—94%
SimpleQA Verified—49.3%
Vectara Hallucination Rate10.9%—

Multimodal Not comparable

C4ai Aya Expanse 32b: —, Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Vision—1263
Blueprint-Bench 2—33.2%
Furniture Assembly—40%
LMArena Document—1452

Multilingual Grok 4.6 leads

C4ai Aya Expanse 32b: 38.4 (#230), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Non-English12131420
LMArena Chinese12111480
LMArena French12481461
LMArena German11991431
LMArena Japanese11631376
LMArena Korean11581397
LMArena Russian12271422
LMArena Spanish11931404

Instruction Following Grok 4.6 leads

C4ai Aya Expanse 32b: 62.6 (#237), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Instruction Following11961431

Long Context Grok 4.6 leads

C4ai Aya Expanse 32b: 37.2 (#220), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Longer Query12281454

Writing & Preference Grok 4.6 leads

C4ai Aya Expanse 32b: 42.2 (#235), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkC4ai Aya Expanse 32bGrok 4.6
LMArena Text12241428
LMArena Creative Writing12001428
LMArena Multi-Turn11901425

Frequently asked questions

Is C4ai Aya Expanse 32b better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 35.9 on the Noometry Index.

Is C4ai Aya Expanse 32b or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 34.8 in the Noometry coding category.

Which has the bigger context window?

Grok 4.6 does, with 500K tokens against 128K.

How many benchmarks do C4ai Aya Expanse 32b and Grok 4.6 share?

17 benchmarks have published results for both models. C4ai Aya Expanse 32b has 18 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper