Model comparison

C4ai Aya Expanse 8b vs Llama 3.1 Tulu 3 8b

C4ai Aya Expanse 8b and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (34.9 vs 35.7), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

C4ai Aya Expanse 8b Cohere

34.9

Rank #229 Confirmed

Summary

  • They share 11 benchmarks with published results for both. C4ai Aya Expanse 8b scores higher in 2 categories and Llama 3.1 Tulu 3 8b in 5 categories; one gap is clear of the uncertainty.

Side by side

C4ai Aya Expanse 8b and Llama 3.1 Tulu 3 8b specifications
C4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
ProviderCohereAllen Institute for AI (Ai2)
Noometry Index34.935.7
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1511

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

C4ai Aya Expanse 8b: 33.8 (#252), Llama 3.1 Tulu 3 8b: 34.4 (#235)

Coding benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Coding11601183

Reasoning Too close to call

C4ai Aya Expanse 8b: 22.4 (#197), Llama 3.1 Tulu 3 8b: 22.8 (#188)

Reasoning benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Hard Prompts11561174

Math Too close to call

C4ai Aya Expanse 8b: 33.3 (#203), Llama 3.1 Tulu 3 8b: 33.9 (#198)

Math benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Math11681195

Knowledge Not comparable

C4ai Aya Expanse 8b: 33.6 (#201), Llama 3.1 Tulu 3 8b: —

Knowledge benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
Vectara Hallucination Rate9.5%—
LMArena Expert1153—

Multilingual Too close to call

C4ai Aya Expanse 8b: 36.1 (#242), Llama 3.1 Tulu 3 8b: 35.4 (#246)

Multilingual benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Non-English11801169
LMArena Chinese11811176
LMArena Russian11971193
LMArena German1188—
LMArena Japanese1120—

Instruction Following Llama 3.1 Tulu 3 8b leads

C4ai Aya Expanse 8b: 60.1 (#253), Llama 3.1 Tulu 3 8b: 61.3 (#246)

Instruction Following benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Instruction Following11551174

Long Context Too close to call

C4ai Aya Expanse 8b: 36.1 (#236), Llama 3.1 Tulu 3 8b: 35.8 (#239)

Long Context benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Longer Query11911181

Writing & Preference Too close to call

C4ai Aya Expanse 8b: 39.0 (#250), Llama 3.1 Tulu 3 8b: 39.7 (#245)

Writing & Preference benchmarks
BenchmarkC4ai Aya Expanse 8bLlama 3.1 Tulu 3 8b
LMArena Text11851193
LMArena Creative Writing11681182
LMArena Multi-Turn11601154

Frequently asked questions

Is C4ai Aya Expanse 8b better than Llama 3.1 Tulu 3 8b?

C4ai Aya Expanse 8b and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (34.9 vs 35.7), so choose on price, context window or the category you care about most.

Is C4ai Aya Expanse 8b or Llama 3.1 Tulu 3 8b better for coding?

They score almost the same on coding (33.8 vs 34.4); test both on your own repository before choosing.

How many benchmarks do C4ai Aya Expanse 8b and Llama 3.1 Tulu 3 8b share?

11 benchmarks have published results for both models. C4ai Aya Expanse 8b has 15 scored results on Noometry and Llama 3.1 Tulu 3 8b has 11.

Related comparisons

Go deeper