Model comparison

Codestral vs Qwen1.5 4b Chat

Codestral is the stronger model overall, scoring 30.6 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Codestral and Qwen1.5 4b Chat specifications
CodestralQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.628.8
Released2024-05-29—
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Codestral: 27.3 (#321), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkCodestralQwen1.5 4b Chat
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—999
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Codestral leads

Codestral: 19.8 (#251), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkCodestralQwen1.5 4b Chat
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—976

Math Not comparable

Codestral: —, Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Math—1026

Knowledge Not comparable

Codestral: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Expert—980

Multilingual Not comparable

Codestral: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Codestral: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Codestral: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Not comparable

Codestral: —, Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkCodestralQwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
LMArena Multi-Turn—977

Frequently asked questions

Is Codestral better than Qwen1.5 4b Chat?

Codestral is the stronger model overall, scoring 30.6 to 28.8 on the Noometry Index.

Is Codestral or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper