Model comparison

Codellama 70b Instruct vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and Qwen Plus in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen Plus leads 45.1 to 24.8.
  • Codellama 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 70b Instruct and Qwen Plus specifications
Codellama 70b InstructQwen Plus
ProviderMetaAlibaba (Qwen)
Noometry Index33.737.1
Released—2024-01-25
WeightsOpenProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked720

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Codellama 70b Instruct: 37.6 (#193), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkCodellama 70b InstructQwen Plus
BigCodeBench Instruct40.7%—
LMArena Coding—1328
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning Qwen Plus leads

Codellama 70b Instruct: 20.1 (#242), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkCodellama 70b InstructQwen Plus
LMArena Hard Prompts10521317
Kagi LLM Benchmark—63.3%
DTBench—81.1%
LMCA—24%

Math Not comparable

Codellama 70b Instruct: —, Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkCodellama 70b InstructQwen Plus
OTIS Mock AIME 2024-2025—17.8%
LMArena Math—1326
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Not comparable

Codellama 70b Instruct: —, Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkCodellama 70b InstructQwen Plus
GPQA Diamond—48.1%
LMArena Expert—1328

Multilingual Qwen Plus leads

Codellama 70b Instruct: 24.8 (#288), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkCodellama 70b InstructQwen Plus
LMArena Non-English9921310
LMArena Chinese—1347
LMArena Japanese—1251
LMArena Russian—1323

Instruction Following Qwen Plus leads

Codellama 70b Instruct: 51.9 (#293), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructQwen Plus
LMArena Instruction Following10241303

Long Context Not comparable

Codellama 70b Instruct: —, Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkCodellama 70b InstructQwen Plus
LMArena Longer Query—1324

Writing & Preference Qwen Plus leads

Codellama 70b Instruct: 33.4 (#277), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructQwen Plus
LMArena Text10571326
LMArena Creative Writing—1293
LMArena Multi-Turn—1336

Frequently asked questions

Is Codellama 70b Instruct better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Qwen Plus share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper