Model comparison

Codellama 70b Instruct vs GLM-4.6V

GLM-4.6V is the stronger model overall, scoring 41.3 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

GLM-4.6V Z.ai (Zhipu)

41.3

Rank #137 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and GLM-4.6V in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GLM-4.6V leads 48.6 to 24.8.

Side by side

Codellama 70b Instruct and GLM-4.6V specifications
Codellama 70b InstructGLM-4.6V
ProviderMetaZ.ai (Zhipu)
Noometry Index33.741.3
Released—2025-12-08
WeightsOpenOpen
Context window—128K
Max output—33K
Input $ / M tokens—$0.30
Output $ / M tokens—$0.90
Results tracked712

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.6V leads

Codellama 70b Instruct: 37.6 (#193), GLM-4.6V: 40.9 (#128)

Coding benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
BigCodeBench Instruct40.7%—
LMArena Coding—1390
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning GLM-4.6V leads

Codellama 70b Instruct: 20.1 (#242), GLM-4.6V: 27.6 (#115)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Hard Prompts10521368

Knowledge Not comparable

Codellama 70b Instruct: —, GLM-4.6V: 38.0 (#149)

Knowledge benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Expert—1371

Multimodal Not comparable

Codellama 70b Instruct: —, GLM-4.6V: 34.8 (#90)

Multimodal benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Vision—1164

Multilingual GLM-4.6V leads

Codellama 70b Instruct: 24.8 (#288), GLM-4.6V: 48.6 (#141)

Multilingual benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Non-English9921359
LMArena Chinese—1425
LMArena Russian—1340

Instruction Following GLM-4.6V leads

Codellama 70b Instruct: 51.9 (#293), GLM-4.6V: 71.4 (#151)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Instruction Following10241352

Long Context Not comparable

Codellama 70b Instruct: —, GLM-4.6V: 41.3 (#143)

Long Context benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Longer Query—1358

Writing & Preference GLM-4.6V leads

Codellama 70b Instruct: 33.4 (#277), GLM-4.6V: 56.6 (#137)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGLM-4.6V
LMArena Text10571377
LMArena Creative Writing—1347
LMArena Multi-Turn—1360

Frequently asked questions

Is Codellama 70b Instruct better than GLM-4.6V?

GLM-4.6V is the stronger model overall, scoring 41.3 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or GLM-4.6V better for coding?

GLM-4.6V scores higher on coding benchmarks: 40.9 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and GLM-4.6V share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and GLM-4.6V has 12.

Related comparisons

Go deeper