Model comparison

Codellama 70b Instruct vs Grok Build 0.1

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 20.1.
  • Codellama 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 70b Instruct and Grok Build 0.1 specifications
Codellama 70b InstructGrok Build 0.1
ProviderMetaxAI
Noometry Index33.736.4
Released—2026-04-16
WeightsOpenProprietary
Context window—256K
Max output—256K
Input $ / M tokens—$1
Output $ / M tokens—$2
Results tracked73

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Codellama 70b Instruct: 37.6 (#193), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
SciCode—50.2%
BigCodeBench Instruct40.7%—
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
GBAEval—2.4%

Reasoning Grok Build 0.1 leads

Codellama 70b Instruct: 20.1 (#242), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
CritPt—9.1%
LMArena Hard Prompts1052—

Multilingual Not comparable

Codellama 70b Instruct: 24.8 (#288), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
LMArena Non-English992—

Instruction Following Not comparable

Codellama 70b Instruct: 51.9 (#293), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
LMArena Instruction Following1024—

Writing & Preference Not comparable

Codellama 70b Instruct: 33.4 (#277), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGrok Build 0.1
LMArena Text1057—

Frequently asked questions

Is Codellama 70b Instruct better than Grok Build 0.1?

Grok Build 0.1 is the stronger model overall, scoring 36.4 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Grok Build 0.1 share?

0 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper