Model comparison

Codellama 70b Instruct vs Gemini 3.1 Pro Preview

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and Gemini 3.1 Pro Preview in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 20.1.
  • Codellama 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 70b Instruct and Gemini 3.1 Pro Preview specifications
Codellama 70b InstructGemini 3.1 Pro Preview
ProviderMetaGoogle
Noometry Index33.756.7
Released—2026-02-19
WeightsOpenProprietary
Context window—1.05M
Max output—66K
Input $ / M tokens—$2
Output $ / M tokens—$12
Results tracked771

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3.1 Pro Preview leads

Codellama 70b Instruct: 37.6 (#193), Gemini 3.1 Pro Preview: 42.5 (#99)

Coding benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
SWE-bench Verified—75.6%
DeepSWE—11.7%
LMArena WebDev—1447
SciCode—58.9%
GSO—22.6%
WeirdML—72.1%
BigCodeBench Instruct40.7%—
LMArena Coding—1484
MirrorCode—8.9%
BigCodeBench Complete49.6%—
ALE-Bench—1,161
AlgoTune—2.02
HumanEval+65.9%—

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Gemini 3.1 Pro Preview: 37.7 (#34)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
Terminal-Bench—80.2%
APEX-Agents—35.3%
τ²-bench Banking—26%
DeepResearch Bench—47.8%
PostTrainBench—22%
BALROG—57%
ExploitBench—26.1%
GBAEval—0.8%
GDP.pdf—17%
LMArena Search—1211
METR Time Horizons—77%
Vending-Bench 2—3,774

Reasoning Gemini 3.1 Pro Preview leads

Codellama 70b Instruct: 20.1 (#242), Gemini 3.1 Pro Preview: 71.7 (#12)

Reasoning benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
LMArena Hard Prompts10521485
ARC-AGI-2—77.1%
SimpleBench—79.6%
NYT Connections (extended)—97.4%
ARC-AGI-1—98%
CritPt—17.7%
Chess Puzzles—55%
EnigmaEval—36.8%
Thematic Generalization—79.4%
EBR-Bench—14.3%
Mystery Game Puzzles—34%
DTBench—97.1%
LMCA—53.8%
Epoch Capabilities Index—154.77
ForecastBench—59

Math Not comparable

Codellama 70b Instruct: —, Gemini 3.1 Pro Preview: 62.1 (#34)

Math benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
FrontierMath (Tiers 1-3)—59.6%
FrontierMath Tier 4—26.8%
MathArena Final-Answer Competitions—86.5%
OTIS Mock AIME 2024-2025—95.6%
ProofBench—26%
LMArena Math—1485
FrontierMath (Feb 2025 set)—36.9%
FrontierMath Tier 4 (v1)—16.7%

Knowledge Not comparable

Codellama 70b Instruct: —, Gemini 3.1 Pro Preview: 71.8 (#3)

Knowledge benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
GPQA Diamond—94.4%
Humanity's Last Exam—46.4%
SimpleQA Verified—73.5%
Vectara Hallucination Rate—10.4%
LMArena Expert—1485

Multimodal Not comparable

Codellama 70b Instruct: —, Gemini 3.1 Pro Preview: 37.9 (#69)

Multimodal benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
LMArena Vision—1296
Blueprint-Bench 2—26.5%
Furniture Assembly—26.7%
LMArena Document—1444

Multilingual Gemini 3.1 Pro Preview leads

Codellama 70b Instruct: 24.8 (#288), Gemini 3.1 Pro Preview: 57.0 (#12)

Multilingual benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
LMArena Non-English9921477
LMArena Chinese—1529
LMArena French—1487
LMArena German—1491
LMArena Japanese—1493
LMArena Korean—1455
LMArena Russian—1498
LMArena Spanish—1479

Instruction Following Gemini 3.1 Pro Preview leads

Codellama 70b Instruct: 51.9 (#293), Gemini 3.1 Pro Preview: 77.0 (#32)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
LMArena Instruction Following10241466

Long Context Not comparable

Codellama 70b Instruct: —, Gemini 3.1 Pro Preview: 47.4 (#18)

Long Context benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
CL-bench—20.8%
CL-bench Life—16.9%
LMArena Longer Query—1483

Writing & Preference Gemini 3.1 Pro Preview leads

Codellama 70b Instruct: 33.4 (#277), Gemini 3.1 Pro Preview: 66.1 (#37)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructGemini 3.1 Pro Preview
LMArena Text10571481
LMArena Creative Writing—1482
EQ-Bench Creative Writing—1491
EQ-Bench 4—1142
LMArena Multi-Turn—1488

Frequently asked questions

Is Codellama 70b Instruct better than Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or Gemini 3.1 Pro Preview better for coding?

Gemini 3.1 Pro Preview scores higher on coding benchmarks: 42.5 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Gemini 3.1 Pro Preview share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Gemini 3.1 Pro Preview has 71.

Related comparisons

Go deeper