Model comparison

GPT-6 Luna vs StarCoder 2 3B

GPT-6 Luna has enough public results to be ranked (#36); StarCoder 2 3B does not yet, so treat this comparison as directional.

Last verified . 1 shared benchmarks.

GPT-6 Luna OpenAI

53.3

Rank #36 Confirmed

StarCoder 2 3B NVIDIA

36.3

Unranked Sparse

Summary

  • They share 1 benchmark with published results for both. GPT-6 Luna scores higher in 1 category and StarCoder 2 3B in 0 categories; one gap is clear of the uncertainty.
  • The widest gap is in coding, where GPT-6 Luna leads 55.5 to 33.5.
  • StarCoder 2 3B has downloadable open weights; the other is API-only.

Side by side

GPT-6 Luna and StarCoder 2 3B specifications
GPT-6 LunaStarCoder 2 3B
ProviderOpenAINVIDIA
Noometry Index53.336.3
Released2026-09-222024-02-22
WeightsProprietaryOpen
Context window1.05M—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.50—
Results tracked427

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Luna leads

GPT-6 Luna: 55.5 (#25), StarCoder 2 3B: 33.5

Coding benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
DeepSWE66.6%—
FrontierCode42.4%—
LMArena WebDev1581—
SciCode54.6%—
LMArena Coding1439—
BigCodeBench Complete—21.4%
ALE-Bench1,577—
HumanEval+—27.4%

Agentic & Tool Use Not comparable

GPT-6 Luna: 33.3 (#54), StarCoder 2 3B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
APEX-Agents44.3%—
GDP.pdf23%—

Reasoning Not comparable

GPT-6 Luna: 48.2 (#41), StarCoder 2 3B: —

Reasoning benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
Epoch Capabilities Index156.2888.76
ARC-AGI-259.3%—
NYT Connections (extended)68.7%—
ARC-AGI-186.7%—
CritPt19.4%—
Chess Puzzles31%—
LMArena Hard Prompts1411—
Mystery Game Puzzles7%—
DTBench90.1%—
LMCA44.5%—
WinoGrande—57.1%

Math Not comparable

GPT-6 Luna: 76.1 (#15), StarCoder 2 3B: —

Math benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
FrontierMath (Tiers 1-3)78.9%—
FrontierMath Tier 456.1%—
OTIS Mock AIME 2024-202598.9%—
ProofBench64%—
LMArena Math1416—
GSM8K—21.6%

Knowledge Not comparable

GPT-6 Luna: 57.0 (#41), StarCoder 2 3B: —

Knowledge benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
GPQA Diamond90.5%—
SimpleQA Verified41.4%—
LMArena Expert1444—
ARC (AI2) Challenge—34.2%
MMLU—36.6%

Multimodal Not comparable

GPT-6 Luna: 42.4 (#30), StarCoder 2 3B: —

Multimodal benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
LMArena Vision1217—
Blueprint-Bench 231.2%—
Furniture Assembly44.2%—

Multilingual Not comparable

GPT-6 Luna: 50.5 (#117), StarCoder 2 3B: —

Multilingual benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
LMArena Non-English1386—
LMArena Chinese1433—
LMArena French1420—
LMArena German1369—
LMArena Japanese1369—
LMArena Korean1360—
LMArena Russian1394—
LMArena Spanish1393—

Instruction Following Not comparable

GPT-6 Luna: 74.3 (#99), StarCoder 2 3B: —

Instruction Following benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
LMArena Instruction Following1409—

Long Context Not comparable

GPT-6 Luna: 43.0 (#111), StarCoder 2 3B: —

Long Context benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
LMArena Longer Query1409—

Writing & Preference Not comparable

GPT-6 Luna: 58.3 (#119), StarCoder 2 3B: —

Writing & Preference benchmarks
BenchmarkGPT-6 LunaStarCoder 2 3B
LMArena Text1391—
LMArena Creative Writing1363—
LMArena Multi-Turn1396—

Frequently asked questions

Is GPT-6 Luna better than StarCoder 2 3B?

GPT-6 Luna has enough public results to be ranked (#36); StarCoder 2 3B does not yet, so treat this comparison as directional.

Is GPT-6 Luna or StarCoder 2 3B better for coding?

GPT-6 Luna scores higher on coding benchmarks: 55.5 versus 33.5 in the Noometry coding category.

How many benchmarks do GPT-6 Luna and StarCoder 2 3B share?

1 benchmark has published results for both models. GPT-6 Luna has 42 scored results on Noometry and StarCoder 2 3B has 7.

Related comparisons

Go deeper