Agentic & Tool Use benchmark

Vending-Bench 2 leaderboard

As of October 2026, GPT-6 Astra has the highest published Vending-Bench 2 score on Noometry at 15,515, out of 60 models with results.

Last verified

About Vending-Bench 2

Running a simulated vending-machine business over a long horizon. The score is the final bank balance.

Category
Agentic & Tool Use
Introduced
2025
Format
Long-horizon simulation
Unit
Raw score
Official site
andonlabs.com

Top 15 models

Top models on Vending-Bench 2
  1. GPT-6 Astra 15,515
  2. GPT-6 Sol 14,428
  3. Gemini 4 Argon 13,718
  4. Claude Opus 5 11,182
  5. Claude Opus 4.7 10,937
  6. Grok 4.7 10,537
  7. GPT-5.6 Sol 9,619
  8. Claude Opus 5.5 9,235
  9. Grok 4.6 9,047
  10. GLM-5.2 8,314
  11. GLM-5.3 8,164
  12. Claude Opus 4.6 8,018
  13. GPT-5.5 7,524
  14. GPT-5.6 Terra 7,343
  15. Claude Sonnet 4.6 7,204

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Vending-Bench 2 results by model
#ModelProviderRatingSettingSourceDate
1GPT-6 Astra OpenAI15,515Epoch AI
2GPT-6 Sol OpenAI14,428Epoch AI
3Gemini 4 Argon Google13,718Epoch AI
4Claude Opus 5 Anthropic11,182Epoch AI
5Claude Opus 4.7 Anthropic10,937Epoch AI
6Grok 4.7 xAI10,537Epoch AI
7GPT-5.6 Sol OpenAI9,619Epoch AI
8Claude Opus 5.5 Anthropic9,235Epoch AI
9Grok 4.6 xAI9,047Epoch AI
10GLM-5.2 Z.ai (Zhipu)8,314Epoch AI
11GLM-5.3 Z.ai (Zhipu)8,164Epoch AI
12Claude Opus 4.6 Anthropic8,018Epoch AI
13GPT-5.5 OpenAI7,524Epoch AI
14GPT-5.6 Terra OpenAI7,343Epoch AI
15Claude Sonnet 4.6 Anthropic7,204Epoch AI
16Muse Spark 1.1 Meta6,520Epoch AI
17Claude Sonnet 5 Anthropic6,378Epoch AI
18Kimi K2.6 Moonshot AI6,205Epoch AI
19GPT-5.4 OpenAI6,144Epoch AI
20GPT-5.3 Codex OpenAI5,940Epoch AI
21Claude Opus 4.8 Anthropic5,787Epoch AI
22Claude Fable 5 Anthropic5,680highEpoch AI
23GLM-5.1 Z.ai (Zhipu)5,634Epoch AI
24Gemini 3 Pro Google5,478Epoch AI
25Claude Fable 5.1 Anthropic5,422Epoch AI
26Gemini 3.5 Flash Google5,396Epoch AI
27Kimi K3 Moonshot AI5,165Epoch AI
28Qwen3.6 Plus Alibaba (Qwen)5,115Epoch AI
29Gemini 3.8 Flash Google5,094Epoch AI
30Kimi K2.7 Code Moonshot AI5,083Epoch AI
31Claude Opus 4.5 Anthropic4,967Epoch AI
32Grok 4.20 (Non-Reasoning) xAI4,663Epoch AI
33GLM-5 Z.ai (Zhipu)4,432Epoch AI
34Qwen3.6 Max Preview Alibaba (Qwen)4,254Epoch AI
35GPT-5.6 Luna OpenAI4,095Epoch AI
36Grok 4.5 xAI3,887Epoch AI
37Claude Sonnet 4.5 Anthropic3,839Epoch AI
38Gemini 3.1 Pro Preview Google3,774Epoch AI
39Gemini 3 Flash Preview Google3,635Epoch AI
40GPT-5.2 OpenAI3,591Epoch AI
41DeepSeek V4 Pro DeepSeek3,285Epoch AI
42GLM-4.7 Z.ai (Zhipu)2,377Epoch AI
43MiniMax-M3 MiniMax2,158Epoch AI
44GPT-5.1 OpenAI1,473Epoch AI
45Kimi K2.5 Moonshot AI1,198Epoch AI
46Grok 4.1 Fast xAI1,107Epoch AI
47DeepSeek-V3.2-Exp DeepSeek1,034Epoch AI
48Gemini 2.5 Pro Google573.64Epoch AI
49Gemini 2.5 Flash Google548.84Epoch AI
50Qwen3.5-Flash Alibaba (Qwen)462.69Epoch AI
51Claude Haiku 4.5 Anthropic458.89Epoch AI
52Qwen3.5 27B Alibaba (Qwen)201.98Epoch AI
53MiniMax-M2 MiniMax160.6Epoch AI
54Qwen3 Max Alibaba (Qwen)71.56Epoch AI
55Grok 4.3 xAI35.26Epoch AI
56Qwen3.5 Plus Alibaba (Qwen)0.54Epoch AI
57Qwen3 235B-A22B Alibaba (Qwen)-11.34Epoch AI
58gpt-oss-120b OpenAI-21.53Epoch AI
59MiniMax-M2.5 MiniMax-23.16Epoch AI
60GPT-5 Mini OpenAI-31.18Epoch AI

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does Vending-Bench 2 measure?

Running a simulated vending-machine business over a long horizon. The score is the final bank balance.

Which model has the highest Vending-Bench 2 score?

As of October 2026, GPT-6 Astra has the highest published Vending-Bench 2 score on Noometry at 15,515, out of 60 models with results.

What is the best open-weight model on Vending-Bench 2?

GLM-5.2 has the highest Vending-Bench 2 score among open-weight models at 8,314, ranking 10 of 60 overall.