Model comparison

GLM-4.7-Flash vs Qwen3.6 Flash

GLM-4.7-Flash and Qwen3.6 Flash score almost the same on the Noometry Index (38.8 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Qwen3.6 Flash Alibaba (Qwen)

38.8

Rank #182 Confirmed

Summary

  • They share 3 benchmarks with published results for both. GLM-4.7-Flash scores higher in 0 categories and Qwen3.6 Flash in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Flash leads 29.0 to 20.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 58.3% for GLM-4.7-Flash and 84.4% for Qwen3.6 Flash.
  • GLM-4.7-Flash is cheaper at $0.06 / $0.40 per million input/output tokens, against $0.19 / $1.13 for Qwen3.6 Flash.
  • Qwen3.6 Flash accepts more context: 1M tokens versus 200K.
  • GLM-4.7-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-4.7-Flash and Qwen3.6 Flash specifications
GLM-4.7-FlashQwen3.6 Flash
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index38.838.8
Released2026-01-192026-04-27
WeightsOpenProprietary
Context window200K1M
Max output131K66K
Input $ / M tokens$0.06$0.19
Output $ / M tokens$0.40$1.13
Results tracked2113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GLM-4.7-Flash: 40.6 (#135), Qwen3.6 Flash: —

Coding benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
LMArena Coding1383—
ALE-Bench—326.4

Reasoning Qwen3.6 Flash leads

GLM-4.7-Flash: 20.9 (#229), Qwen3.6 Flash: 29.0 (#96)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
Chess Puzzles0%20%
SimpleBench—35.2%
LMArena Hard Prompts1356—
Mystery Game Puzzles—18%
DTBench—77.1%
LMCA—31%
Epoch Capabilities Index—143.26

Math Qwen3.6 Flash leads

GLM-4.7-Flash: 36.1 (#173), Qwen3.6 Flash: 39.0 (#117)

Math benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
OTIS Mock AIME 2024-202558.3%84.4%
FrontierMath (Tiers 1-3)—22.5%
LMArena Math1355—
FrontierMath (Feb 2025 set)—10.3%
FrontierMath Tier 4 (v1)—0%

Knowledge Qwen3.6 Flash leads

GLM-4.7-Flash: 35.5 (#184), Qwen3.6 Flash: 42.1 (#100)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
GPQA Diamond60.5%83.3%
SimpleQA Verified—15.9%
Vectara Hallucination Rate9.3%—
LMArena Expert1357—

Multilingual Not comparable

GLM-4.7-Flash: 46.5 (#158), Qwen3.6 Flash: —

Multilingual benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
LMArena Non-English1330—
LMArena Chinese1403—
LMArena French1332—
LMArena German1337—
LMArena Korean1283—
LMArena Russian1332—
LMArena Spanish1350—

Instruction Following Not comparable

GLM-4.7-Flash: 70.1 (#167), Qwen3.6 Flash: —

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
LMArena Instruction Following1327—

Long Context Not comparable

GLM-4.7-Flash: 40.9 (#148), Qwen3.6 Flash: —

Long Context benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
LMArena Longer Query1345—

Writing & Preference Not comparable

GLM-4.7-Flash: 47.4 (#210), Qwen3.6 Flash: —

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashQwen3.6 Flash
LMArena Text1351—
LMArena Creative Writing1297—
EQ-Bench Creative Writing1125—
LMArena Multi-Turn1342—

Frequently asked questions

Is GLM-4.7-Flash better than Qwen3.6 Flash?

GLM-4.7-Flash and Qwen3.6 Flash score almost the same on the Noometry Index (38.8 vs 38.8), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-4.7-Flash or Qwen3.6 Flash?

GLM-4.7-Flash is cheaper. It lists at $0.06 per million input tokens and $0.40 per million output tokens; Qwen3.6 Flash lists at $0.19 and $1.13.

Which has the bigger context window?

Qwen3.6 Flash does, with 1M tokens against 200K.

How many benchmarks do GLM-4.7-Flash and Qwen3.6 Flash share?

3 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Qwen3.6 Flash has 13.

Related comparisons

Go deeper