The AI Model Scout
Two leaderboards: every model ranked on intelligence per dollar, and on how well it searches the web, against what it actually costs to run a task.
Last updated: August 31, 2026 · Source: https://artificialanalysis.ai/
GPT-5.6 Luna (max)
OpenAICheapest model that still keeps near-frontier intelligence
- Intelligence
- 52.3
- Web research
- 54.6
- Cost/task
- $0.049
GPT-5.6 Luna (max) scores 52.3 intelligence at $0.049 per task, 48x cheaper than the leader Claude Opus 5 (max).
GLM-5.3-Flash
Z AISmartest web researcher for the least money
- Web research
- 94.2
- Intelligence
- 57.5
- Cost/task
- $0.087
GLM-5.3-Flash tops web research at 94.2, with 57.5 intelligence for just $0.087 per task. Smarter than the value pick and nearly as cheap.
Best value right now
GPT-5.6 Luna (max)
OpenAICheapest model that still keeps near-frontier intelligence
- Intelligence
- 52.3
- Cost/task
- $0.049
- Input
- $0.20/M tokens
- Output
- $1.20/M tokens
- Speed
- 129 tok/s
- Web search
- 54.6
- Value score
- 1074.3
Why this one
GPT-5.6 Luna (max) scores 52.3 on the Intelligence Index, 83% of the leader Claude Opus 5 (max), while costing $0.049 per task. That is 48x cheaper than the leader. It sits on the Pareto frontier, so no model is both smarter and cheaper.
Best for web search right now
GLM-5.3-Flash
Z AISmartest web researcher for the least money
- Web search
- 94.2
- Intelligence
- 57.5
- Cost/task
- $0.087
- Context
- 1M
- Most accurate (least hallucination)
- 72%
Why this one
GLM-5.3-Flash scores 94.2 on the deep-research benchmark, the highest on the board, while keeping 57.5 general intelligence and costing only $0.087 per task. For search-heavy work it beats far pricier models at a fraction of the cost.
What changed
Since the previous scan
Nothing moved since the last scan.
Most intelligence per dollar
Ranked by Intelligence Index against what one full task actually costs.
| Model | Provider | Intelligence | Cost/task | Input | Output | Speed | Value score |
|---|---|---|---|---|---|---|---|
| Claude Opus 5 (max)Best in class | Anthropic | 63.0 | $2.34 | $5.00 | $25.00 | 53 | 27.0 |
| Claude Fable 5 (with fallback)Claude Opus 5 (max) is smarter and cheaper | Anthropic | 62.1 | $3.14 | $10.00 | $50.00 | 67 | 19.8 |
| GPT-5.6 Sol (max)Best in class | OpenAI | 60.9 | $0.95 | $4.00 | $20.00 | 79 | 63.9 |
| Grok 4.6 (high)Best in class | SpaceXAI | 60.9 | $0.94 | $2.00 | $6.00 | 59 | 65.0 |
| Kimi K3 (max)Best in class | Kimi | 59.7 | $0.84 | $3.00 | $15.00 | 40 | 71.3 |
| GLM-5.3 (max)Best in class | Z AI | 59.5 | $0.68 | $1.40 | $4.40 | 78 | 87.1 |
| Qwen3.8 2.4T A95BGLM-5.3 (max) is smarter and cheaper | Alibaba | 57.7 | $0.81 | $2.00 | $6.00 | not published | 71.5 |
| GLM-5.3-FlashBest in class | Z AI | 57.5 | $0.087 | $0.15 | $0.50 | not published | 661.2 |
| Muse Spark 1.2 (xhigh)GLM-5.3-Flash is smarter and cheaper | Meta | 56.8 | $0.40 | $1.25 | $4.25 | not published | 142.2 |
| GPT-5.6 Terra (max)GLM-5.3-Flash is smarter and cheaper | OpenAI | 56.6 | $0.53 | $2.00 | $12.00 | 122 | 107.6 |
| Gemini 3.7 Flash (high)GLM-5.3-Flash is smarter and cheaper | 56.0 | $0.40 | $0.75 | $3.75 | 315 | 139.3 | |
| DeepSeek V4 Pro 0813 (max)GLM-5.3-Flash is smarter and cheaper | DeepSeek | 53.2 | $0.27 | $1.32 | $3.96 | 54 | 200.5 |
| GPT-5.6 Luna (max)Best price-to-quality | OpenAI | 52.3 | $0.049 | $0.20 | $1.20 | 129 | 1074.3 |
| Qwen3.8 27B (xhigh)GLM-5.3-Flash is smarter and cheaper | Alibaba | 52.0 | $0.37 | $0.50 | $3.00 | 47 | 141.0 |
| Motif 3 | Motif Technologies | 47.4 | not published | not published | not published | not published | not published |
| MiniMax-M3GLM-5.3-Flash is smarter and cheaper | MiniMax | 45.4 | $0.14 | $0.30 | $1.20 | 150 | 327.3 |
| InklingGLM-5.3-Flash is smarter and cheaper | Thinking Machines | 42.3 | $0.34 | $1.00 | $4.05 | 77 | 124.8 |
| Nemotron 3 UltraGLM-5.3-Flash is smarter and cheaper | NVIDIA | 38.3 | $0.38 | $0.60 | $2.75 | 70 | 100.1 |
| Gemini 3.5 Flash-LiteGLM-5.3-Flash is smarter and cheaper | 37.4 | $0.097 | $0.30 | $2.50 | 358 | 388.0 | |
| Solar Open2 250B | Upstage | 37.4 | not published | not published | not published | not published | not published |
Best at web search
Ranked by the deep-research benchmark: how well each model searches and synthesizes the web.
| Model | Provider | Web search | Intelligence | Cost/task | Input | Output | Context |
|---|---|---|---|---|---|---|---|
| GLM-5.3-FlashBest researcher | Z AI | 94.2 | 57.5 | $0.087 | $0.15 | $0.50 | 1M |
| GLM-5.3 (max) | Z AI | 93.9 | 59.5 | $0.68 | $1.40 | $4.40 | 1M |
| Grok 4.6 (high) | SpaceXAI | 90.6 | 60.9 | $0.94 | $2.00 | $6.00 | 500K |
| Qwen3.8 2.4T A95B | Alibaba | 87.7 | 57.7 | $0.81 | $2.00 | $6.00 | 984K |
| Qwen3.8 27B (xhigh) | Alibaba | 84.0 | 52.0 | $0.37 | $0.50 | $3.00 | 256K |
| Muse Spark 1.2 (xhigh) | Meta | 83.0 | 56.8 | $0.40 | $1.25 | $4.25 | 1M |
| Claude Opus 5 (max) | Anthropic | 81.8 | 63.0 | $2.34 | $5.00 | $25.00 | 1M |
| Kimi K3 (max) | Kimi | 78.4 | 59.7 | $0.84 | $3.00 | $15.00 | 1M |
| Claude Fable 5 (with fallback) | Anthropic | 77.0 | 62.1 | $3.14 | $10.00 | $50.00 | 1M |
| MiniMax-M3 | MiniMax | 76.0 | 45.4 | $0.14 | $0.30 | $1.20 | 1M |
| GPT-5.6 Sol (max) | OpenAI | 65.3 | 60.9 | $0.95 | $4.00 | $20.00 | 1M |
| Gemini 3.7 Flash (high) | 64.3 | 56.0 | $0.40 | $0.75 | $3.75 | 1M | |
| Nemotron 3 Ultra | NVIDIA | 61.3 | 38.3 | $0.38 | $0.60 | $2.75 | 262K |
| GPT-5.6 Terra (max) | OpenAI | 59.3 | 56.6 | $0.53 | $2.00 | $12.00 | 1M |
| Gemini 3.5 Flash-Lite | 58.0 | 37.4 | $0.097 | $0.30 | $2.50 | 1M | |
| DeepSeek V4 Pro 0813 (max) | DeepSeek | 56.1 | 53.2 | $0.27 | $1.32 | $3.96 | 1M |
| GPT-5.6 Luna (max) | OpenAI | 54.6 | 52.3 | $0.049 | $0.20 | $1.20 | 1M |
| Inkling | Thinking Machines | 50.8 | 42.3 | $0.34 | $1.00 | $4.05 | 1M |
Token prices are USD per million tokens. Cost per task is the USD cost of one full Intelligence Index run. Speed is output tokens per second. Web search is the deep-research benchmark score, and context is the maximum window in tokens.
At a glance
- Highest intelligence:
- Claude Opus 5 (max) (63.0)
- Best at web search:
- GLM-5.3-Flash (94.2)
- Most accurate (least hallucination):
- MiniMax-M3 (82%)
- Fastest output:
- GPT-5.6 Luna (max)
- Lowest cost per task:
- GPT-5.6 Luna (max)
- Average cost per task:
- $0.71
- On the Pareto frontier:
- 7 of 20 models
How this is ranked
Value score is Intelligence Index divided by cost per task. The best value pick is the cheapest model still scoring at least 50.4, which is 80% of the leader Claude Opus 5 (max). Figures come straight from the source's published datasets, not from an AI summary.
Frequently Asked Questions
How is the best value model chosen?
We take every model that still scores at least 80% of the highest Intelligence Index on the board, then pick the cheapest of those by cost per task. That threshold keeps a cheap but weak model from winning on price alone.
Which model is best for web search?
We rank the deep-research benchmark, which measures how well a model gathers, cites and synthesizes information from the web, then read off the leader. The web-search champion is often a different, cheaper model than the raw intelligence leader, smarter at research and far less expensive to run.
Why cost per task instead of price per token?
Reasoning models burn large numbers of hidden thinking tokens, so a low per-token price can still produce an expensive request. Cost per task measures what one full benchmark run actually costs, which is what you pay in practice.
Where do these numbers come from?
Directly from the benchmark datasets Artificial Analysis publishes with its charts. We read them as data rather than asking a language model to transcribe them, so the figures here match the source exactly.
What does the Pareto frontier mean here?
A model is on the frontier when no other model is both smarter and cheaper than it. Anything off the frontier is strictly beaten by something else, and we name what beats it.