AI Models

The AI Model Scout

Two leaderboards: every model ranked on intelligence per dollar, and on how well it searches the web, against what it actually costs to run a task.

Last updated: August 31, 2026 · Source: https://artificialanalysis.ai/

Best price-to-quality

GPT-5.6 Luna (max)

OpenAI

Cheapest model that still keeps near-frontier intelligence

Intelligence
52.3
Web research
54.6
Cost/task
$0.049

GPT-5.6 Luna (max) scores 52.3 intelligence at $0.049 per task, 48x cheaper than the leader Claude Opus 5 (max).

Best for web search

GLM-5.3-Flash

Z AI

Smartest web researcher for the least money

Web research
94.2
Intelligence
57.5
Cost/task
$0.087

GLM-5.3-Flash tops web research at 94.2, with 57.5 intelligence for just $0.087 per task. Smarter than the value pick and nearly as cheap.

Best value right now

GPT-5.6 Luna (max)

OpenAI

Cheapest model that still keeps near-frontier intelligence

Intelligence
52.3
Cost/task
$0.049
Input
$0.20/M tokens
Output
$1.20/M tokens
Speed
129 tok/s
Web search
54.6
Value score
1074.3

Why this one

GPT-5.6 Luna (max) scores 52.3 on the Intelligence Index, 83% of the leader Claude Opus 5 (max), while costing $0.049 per task. That is 48x cheaper than the leader. It sits on the Pareto frontier, so no model is both smarter and cheaper.

Best for web search right now

GLM-5.3-Flash

Z AI

Smartest web researcher for the least money

Web search
94.2
Intelligence
57.5
Cost/task
$0.087
Context
1M
Most accurate (least hallucination)
72%

Why this one

GLM-5.3-Flash scores 94.2 on the deep-research benchmark, the highest on the board, while keeping 57.5 general intelligence and costing only $0.087 per task. For search-heavy work it beats far pricier models at a fraction of the cost.

What changed

Since the previous scan

Nothing moved since the last scan.

Most intelligence per dollar

Ranked by Intelligence Index against what one full task actually costs.

ModelProviderIntelligenceCost/taskInputOutputSpeedValue score
Claude Opus 5 (max)Best in classAnthropic63.0$2.34$5.00$25.005327.0
Claude Fable 5 (with fallback)Claude Opus 5 (max) is smarter and cheaperAnthropic62.1$3.14$10.00$50.006719.8
GPT-5.6 Sol (max)Best in classOpenAI60.9$0.95$4.00$20.007963.9
Grok 4.6 (high)Best in classSpaceXAI60.9$0.94$2.00$6.005965.0
Kimi K3 (max)Best in classKimi59.7$0.84$3.00$15.004071.3
GLM-5.3 (max)Best in classZ AI59.5$0.68$1.40$4.407887.1
Qwen3.8 2.4T A95BGLM-5.3 (max) is smarter and cheaperAlibaba57.7$0.81$2.00$6.00not published71.5
GLM-5.3-FlashBest in classZ AI57.5$0.087$0.15$0.50not published661.2
Muse Spark 1.2 (xhigh)GLM-5.3-Flash is smarter and cheaperMeta56.8$0.40$1.25$4.25not published142.2
GPT-5.6 Terra (max)GLM-5.3-Flash is smarter and cheaperOpenAI56.6$0.53$2.00$12.00122107.6
Gemini 3.7 Flash (high)GLM-5.3-Flash is smarter and cheaperGoogle56.0$0.40$0.75$3.75315139.3
DeepSeek V4 Pro 0813 (max)GLM-5.3-Flash is smarter and cheaperDeepSeek53.2$0.27$1.32$3.9654200.5
GPT-5.6 Luna (max)Best price-to-qualityOpenAI52.3$0.049$0.20$1.201291074.3
Qwen3.8 27B (xhigh)GLM-5.3-Flash is smarter and cheaperAlibaba52.0$0.37$0.50$3.0047141.0
Motif 3Motif Technologies47.4not publishednot publishednot publishednot publishednot published
MiniMax-M3GLM-5.3-Flash is smarter and cheaperMiniMax45.4$0.14$0.30$1.20150327.3
InklingGLM-5.3-Flash is smarter and cheaperThinking Machines42.3$0.34$1.00$4.0577124.8
Nemotron 3 UltraGLM-5.3-Flash is smarter and cheaperNVIDIA38.3$0.38$0.60$2.7570100.1
Gemini 3.5 Flash-LiteGLM-5.3-Flash is smarter and cheaperGoogle37.4$0.097$0.30$2.50358388.0
Solar Open2 250BUpstage37.4not publishednot publishednot publishednot publishednot published

Best at web search

Ranked by the deep-research benchmark: how well each model searches and synthesizes the web.

ModelProviderWeb searchIntelligenceCost/taskInputOutputContext
GLM-5.3-FlashBest researcherZ AI94.257.5$0.087$0.15$0.501M
GLM-5.3 (max)Z AI93.959.5$0.68$1.40$4.401M
Grok 4.6 (high)SpaceXAI90.660.9$0.94$2.00$6.00500K
Qwen3.8 2.4T A95BAlibaba87.757.7$0.81$2.00$6.00984K
Qwen3.8 27B (xhigh)Alibaba84.052.0$0.37$0.50$3.00256K
Muse Spark 1.2 (xhigh)Meta83.056.8$0.40$1.25$4.251M
Claude Opus 5 (max)Anthropic81.863.0$2.34$5.00$25.001M
Kimi K3 (max)Kimi78.459.7$0.84$3.00$15.001M
Claude Fable 5 (with fallback)Anthropic77.062.1$3.14$10.00$50.001M
MiniMax-M3MiniMax76.045.4$0.14$0.30$1.201M
GPT-5.6 Sol (max)OpenAI65.360.9$0.95$4.00$20.001M
Gemini 3.7 Flash (high)Google64.356.0$0.40$0.75$3.751M
Nemotron 3 UltraNVIDIA61.338.3$0.38$0.60$2.75262K
GPT-5.6 Terra (max)OpenAI59.356.6$0.53$2.00$12.001M
Gemini 3.5 Flash-LiteGoogle58.037.4$0.097$0.30$2.501M
DeepSeek V4 Pro 0813 (max)DeepSeek56.153.2$0.27$1.32$3.961M
GPT-5.6 Luna (max)OpenAI54.652.3$0.049$0.20$1.201M
InklingThinking Machines50.842.3$0.34$1.00$4.051M

Token prices are USD per million tokens. Cost per task is the USD cost of one full Intelligence Index run. Speed is output tokens per second. Web search is the deep-research benchmark score, and context is the maximum window in tokens.

At a glance

Highest intelligence:
Claude Opus 5 (max) (63.0)
Best at web search:
GLM-5.3-Flash (94.2)
Most accurate (least hallucination):
MiniMax-M3 (82%)
Fastest output:
GPT-5.6 Luna (max)
Lowest cost per task:
GPT-5.6 Luna (max)
Average cost per task:
$0.71
On the Pareto frontier:
7 of 20 models

How this is ranked

Value score is Intelligence Index divided by cost per task. The best value pick is the cheapest model still scoring at least 50.4, which is 80% of the leader Claude Opus 5 (max). Figures come straight from the source's published datasets, not from an AI summary.

Frequently Asked Questions

How is the best value model chosen?

We take every model that still scores at least 80% of the highest Intelligence Index on the board, then pick the cheapest of those by cost per task. That threshold keeps a cheap but weak model from winning on price alone.

Which model is best for web search?

We rank the deep-research benchmark, which measures how well a model gathers, cites and synthesizes information from the web, then read off the leader. The web-search champion is often a different, cheaper model than the raw intelligence leader, smarter at research and far less expensive to run.

Why cost per task instead of price per token?

Reasoning models burn large numbers of hidden thinking tokens, so a low per-token price can still produce an expensive request. Cost per task measures what one full benchmark run actually costs, which is what you pay in practice.

Where do these numbers come from?

Directly from the benchmark datasets Artificial Analysis publishes with its charts. We read them as data rather than asking a language model to transcribe them, so the figures here match the source exactly.

What does the Pareto frontier mean here?

A model is on the frontier when no other model is both smarter and cheaper than it. Anything off the frontier is strictly beaten by something else, and we name what beats it.