Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

THE MODEL SELECTION SERIES

LLM API pricing calculator

API prices are usually quoted per million tokens, but your budget depends on how much context you send and how much output the model bills. Change the calculator to compare the same workload across providers.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗ · Download snapshot ↓

01 / WHAT TO CHECK

Read input and output prices together

A model priced at $2 input and $10 output per million tokens costs $40 for 1,000 calls of 10,000 input and 2,000 output tokens each. Different tokenizers can turn the same text into different token counts.

02 / WHAT TO CHECK

Reasoning, caching and long context change the bill

Include billed reasoning tokens in output. Cache reads may be cheaper while cache writes can cost more. Some providers charge higher rates for large prompts or premium service tiers. Our estimate excludes those adjustments; check the linked source before committing a budget.

03 / WHAT TO CHECK

Direct API and routed prices are different offers

We use verified direct prices for selected OpenAI and Anthropic models and explicitly labeled OpenRouter catalog rates for other models. A catalog entry can route across endpoints with different prices and limits. Each profile shows the exact catalog identifier and any observed price difference.

04 / WHAT TO CHECK

Use Practical Value as a preference, not a financial forecast

The default score combines 70% task-weighted capability with 30% affordability rank within the snapshot. Move the slider to express your tradeoff. It does not predict task success, retries, hosting charges or actual monthly savings.

YOUR WORKLOAD. YOUR PRIORITIES.

Find your model sweet spot.

How our ratings work ↗

Make the numbers yours

Standard text API estimate. Include reasoning tokens in output.

A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.

THE MARKET MAP

Capability meets cost.

Sweet-spot zoneEfficient frontierTap a dot to inspect
SWEET SPOTCAPABILITY FIRST020406080100$0.1$1$10$100$1,000Agent Fit / 100 ↑← Lower cost · USD · log scale321SWEET SPOT020406080100$0.1$1$10$100$1,000Agent Fit / 100 ↑← Lower cost · USD · log scale321
78.9Agent Fit$5.4your workload
YOUR FAST SHORTLIST

Closest to the sweet spot

How we choose ↗
  1. 1
    DeepSeek V4.1 Flash ↗78.9 Agent Fit · $5.4
  2. 2
    Muse Spark 1.3 ↗79.3 Agent Fit · $21
  3. 3
    DeepSeek V4 Flash Vision Exp ↗75.1 Agent Fit · $3.45

Shaded zone: 75.2+ points and $21.9 or less. Numbers mark efficient models in or nearest this zone. This shortlist changes with the metric, workload and filters.

THE EVIDENCE, SIDE BY SIDE

AI model comparison table

58 of 58 models · select up to 4 to comparePrices in USD · tap ? for instant explanations
01
Agent Fit 78.9
Value 80.6
Your cost $5.4
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
81.1
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.3 / $1.2
Price observation
OpenRouter catalog
02
DeepSeek V4 Flash Vision Exp ↗

DeepSeek · experimental · source default

Agent Fit 75.1
Value 79.9
Your cost $3.45
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
76.8
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.2156 / $0.6468
Price observation
OpenRouter catalog
03
DeepSeek V4 Flash 0731 ↗

DeepSeek · source default

Agent Fit 71.1
Value 78.1
Your cost $2.64
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
74.2
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.0077 / $1.28
Price observation
OpenRouter catalog
04
GPT-6 Luna ↗

OpenAI · max

Agent Fit 67.7
Value 76.8
Your cost $2
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
72.0
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.1 / $0.5
Price observation
Direct API
05
GPT-5.6 Luna ↗

OpenAI · source ID max; display xhigh

Agent Fit 70.4
Value 76.0
Your cost $4.4
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
73.6
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.2 / $1.2
Price observation
Direct API
06
GLM-5.3 Flash ↗

Z.ai · source default

Agent Fit 67.2
Value 76.0
Your cost $2.5
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
71.6
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.15 / $0.5
Price observation
OpenRouter catalog
07
DeepSeek V4 Pro ↗

DeepSeek · source default

Agent Fit 67.1
Value 74.8
Your cost $2.92
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
71.6
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.2088 / $0.4176
Price observation
OpenRouter catalog
08
GPT-5.4 nano ↗

OpenAI · xhigh

Agent Fit 67.7
Value 73.6
Your cost $4.5
400K contextTool calling
Benchmarks & API pricing +
LiveBench
69.6
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.2 / $1.25
Price observation
Direct API
09
Qwen3.8 27B ↗

Alibaba · source default

Agent Fit 73.5
Value 73.3
Your cost $10.2
1.00M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
75.3
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.42 / $3
Price observation
OpenRouter catalog
10
DeepSeek V4 Flash ↗

DeepSeek · source default

Agent Fit 61.6
Value 73.1
Your cost $0.392
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
65.5
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.028 / $0.056
Price observation
OpenRouter catalog
11
DeepSeek V4 Pro 0813 ↗

DeepSeek · source default

Agent Fit 73.3
Value 72.6
Your cost $10.56
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
77.4
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.66 / $1.98
Price observation
OpenRouter catalog
12
Gemini 3.7 Flash ↗

Google · high

Agent Fit 76.1
Value 72.3
Your cost $15
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
78.8
Evidence Consensus
50.0
Input / output per 1M tokens
$0.75 / $3.75
Price observation
OpenRouter catalog
AI model benchmarks, aggregate ratings, API pricing and context windows. Snapshot 2026-10-03.
CompareInput / outputUSD / 1M tokens
01
DeepSeek V4.1 Flash ↗DeepSeek · maxOpen weights
78.980.6—81.1—$5.41.05M$0.3 / $1.2OpenRouter catalog
02
DeepSeek V4 Flash Vision Exp ↗DeepSeek · experimental · source defaultOpen weightsPreview
75.179.9—76.8—$3.451.05M$0.2156 / $0.6468OpenRouter catalog
03
DeepSeek V4 Flash 0731 ↗DeepSeek · source defaultOpen weights
71.178.1—74.2—$2.641.05M$0.0077 / $1.28OpenRouter catalog
04
GPT-6 Luna ↗OpenAI · max
67.776.8—72.0—$21.05M$0.1 / $0.5Direct API
05
GPT-5.6 Luna ↗OpenAI · source ID max; display xhighConfiguration discrepancy
70.476.0—73.6—$4.41.05M$0.2 / $1.2Direct API
06
GLM-5.3 Flash ↗Z.ai · source defaultOpen weights
67.276.0—71.6—$2.51.05M$0.15 / $0.5OpenRouter catalog
07
DeepSeek V4 Pro ↗DeepSeek · source defaultOpen weights
67.174.8—71.6—$2.921.05M$0.2088 / $0.4176OpenRouter catalog
08
GPT-5.4 nano ↗OpenAI · xhigh
67.773.6—69.6—$4.5400K$0.2 / $1.25Direct API
09
Qwen3.8 27B ↗Alibaba · source defaultOpen weights
73.573.3—75.3—$10.21.00M$0.42 / $3OpenRouter catalog
10
DeepSeek V4 Flash ↗DeepSeek · source defaultOpen weights
61.673.1—65.5—$0.3921.05M$0.028 / $0.056OpenRouter catalog
11
DeepSeek V4 Pro 0813 ↗DeepSeek · source defaultOpen weights
73.372.6—77.4—$10.561.05M$0.66 / $1.98OpenRouter catalog
12
Gemini 3.7 Flash ↗Google · high
76.172.350.02 sources78.81,488±5 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
13
Muse Spark 1.3 ↗Meta · xhigh
79.371.3—81.6—$211.05M$1.25 / $4.25OpenRouter catalog
14
Kimi K2.6 ↗Moonshot AI · thinkingOpen weights
66.970.8—70.5—$8262K$0.4342 / $1.83OpenRouter catalog
15
Gemini 3.8 Flash ↗Google · high
73.370.455.62 sources75.81,495±5 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
16
MiniMax M3 ↗MiniMax · source default
63.169.5—67.3—$5.41.05M$0.3 / $1.2OpenRouter catalog
17
Qwen3.6 Plus ↗Alibaba · source default
63.969.3—68.9—$7.151.00M$0.325 / $1.95OpenRouter catalog
18
Muse Spark 1.2 ↗Meta · xhigh
76.369.255.62 sources78.01,494±9$211.05M$1.25 / $4.25OpenRouter catalog
19
GLM-5.2 ↗Z.ai · source defaultOpen weights
68.568.7—73.2—$12.081.05M$0.41 / $3.99OpenRouter catalog
20
Gemini 3.6 Flash ↗Google · high
70.368.313.92 sources73.61,483±4$151.05M$0.75 / $3.75OpenRouter catalog
21
Muse Spark 1.1 ↗Meta · source ID xhigh; display highConfiguration discrepancy
74.067.6—75.3—$211.05M$1.25 / $4.25OpenRouter catalog
22
Nemotron 3 Ultra ↗NVIDIA · 550B A55B · source defaultOpen weights
63.867.5—67.4—$9.4262K$0.5 / $2.2OpenRouter catalog
23
GLM-5.3 ↗Z.ai · source defaultOpen weights
73.766.3—76.1—$22.81.05M$1.4 / $4.4OpenRouter catalog
24
Grok 4.6 ↗xAI · source default
75.365.8—78.0—$32500K$2 / $6OpenRouter catalog
25
Kimi K2.7 Code ↗Moonshot AI · source defaultOpen weights
64.865.6—68.4—$13.41262K$0.6712 / $3.35OpenRouter catalog
26
Inkling ↗Thinking Machines · xhighOpen weights
68.965.2—71.9—$17.6524K$0.95 / $4.05OpenRouter catalog
59.565.1—63.9—$81.05M$0.3 / $2.5OpenRouter catalog
28
GPT-6.1 Sol ↗OpenAI · max
77.765.052.82 sources81.61,483±11$401.05M$2 / $10Direct API
29
Grok 4.7 ↗xAI · xhigh
73.764.7—77.4—$32500K$2 / $6OpenRouter catalog
30
Qwen3.6 27B ↗Alibaba · source defaultOpen weights
60.064.4—64.0—$9.6262K$0.32 / $3.2OpenRouter catalog
31
Grok 4.5 ↗xAI · source default
73.164.3—75.8—$32500K$2 / $6OpenRouter catalog
32
Qwen3.7 Max ↗Alibaba · source default
70.463.5—73.1—$23.61.00M$1.48 / $4.43OpenRouter catalog
33
GPT-6 Sol ↗OpenAI · max
74.762.9—79.2—$401.05M$2 / $10Direct API
34
Claude Sonnet 5 ↗Anthropic · xhigh
73.361.9—76.0—$401.00M$2 / $10Direct API
35
GPT-5.4 mini ↗OpenAI · xhigh
62.561.7—66.4—$16.5400K$0.75 / $4.5Direct API
36
Gemini 3.5 Flash ↗Google · high
70.861.613.92 sources74.61,477±4$331.05M$1.5 / $9OpenRouter catalog
37
Kimi K3 ↗Moonshot AI · source defaultOpen weights
77.461.3—79.2—$541.05M$2.7 / $13.5OpenRouter catalog
38
GPT-5.6 Terra ↗OpenAI · source ID max; display xhighConfiguration discrepancy
74.160.8—77.9—$441.05M$2 / $12Direct API
39
Claude Opus 5.5 ↗Anthropic · max
79.460.8—83.2—$801.00M$4 / $20Direct API
40
Claude Sonnet 5.5 ↗Anthropic · max
71.060.3—75.7—$401.00M$2 / $10Direct API
41
Gemini 3.1 Pro Preview ↗Google · highPreview
73.260.3—77.0—$441.05M$2 / $12OpenRouter catalog
42
GPT-5.6 Sol ↗OpenAI · source ID max; display xhighConfiguration discrepancy
77.159.1—81.1—$801.05M$4 / $20Direct API
43
GPT-5.4 ↗OpenAI · xhigh
74.458.6—78.0—$551.05M$2.5 / $15Direct API
44
GPT-5.2 Codex ↗OpenAI · source default
69.956.8—74.0—$45.5400K$1.75 / $14OpenRouter catalog
45
GPT-5.2 ↗OpenAI · high · 2025-12-11
69.856.8—74.6—$45.5400K$1.75 / $14OpenRouter catalog
46
Grok 4.3 ↗xAI · source default
56.056.7—62.2—$17.51.00M$1.25 / $2.5OpenRouter catalog
47
Claude Fable 5.1 ↗Anthropic · max
79.656.394.42 sources83.41,501±6$2001.00M$10 / $50Direct API
48
Claude Opus 5 ↗Anthropic · max
75.756.261.12 sources80.11,489±5$1001.00M$5 / $25Direct API
49
Claude Fable 5 ↗Anthropic · source ID max; display xhighConfiguration discrepancy
79.055.8—83.0—$2001.00M$10 / $50Direct API
50
GPT-6 Astra ↗OpenAI · max
78.655.647.22 sources82.21,477±7$2001.05M$10 / $50Direct API
51
GPT-5.5 ↗OpenAI · xhigh
75.854.7—80.2—$1101.05M$5 / $30Direct API
52
Claude Sonnet 4.6 ↗Anthropic · medium · adaptive thinking
69.454.6—73.0—$601.00M$3 / $15Direct API
53
Claude Opus 4.8 ↗Anthropic · max
73.054.3—76.2—$1001.00M$5 / $25Direct API
54
Claude Opus 4.7 ↗Anthropic · xhigh
72.954.3—76.5—$1001.00M$5 / $25Direct API
55
Claude Opus 4.6 ↗Anthropic · high · adaptive thinking
70.552.655.62 sources74.51,505±3$1001.00M$5 / $25Direct API
56
Claude Opus 4.5 ↗Anthropic · high · 64K thinking
66.750.0—72.6—$100200K$5 / $25Direct API
57
Qwen3.8 Flash Next ↗Alibaba · source defaultOpen weights
76.2——76.2——Not verifiedNot verified
58
Qwen3.8 Max ↗Alibaba · source default · version mapping unconfirmed
77.0——78.5——Not verifiedNot verified

— means missing comparable evidence, never zero. Consensus covers only 10 matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.

LOOK BEYOND ONE NUMBER

Your shortlist, under the microscope.

Change models ↑
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5GPT-6 SolDeepSeek V4.1 Flash
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxGPT-6 SolmaxDeepSeek V4.1 Flashmax
Agent Fit79.474.778.9
Practical Value60.862.980.6
Evidence ConsensusInsufficient evidenceInsufficient evidenceInsufficient evidence
Workload cost$80$40$5.4
Context window1.00M1.05M1.05M
Tool callingSupportedSupportedSupported
Reasoning92.288.786.7
Coding89.381.880.0
Agentic coding71.752.977.3
Mathematics97.196.493.3
Data analysis80.381.279.3
Language86.385.381.2
Instruction following65.768.670.0
Input / output per 1M$4 / $20$2 / $10$0.3 / $1.2
AI Agent Store · Model intelligence · Snapshot 2026-10-03Sources, limitations & corrections ↗