Skip to main content
Under the hood

The AI behind every stock report

Every reviewed stock runs through a deep, multi-model AI pipeline — dozens of independent analyses, several models, and an adversarial cross-examination before a verdict. This is a plain-English look at the model landscape, how AI usage is measured, and how we keep it reliable. You don't manage any of it — no keys, no quotas, no setup. We run it all.

The full model lineup

Verified 2026-09-06

No single model is best at everything, so each step of the analysis is routed to the model that fits it. Here's the current lineup from each maker with its public list pricing, broken out model by model — newest first.

Anthropic Claude USD / 1M tokens · list price Verified 2026-09-06
Claude Fable 5.1 Most capable · reasoning Released Sep 2026
Input $10
Output $50
Cache write $12.50
Cache read $0.25
Batch (in/out) $5 / $25
Context 1M
Max out 128k
Claude Opus 5 Flagship Opus Released Jul 2026
Input $5
Output $25
Cache write $6.25
Cache read $0.50
Batch (in/out) $2.50 / $12.50
Context 1M
Max out 128k
Claude Fable 5 Legacy · Fable tier Released Jun 2026
Input $10
Output $50
Cache write $12.50
Cache read $1
Batch (in/out) $5 / $25
Context 1M
Max out 128k
Claude Sonnet 5 Speed + intelligence Released Jun 2026
Input $2
Output $10
Cache write $2.50
Cache read $0.20
Batch (in/out) $1 / $5
Context 1M
Max out 128k
Claude Opus 4.8 Prior Opus Released May 2026
Input $5
Output $25
Cache write $6.25
Cache read $0.50
Batch (in/out) $2.50 / $12.50
Context 1M
Max out 128k
Claude Opus 4.7 Prior Opus Released Apr 2026
Input $5
Output $25
Cache write $6.25
Cache read $0.50
Batch (in/out) $2.50 / $12.50
Context 1M
Max out 128k
Claude Opus 4.6 Legacy Released Feb 2026
Input $5
Output $25
Cache write $6.25
Cache read $0.50
Batch (in/out) $2.50 / $12.50
Context 1M
Max out 128k
Claude Sonnet 4.6 Legacy Released Feb 2026
Input $3
Output $15
Cache write $3.75
Cache read $0.30
Batch (in/out) $1.50 / $7.50
Context 1M
Max out 128k
Claude Opus 4.5 Legacy Released Nov 2025
Input $5
Output $25
Cache write $6.25
Cache read $0.50
Batch (in/out) $2.50 / $12.50
Context 200k
Max out 64k
Claude Haiku 4.5 Fastest · near-frontier Released Oct 2025
Input $1
Output $5
Cache write $1.25
Cache read $0.10
Batch (in/out) $0.50 / $2.50
Context 200k
Max out 64k
Claude Sonnet 4.5 Prior Sonnet Released Sep 2025
Input $3
Output $15
Cache write $3.75
Cache read $0.30
Batch (in/out) $1.50 / $7.50
Context 200k
Max out 64k

Cache 5-min write = 1.25× input · 1-hr write = 2× · read = 0.1× (0.025× on Fable 5.1). Batch = 50% off in+out. Every 4.6+ model ships the full 1M context at standard pricing. Fast mode (Opus 5 / 4.8 only) = $10 / $50. Sonnet 5’s $2 / $10 launch price became the standard price — the Sep 1 increase was cancelled. Retirement floors: Sonnet 4.5 not before 2026-09-29, Haiku 4.5 not before 2026-10-15, Opus 4.5 not before 2026-11-24.

OpenAI GPT USD / 1M tokens · list price Verified 2026-09-06
GPT-6 Astra Flagship · trusted access Released Sep 2026
Input $10
Output $50
Cached in $1
Long ctx (in/out) $20 / $75
Batch (in/out) $5 / $25
GPT-5.6 Sol Top 5.6 · promo price Released Jul 2026
Input $4
Output $20
Cached in $0.40
Long ctx (in/out) $8 / $30
Batch (in/out) $2 / $10
GPT-5.6 Terra Mid 5.6 Released Jul 2026
Input $2
Output $12
Cached in $0.20
Long ctx (in/out) $4 / $18
Batch (in/out) $1 / $6
GPT-5.6 Luna Small 5.6 Released Jul 2026
Input $0.20
Output $1.20
Cached in $0.02
Long ctx (in/out) $0.40 / $1.80
Batch (in/out) $0.10 / $0.60
GPT-5.5-pro Max compute Released Apr 2026
Input $30
Output $180
Cached in n/a
Batch (in/out) $15 / $90
Context 1.05M
Max out 128k
GPT-5.5 Prior flagship Released Apr 2026
Input $5
Output $30
Cached in $0.50
Long ctx (in/out) $10 / $45
Batch (in/out) $2.50 / $15
Context 1.05M
Max out 128k
GPT-5.4-pro Pro reasoning Released Mar 2026
Input $30
Output $180
Cached in n/a
Batch (in/out) $15 / $90
Context 1.05M
Max out 128k
GPT-5.4 Reasoning Released Mar 2026
Input $2.50
Output $15
Cached in $0.25
Long ctx (in/out) $5 / $22.50
Batch (in/out) $1.25 / $7.50
Context 1.05M
Max out 128k
GPT-5.4-mini Small reasoning Released Mar 2026
Input $0.75
Output $4.50
Cached in $0.075
Batch (in/out) $0.375 / $2.25
Context 400k
Max out 128k
GPT-5.4-nano Cheapest reasoning Released Mar 2026
Input $0.20
Output $1.25
Cached in $0.02
Batch (in/out) $0.10 / $0.625
Context 400k
Max out 128k
GPT-5.2 Legacy Released Dec 2025
Input $1.75
Output $14
Cached in $0.175
Batch (in/out) $0.875 / $7
GPT-5.1 Legacy Released Nov 2025
Input $1.25
Output $10
Cached in $0.125
Batch (in/out) $0.625 / $5

One discounted “cached input” rate · no separate cache-write. Batch = 50% off. Prompts above 272k tokens are billed at the model’s long-context rate on every token. GPT-5.6 Sol carries promotional pricing through Nov 21, 2026. Fast mode = 2× standard. Regional (data-residency) endpoints add 10% on models released after Mar 5, 2026.

Google Gemini USD / 1M tokens · list price Verified 2026-09-06
Gemini 3.8 Flash Latest Flash · promo Released Sep 2026
Input $0.75
Output $3.75
Cached in $0.075
Cache store $1/hr
Batch (in/out) $0.375 / $1.875
From 2027 $1.50 / $7.50
Gemini 3.7 Flash Flash · promo Released Aug 2026
Input $0.75
Output $3.75
Cached in $0.075
Cache store $1/hr
Batch (in/out) $0.375 / $1.875
From 2027 $1.50 / $7.50
Gemini 3.6 Flash Flash · promo Released Jul 2026
Input $0.75
Output $3.75
Cached in $0.075
Cache store $1/hr
Batch (in/out) $0.375 / $1.875
From 2027 $1.50 / $7.50
Gemini 3.5 Flash-Lite Cheap high-volume Released Jun 2026
Input $0.30
Output $2.50
Batch (in/out) $0.15 / $1.25
Gemini 3.5 Flash Prior Flash Released May 2026
Input $1.50
Output $9
Cached in $0.15
Cache store $1/hr
Batch (in/out) $0.75 / $4.50
Gemini 3.1 Flash-Lite Cheap high-volume Released May 2026
Input $0.25
Output $1.50
Cached in $0.025
Cache store $1/hr
Batch (in/out) $0.125 / $0.75
Gemma 4 Open weights Released Apr 2026
Input free
Output free
Gemini 3.1 Pro Flagship · preview Released Feb 2026
Input $2 / $4
Output $12 / $18
Cached in $0.20 / $0.40
Cache store $4.50/hr
Batch (in/out) $1 / $6 · $2 / $9
Tier ≤200k / >200k
Gemini 2.5 Flash-Lite Cheapest 2.5 Released Jul 2025
Input $0.10
Output $0.40
Cached in $0.01
Cache store $1/hr
Batch (in/out) $0.05 / $0.20
Gemini 2.5 Pro Prior flagship Released Jun 2025
Input $1.25 / $2.50
Output $10 / $15
Cached in $0.125 / $0.25
Cache store $4.50/hr
Batch (in/out) $0.625 / $5 · $1.25 / $7.50
Tier ≤200k / >200k
Gemini 2.5 Flash Prior-gen fast Released Jun 2025
Input $0.30
Output $2.50
Cached in $0.03
Cache store $1/hr
Batch (in/out) $0.15 / $1.25

The 3.6 / 3.7 / 3.8 Flash line is half price through Dec 31, 2026 ($0.75 / $3.75), doubling to $1.50 / $7.50 on Jan 1, 2027. Pro tiers are context-tiered (≤200k / >200k). Cache adds a per-1M-tokens/hour storage charge. Context windows live on the Models page.

xAI Grok USD / 1M tokens · list price Verified 2026-09-06
Grok 4.6 Flagship Released Aug 2026
Input $2
Output $6
Cached in $0.50
≥200k (in/out) $4 / $12
Context 500k
Grok 4.5 Prior flagship Released Jul 2026
Input $2
Output $6
Cached in $0.30
≥200k (in/out) $4 / $12
Context 500k
Grok Build 0.1 Code agent · beta Released May 2026
Input $1
Output $2
Cached in $0.20
≥200k (in/out) $2 / $4
Context 256k
Grok 4.3 Prior flagship Released Apr 2026
Input $1.25
Output $2.50
Cached in $0.20
≥200k (in/out) $2.50 / $5
Context 1M
Grok 4.20 · reasoning Reasoning Released Mar 2026
Input $1.25
Output $2.50
Cached in $0.20
≥200k (in/out) $2.50 / $5
Context 1M
Grok 4.20 · non-reasoning Non-reasoning Released Mar 2026
Input $1.25
Output $2.50
Cached in $0.20
≥200k (in/out) $2.50 / $5
Context 1M
Grok 4.20 · multi-agent Multi-agent Released Mar 2026
Input $1.25
Output $2.50
Cached in $0.20
≥200k (in/out) $2.50 / $5
Context 1M

Cached-input read published; no separate cache-write, and xAI publishes no batch price. Prompts of 200k tokens or more are billed at double the base rate on every token.

DeepSeek V4 USD / 1M tokens · list price Verified 2026-09-06
DeepSeek V4 Pro High-capability · 1.6T Released Apr 2026
Input (miss) $0.66
Input (hit) $0.022
Output $1.98
Peak 2×
Context 1M
DeepSeek V4 Flash Default · 284B/13B Released Apr 2026
Input (miss) $0.22
Input (hit) $0.007
Output $0.66
Peak 2×
Context 1M

Separate cache-HIT vs cache-MISS input rate (≈97% cache discount). Since Aug 16, 2026 pricing is also time-of-day tiered: the figures shown are OFF-PEAK; peak hours (01:00–04:00 and 06:00–10:00 UTC) are double. Read the official pricing page before any billing math.

Mistral La Plateforme USD / 1M tokens · list price Verified 2026-09-06
GLM 5.2 Third-party · agentic Released Aug 2026
Input $1.40
Output $4.40
Cached in $0.14
Batch (in/out) $0.70 / $2.20
Mistral Medium 3.5 Flagship · multimodal Released Apr 2026
Input $1.50
Output $7.50
Batch (in/out) $0.75 / $3.75
Context 256k
Mistral Small 4 Lightweight multimodal Released Mar 2026
Input $0.15
Output $0.60
Batch (in/out) $0.075 / $0.30
Context 256k
Mistral Large 3 Open-weight flagship Released Dec 2025
Input $0.50
Output $1.50
Batch (in/out) $0.25 / $0.75
Context 256k
Ministral 3 (3/8/14B) Edge / small Released Dec 2025
Input $0.10–0.20
Output $0.10–0.20
Context 128k
Codestral Code completion Released Aug 2025
Input $0.30
Output $0.90
Batch (in/out) $0.15 / $0.45
Context 128k

Batch = 50% off; cached input tokens up to 90% off (per-model cache rates not all published). Context from model cards. Embeddings, OCR ($4/1k pages) and voice are priced separately. GLM 5.2 is a third-party model served on the platform.

List prices are each provider's own published per-million-token rates (public info) — shown for context, not Delvantic's internal cost. Verified September 2026; providers change these often.

How usage is measured: tokens

Verified 2026-09-06

LLMs don't bill per question — they bill per token, the small chunks of text a model reads and writes. Roughly one token ≈ 4 characters, or about ¾ of a word.

~1,300
tokens in a 1,000-word document
input + output
both are counted — the data sent in and the analysis written back
tens of thousands
of tokens flow through a single full stock report
Why it matters: token usage is what makes deep analysis affordable. We feed each model only what it needs, cache repeated context (the discounted cache rates above), and reuse shared research across tickers — so a full report stays a fraction of what a brute-force "one giant prompt" approach would burn.

Rate limits & reliability

Verified 2026-09-06

AI providers cap how fast anyone can call their models — requests per minute, tokens per minute. Hit a cap and the API briefly says "slow down" (an HTTP 429). On most platforms that's your problem to handle. On Delvantic, it's ours.

We queue & pace the work

Runs are scheduled and throttled under the hood so the pipeline stays inside provider limits instead of slamming into them.

We retry automatically

If a provider throttles a step, it backs off and retries on its own — a momentary limit doesn't sink a report.

We pick the right model

Routing routine passes to faster models keeps throughput high and leaves headroom for the heavy reasoning steps.

Nothing for you to configure. No API keys, no tiers, no quotas to upgrade. Delvantic owns the provider relationships and the infrastructure — you just read the research.