Under the hood
The AI behind every stock report
Every reviewed stock runs through a deep, multi-model AI pipeline — dozens of independent
analyses, several models, and an adversarial cross-examination before a verdict. This is a
plain-English look at the model landscape, how AI usage is measured, and how we keep it reliable.
You don't manage any of it — no keys, no quotas, no setup. We run it all.
The full model lineup
Verified 2026-09-06
No single model is best at everything, so each step of the analysis is routed to the model that
fits it. Here's the current lineup from each maker with its public list pricing, broken out
model by model — newest first.
Anthropic Claude
USD / 1M tokens · list price Verified 2026-09-06
Claude Fable 5.1
Most capable · reasoning
Released Sep 2026
| Input |
$10 |
| Output |
$50 |
| Cache write |
$12.50 |
| Cache read |
$0.25 |
| Batch (in/out) |
$5 / $25 |
| Context |
1M |
| Max out |
128k |
Claude Opus 5
Flagship Opus
Released Jul 2026
| Input |
$5 |
| Output |
$25 |
| Cache write |
$6.25 |
| Cache read |
$0.50 |
| Batch (in/out) |
$2.50 / $12.50 |
| Context |
1M |
| Max out |
128k |
Claude Fable 5
Legacy · Fable tier
Released Jun 2026
| Input |
$10 |
| Output |
$50 |
| Cache write |
$12.50 |
| Cache read |
$1 |
| Batch (in/out) |
$5 / $25 |
| Context |
1M |
| Max out |
128k |
Claude Sonnet 5
Speed + intelligence
Released Jun 2026
| Input |
$2 |
| Output |
$10 |
| Cache write |
$2.50 |
| Cache read |
$0.20 |
| Batch (in/out) |
$1 / $5 |
| Context |
1M |
| Max out |
128k |
Claude Opus 4.8
Prior Opus
Released May 2026
| Input |
$5 |
| Output |
$25 |
| Cache write |
$6.25 |
| Cache read |
$0.50 |
| Batch (in/out) |
$2.50 / $12.50 |
| Context |
1M |
| Max out |
128k |
Claude Opus 4.7
Prior Opus
Released Apr 2026
| Input |
$5 |
| Output |
$25 |
| Cache write |
$6.25 |
| Cache read |
$0.50 |
| Batch (in/out) |
$2.50 / $12.50 |
| Context |
1M |
| Max out |
128k |
Claude Opus 4.6
Legacy
Released Feb 2026
| Input |
$5 |
| Output |
$25 |
| Cache write |
$6.25 |
| Cache read |
$0.50 |
| Batch (in/out) |
$2.50 / $12.50 |
| Context |
1M |
| Max out |
128k |
Claude Sonnet 4.6
Legacy
Released Feb 2026
| Input |
$3 |
| Output |
$15 |
| Cache write |
$3.75 |
| Cache read |
$0.30 |
| Batch (in/out) |
$1.50 / $7.50 |
| Context |
1M |
| Max out |
128k |
Claude Opus 4.5
Legacy
Released Nov 2025
| Input |
$5 |
| Output |
$25 |
| Cache write |
$6.25 |
| Cache read |
$0.50 |
| Batch (in/out) |
$2.50 / $12.50 |
| Context |
200k |
| Max out |
64k |
Claude Haiku 4.5
Fastest · near-frontier
Released Oct 2025
| Input |
$1 |
| Output |
$5 |
| Cache write |
$1.25 |
| Cache read |
$0.10 |
| Batch (in/out) |
$0.50 / $2.50 |
| Context |
200k |
| Max out |
64k |
Claude Sonnet 4.5
Prior Sonnet
Released Sep 2025
| Input |
$3 |
| Output |
$15 |
| Cache write |
$3.75 |
| Cache read |
$0.30 |
| Batch (in/out) |
$1.50 / $7.50 |
| Context |
200k |
| Max out |
64k |
Cache 5-min write = 1.25× input · 1-hr write = 2× · read = 0.1× (0.025× on Fable 5.1). Batch = 50% off in+out. Every 4.6+ model ships the full 1M context at standard pricing. Fast mode (Opus 5 / 4.8 only) = $10 / $50. Sonnet 5’s $2 / $10 launch price became the standard price — the Sep 1 increase was cancelled. Retirement floors: Sonnet 4.5 not before 2026-09-29, Haiku 4.5 not before 2026-10-15, Opus 4.5 not before 2026-11-24.
OpenAI GPT
USD / 1M tokens · list price Verified 2026-09-06
GPT-6 Astra
Flagship · trusted access
Released Sep 2026
| Input |
$10 |
| Output |
$50 |
| Cached in |
$1 |
| Long ctx (in/out) |
$20 / $75 |
| Batch (in/out) |
$5 / $25 |
GPT-5.6 Sol
Top 5.6 · promo price
Released Jul 2026
| Input |
$4 |
| Output |
$20 |
| Cached in |
$0.40 |
| Long ctx (in/out) |
$8 / $30 |
| Batch (in/out) |
$2 / $10 |
GPT-5.6 Terra
Mid 5.6
Released Jul 2026
| Input |
$2 |
| Output |
$12 |
| Cached in |
$0.20 |
| Long ctx (in/out) |
$4 / $18 |
| Batch (in/out) |
$1 / $6 |
GPT-5.6 Luna
Small 5.6
Released Jul 2026
| Input |
$0.20 |
| Output |
$1.20 |
| Cached in |
$0.02 |
| Long ctx (in/out) |
$0.40 / $1.80 |
| Batch (in/out) |
$0.10 / $0.60 |
GPT-5.5-pro
Max compute
Released Apr 2026
| Input |
$30 |
| Output |
$180 |
| Cached in |
n/a |
| Batch (in/out) |
$15 / $90 |
| Context |
1.05M |
| Max out |
128k |
GPT-5.5
Prior flagship
Released Apr 2026
| Input |
$5 |
| Output |
$30 |
| Cached in |
$0.50 |
| Long ctx (in/out) |
$10 / $45 |
| Batch (in/out) |
$2.50 / $15 |
| Context |
1.05M |
| Max out |
128k |
GPT-5.4-pro
Pro reasoning
Released Mar 2026
| Input |
$30 |
| Output |
$180 |
| Cached in |
n/a |
| Batch (in/out) |
$15 / $90 |
| Context |
1.05M |
| Max out |
128k |
GPT-5.4
Reasoning
Released Mar 2026
| Input |
$2.50 |
| Output |
$15 |
| Cached in |
$0.25 |
| Long ctx (in/out) |
$5 / $22.50 |
| Batch (in/out) |
$1.25 / $7.50 |
| Context |
1.05M |
| Max out |
128k |
GPT-5.4-mini
Small reasoning
Released Mar 2026
| Input |
$0.75 |
| Output |
$4.50 |
| Cached in |
$0.075 |
| Batch (in/out) |
$0.375 / $2.25 |
| Context |
400k |
| Max out |
128k |
GPT-5.4-nano
Cheapest reasoning
Released Mar 2026
| Input |
$0.20 |
| Output |
$1.25 |
| Cached in |
$0.02 |
| Batch (in/out) |
$0.10 / $0.625 |
| Context |
400k |
| Max out |
128k |
GPT-5.2
Legacy
Released Dec 2025
| Input |
$1.75 |
| Output |
$14 |
| Cached in |
$0.175 |
| Batch (in/out) |
$0.875 / $7 |
GPT-5.1
Legacy
Released Nov 2025
| Input |
$1.25 |
| Output |
$10 |
| Cached in |
$0.125 |
| Batch (in/out) |
$0.625 / $5 |
One discounted “cached input” rate · no separate cache-write. Batch = 50% off. Prompts above 272k tokens are billed at the model’s long-context rate on every token. GPT-5.6 Sol carries promotional pricing through Nov 21, 2026. Fast mode = 2× standard. Regional (data-residency) endpoints add 10% on models released after Mar 5, 2026.
Google Gemini
USD / 1M tokens · list price Verified 2026-09-06
Gemini 3.8 Flash
Latest Flash · promo
Released Sep 2026
| Input |
$0.75 |
| Output |
$3.75 |
| Cached in |
$0.075 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.375 / $1.875 |
| From 2027 |
$1.50 / $7.50 |
Gemini 3.7 Flash
Flash · promo
Released Aug 2026
| Input |
$0.75 |
| Output |
$3.75 |
| Cached in |
$0.075 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.375 / $1.875 |
| From 2027 |
$1.50 / $7.50 |
Gemini 3.6 Flash
Flash · promo
Released Jul 2026
| Input |
$0.75 |
| Output |
$3.75 |
| Cached in |
$0.075 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.375 / $1.875 |
| From 2027 |
$1.50 / $7.50 |
Gemini 3.5 Flash-Lite
Cheap high-volume
Released Jun 2026
| Input |
$0.30 |
| Output |
$2.50 |
| Batch (in/out) |
$0.15 / $1.25 |
Gemini 3.5 Flash
Prior Flash
Released May 2026
| Input |
$1.50 |
| Output |
$9 |
| Cached in |
$0.15 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.75 / $4.50 |
Gemini 3.1 Flash-Lite
Cheap high-volume
Released May 2026
| Input |
$0.25 |
| Output |
$1.50 |
| Cached in |
$0.025 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.125 / $0.75 |
Gemma 4
Open weights
Released Apr 2026
Gemini 3.1 Pro
Flagship · preview
Released Feb 2026
| Input |
$2 / $4 |
| Output |
$12 / $18 |
| Cached in |
$0.20 / $0.40 |
| Cache store |
$4.50/hr |
| Batch (in/out) |
$1 / $6 · $2 / $9 |
| Tier |
≤200k / >200k |
Gemini 2.5 Flash-Lite
Cheapest 2.5
Released Jul 2025
| Input |
$0.10 |
| Output |
$0.40 |
| Cached in |
$0.01 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.05 / $0.20 |
Gemini 2.5 Pro
Prior flagship
Released Jun 2025
| Input |
$1.25 / $2.50 |
| Output |
$10 / $15 |
| Cached in |
$0.125 / $0.25 |
| Cache store |
$4.50/hr |
| Batch (in/out) |
$0.625 / $5 · $1.25 / $7.50 |
| Tier |
≤200k / >200k |
Gemini 2.5 Flash
Prior-gen fast
Released Jun 2025
| Input |
$0.30 |
| Output |
$2.50 |
| Cached in |
$0.03 |
| Cache store |
$1/hr |
| Batch (in/out) |
$0.15 / $1.25 |
The 3.6 / 3.7 / 3.8 Flash line is half price through Dec 31, 2026 ($0.75 / $3.75), doubling to $1.50 / $7.50 on Jan 1, 2027. Pro tiers are context-tiered (≤200k / >200k). Cache adds a per-1M-tokens/hour storage charge. Context windows live on the Models page.
xAI Grok
USD / 1M tokens · list price Verified 2026-09-06
Grok 4.6
Flagship
Released Aug 2026
| Input |
$2 |
| Output |
$6 |
| Cached in |
$0.50 |
| ≥200k (in/out) |
$4 / $12 |
| Context |
500k |
Grok 4.5
Prior flagship
Released Jul 2026
| Input |
$2 |
| Output |
$6 |
| Cached in |
$0.30 |
| ≥200k (in/out) |
$4 / $12 |
| Context |
500k |
Grok Build 0.1
Code agent · beta
Released May 2026
| Input |
$1 |
| Output |
$2 |
| Cached in |
$0.20 |
| ≥200k (in/out) |
$2 / $4 |
| Context |
256k |
Grok 4.3
Prior flagship
Released Apr 2026
| Input |
$1.25 |
| Output |
$2.50 |
| Cached in |
$0.20 |
| ≥200k (in/out) |
$2.50 / $5 |
| Context |
1M |
Grok 4.20 · reasoning
Reasoning
Released Mar 2026
| Input |
$1.25 |
| Output |
$2.50 |
| Cached in |
$0.20 |
| ≥200k (in/out) |
$2.50 / $5 |
| Context |
1M |
Grok 4.20 · non-reasoning
Non-reasoning
Released Mar 2026
| Input |
$1.25 |
| Output |
$2.50 |
| Cached in |
$0.20 |
| ≥200k (in/out) |
$2.50 / $5 |
| Context |
1M |
Grok 4.20 · multi-agent
Multi-agent
Released Mar 2026
| Input |
$1.25 |
| Output |
$2.50 |
| Cached in |
$0.20 |
| ≥200k (in/out) |
$2.50 / $5 |
| Context |
1M |
Cached-input read published; no separate cache-write, and xAI publishes no batch price. Prompts of 200k tokens or more are billed at double the base rate on every token.
DeepSeek V4
USD / 1M tokens · list price Verified 2026-09-06
DeepSeek V4 Pro
High-capability · 1.6T
Released Apr 2026
| Input (miss) |
$0.66 |
| Input (hit) |
$0.022 |
| Output |
$1.98 |
| Peak |
2× |
| Context |
1M |
DeepSeek V4 Flash
Default · 284B/13B
Released Apr 2026
| Input (miss) |
$0.22 |
| Input (hit) |
$0.007 |
| Output |
$0.66 |
| Peak |
2× |
| Context |
1M |
Separate cache-HIT vs cache-MISS input rate (≈97% cache discount). Since Aug 16, 2026 pricing is also time-of-day tiered: the figures shown are OFF-PEAK; peak hours (01:00–04:00 and 06:00–10:00 UTC) are double. Read the official pricing page before any billing math.
Mistral La Plateforme
USD / 1M tokens · list price Verified 2026-09-06
GLM 5.2
Third-party · agentic
Released Aug 2026
| Input |
$1.40 |
| Output |
$4.40 |
| Cached in |
$0.14 |
| Batch (in/out) |
$0.70 / $2.20 |
Mistral Medium 3.5
Flagship · multimodal
Released Apr 2026
| Input |
$1.50 |
| Output |
$7.50 |
| Batch (in/out) |
$0.75 / $3.75 |
| Context |
256k |
Mistral Small 4
Lightweight multimodal
Released Mar 2026
| Input |
$0.15 |
| Output |
$0.60 |
| Batch (in/out) |
$0.075 / $0.30 |
| Context |
256k |
Mistral Large 3
Open-weight flagship
Released Dec 2025
| Input |
$0.50 |
| Output |
$1.50 |
| Batch (in/out) |
$0.25 / $0.75 |
| Context |
256k |
Ministral 3 (3/8/14B)
Edge / small
Released Dec 2025
| Input |
$0.10–0.20 |
| Output |
$0.10–0.20 |
| Context |
128k |
Codestral
Code completion
Released Aug 2025
| Input |
$0.30 |
| Output |
$0.90 |
| Batch (in/out) |
$0.15 / $0.45 |
| Context |
128k |
Batch = 50% off; cached input tokens up to 90% off (per-model cache rates not all published). Context from model cards. Embeddings, OCR ($4/1k pages) and voice are priced separately. GLM 5.2 is a third-party model served on the platform.
List prices are each provider's own published per-million-token rates (public info) — shown for
context, not Delvantic's internal cost. Verified September 2026;
providers change these often.
How usage is measured: tokens
Verified 2026-09-06
LLMs don't bill per question — they bill per token, the small chunks of text a
model reads and writes. Roughly one token ≈ 4 characters, or about ¾ of a word.
~1,300
tokens in a 1,000-word document
input + output
both are counted — the data sent in and the analysis written back
tens of thousands
of tokens flow through a single full stock report
Why it matters: token usage is what makes deep analysis affordable. We feed
each model only what it needs, cache repeated context (the discounted cache rates above), and
reuse shared research across tickers — so a full report stays a fraction of what a brute-force
"one giant prompt" approach would burn.
Rate limits & reliability
Verified 2026-09-06
AI providers cap how fast anyone can call their models — requests per minute, tokens per
minute. Hit a cap and the API briefly says "slow down" (an HTTP 429). On most platforms that's
your problem to handle. On Delvantic, it's ours.
We queue & pace the work
Runs are scheduled and throttled under the hood so the pipeline stays inside provider limits instead of slamming into them.
We retry automatically
If a provider throttles a step, it backs off and retries on its own — a momentary limit doesn't sink a report.
We pick the right model
Routing routine passes to faster models keeps throughput high and leaves headroom for the heavy reasoning steps.
Nothing for you to configure. No API keys, no tiers, no quotas to upgrade.
Delvantic owns the provider relationships and the infrastructure — you just read the research.