Customer support assistant
100k messages / month
$113–1,900/mo
Open-weights model via API
Frontier model via API
Details
SourcesOpenRouter model pricing
100k messages / month
$113–1,900/mo
Open-weights model via API
Frontier model via API
SourcesOpenRouter model pricing
10k documents / month
$98–158/mo
Batch OCR API + mid-tier extraction
Standard OCR API + mid-tier extraction
sustained serving, 1 GPU
$204–3,285/mo
Spot marketplace RTX 5090
Dedicated neocloud H200
List price, $ per M output tokens
34×
SourcesOpenRouter model pricing
USD per GPU-hour, observed offers
Spot offers observed on the open marketplace; committed pricing differs.
SourcesVast.ai marketplace pricing· Raw offers API (Vast.ai)
USD per GPU-hour, observed offers
Public on-demand list prices; committed pricing differs.
SourcesLambda pricing· RunPod pricing· Nebius pricing
Customer support assistant (100k messages / month). RAG over company docs. Per message: ~2,000 input tokens (system prompt + retrieved chunks + short history), ~300 output tokens. API pricing only; excludes engineering and vector-store hosting.
Document intake pipeline (10k documents / month). OCR, field extraction, and classification of business documents (invoices, claims, contracts). Average 3 pages per document. OCR priced per page, extraction/classification on a mid-tier model at ~1,500 input / 200 output tokens per document.
Fine-tuned 8B model, self-hosted (sustained serving, 1 GPU). Open 8B model fine-tuned quarterly, served 24/7 on a single GPU. Low end: spot marketplace GPU. High end: dedicated neocloud instance. Excludes engineering and fine-tune compute (amortized separately).
GPU and workload bands are ranges across observed offers within each tier. GPU rows merge board variants (NVL, SXM, PCIE) into a single class; where only one offer was observed, the row shows a single price. We do not publish named provider comparisons; model token prices are public list prices. Build-cost bands: € under 25k · €€ 25–100k · €€€ over 100k.
Sources are linked on each tile · captured Jul 20, 2026.

Moonshot priced its largest model at Anthropic's Sonnet level, and our own Qwen quote rose 18% in a week while every neocloud list price held flat. When the machine is scarcer than the weights, open models stop being the floor.
Jul 13, 2026
AWS's reserved Blackwell sits far above marketplace spot, and this week the Vast.ai B200 asking range rose while last-generation Hopper eased. But hourly rent is not cost per token, and that gap is the real decision for a builder.
Jul 6, 2026
Claude Sonnet 5's intro pricing and Meituan's free 1.6T coder pour the applied-AI cost floor from both sides. Underneath, agentic spend controls and air-gapped open models are quietly moving where the money lands.
Jun 29, 2026
OpenAI shipped its cheapest tier ever and made it almost impossible to buy. When access is the binding constraint, the number on the pricing page is theatre.
Jun 22, 2026
Most companies don't have a data problem. They have a knowledge problem. This is the story of how a business turns scattered files, records, and signals into a system that can think and act.
May 13, 2026
Most teams reach for Notion or Jira by reflex. The structure they're looking for has to come first.
Apr 7, 2026
In 1955, four scientists wrote down seven problems. Seventy years later, the biggest companies on Earth are still trying to answer them.
Mar 13, 2026