Radar

Workload cost index

Customer support assistant

100k messages / month

$113–1,900/mo

Floor

Open-weights model via API

Ceiling

Frontier model via API

Details
Token prices

Frontier vs open spread

List price, $ per M output tokens

34×

  • Frontier$25–30 /M out
  • Mid$7.50–10 /M out
  • Open$0.87–4.43 /M out
Details
GPU compute

Marketplace tier

USD per GPU-hour, observed offers

  • H200$2.80–4.33 /hr
  • B200$5.63–10.63 /hr
  • H100$1.67–4.46 /hr
  • RTX 5090$0.28–0.48 /hr
Details

Spot offers observed on the open marketplace; committed pricing differs.

SourcesVast.ai marketplace pricing· Raw offers API (Vast.ai)

GPU compute

Neocloud tier

USD per GPU-hour, observed offers

  • H200$4.39–4.50 /hr
  • B200$5.89–7.15 /hr
  • H100$2.89–4.29 /hr
  • RTX 5090$0.99 /hr
Details

Public on-demand list prices; committed pricing differs.

SourcesLambda pricing· RunPod pricing· Nebius pricing

Methodology & sources

Customer support assistant (100k messages / month). RAG over company docs. Per message: ~2,000 input tokens (system prompt + retrieved chunks + short history), ~300 output tokens. API pricing only; excludes engineering and vector-store hosting.

Document intake pipeline (10k documents / month). OCR, field extraction, and classification of business documents (invoices, claims, contracts). Average 3 pages per document. OCR priced per page, extraction/classification on a mid-tier model at ~1,500 input / 200 output tokens per document.

Fine-tuned 8B model, self-hosted (sustained serving, 1 GPU). Open 8B model fine-tuned quarterly, served 24/7 on a single GPU. Low end: spot marketplace GPU. High end: dedicated neocloud instance. Excludes engineering and fine-tune compute (amortized separately).

GPU and workload bands are ranges across observed offers within each tier. GPU rows merge board variants (NVL, SXM, PCIE) into a single class; where only one offer was observed, the row shows a single price. We do not publish named provider comparisons; model token prices are public list prices. Build-cost bands: € under 25k · €€ 25–100k · €€€ over 100k.

Sources are linked on each tile · captured Jul 20, 2026.

Last updated: Jul 20, 2026

Insights

Technique
Industry
Market
Practice
Weekly digest

The cheap end of the market raised its prices

Moonshot priced its largest model at Anthropic's Sonnet level, and our own Qwen quote rose 18% in a week while every neocloud list price held flat. When the machine is scarcer than the weights, open models stop being the floor.

PredictionEnergyAI business
Jul 13, 2026