Insights
A finished silicon wafer resting on a pale bench, its die grid catching colour.
Rob Bulmahn / Wikimedia Commons · CC BY 2.0
Weekly digest · Week of Jul 20, 2026

The price list froze. The bill kept falling.

Opus 5 shipped at Opus 4.8's exact price, Gemini 3.6 Flash discounted the same job twice over, and AMD and Nvidia launched racks priced in tokens per dollar and tokens per megawatt. Every model in our basket came back unchanged, which is now the least interesting number on the page.

Economics

Five dollars and twenty five dollars. Anthropic shipped Claude Opus 5 on Friday at exactly what Opus 4.8 cost, to the cent, and claims it more than doubles the older model's benchmark score at a lower cost per task. The price card did not move. The bill for a finished job did.

Google's Gemini 3.6 Flash dropped output to $7.50 a million from $9.00 and says it needs 17% fewer output tokens to do the work, discounting the same job twice. AMD launched its Helios rack on Thursday advertising up to 30% more tokens per dollar than the rival cabinet, on its own lab's estimate. Nvidia's Vera Rubin went into production at five clouds, sold on tokens per megawatt. Not one of them is competing on dollars per hour.

Our basket agrees by standing still. All seven models we track came back at last week's price, and no neocloud list rate changed. The one number that moved was the B200 ceiling on the spot marketplace, $10.63 an hour down to $7.48. Which leaves an awkward column on the buyer's comparison sheet. If price per million tokens is the constant now, what is it still measuring?

Metric move

All seven models in the token basket and all four neocloud GPU ranges came back unchanged from the 20 July snapshot; the only move was the marketplace B200 ceiling, $10.63 to $7.48 per GPU-hour.

Spotlight
A loader resting on the tailgate of a truck stacked with sacks of grain.
Thilina Alagiyawanna / Unsplash

Grain that stopped taking the long way round

India's Public Distribution System moves subsidised food grain across states for more than 810 million people, and the routing between depots and destinations was decided without a cost-optimal model behind it.

How it works

A decision-support layer runs operations-research optimisation over state-level grain movement, choosing the sources and routes that shorten the haul, and hands planners a recommended allocation instead of a blank map.

The economicsClaimed

The government puts national deployment at INR 250 crore (2.5 billion rupees) of savings a year and a 35% cut in emissions; the project was one of six finalists for the 2026 Franz Edelman Award.

Transferable pattern

This is network-flow optimisation over a physical estate you already own, and it is the cheapest category of AI a business can buy: no GPUs, no tokens, no model to serve. Anyone shipping the same goods from many depots to many destinations, spare parts, building materials, medical supplies, packaged food, has this problem sitting in their freight ledger. The tell is a routing decision currently made by habit, seniority, or a spreadsheet nobody wants to touch.

Who it’s for
  • Supply chain lead
  • Operations director
  • Logistics planner
  • Finance lead
Players to watch
  • Gurobi commercial solver behind many logistics models
  • Hexaly French solver built for routing problems
  • Google OR-Tools free open-source routing and scheduling solver
  • IBM CPLEX enterprise optimisation platform, long deployment history
  • ORTEC Dutch route and load optimisation specialist
  • Kinaxis supply-chain planning with optimisation built in
  • o9 Solutions planning platform for supply and demand networks
  • Blue Yonder warehouse and transport optimisation suite
  • WFP UN partner on the Anna Chakra build
  • IIT Delhi academic partner that modelled the network
Signals
Jul 23

AMD sells Helios on tokens per dollar

AMD launched the Helios rack, 72 Instinct MI455X GPUs with 18 sixth-generation EPYC CPUs, claiming up to 30% more tokens per dollar than the leading competitive rack on its own lab's estimate, which puts a second credible vendor into rack-scale price competition.

Jul 21

Gemini 3.6 Flash cuts the same bill twice

Google set Gemini 3.6 Flash at $1.50 in and $7.50 out per million tokens, down from $9.00 out on 3.5 Flash, and says it uses 17% fewer output tokens, compounding a list-price cut with a consumption cut on output-heavy work.

Jul 21

Vera Rubin racks reach production at five clouds

NVIDIA says Vera Rubin NVL72 is now running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius, claiming up to 10x more tokens per megawatt and one-tenth the cost per million tokens against GB200 NVL72, which is what would loosen Blackwell rental pricing.

Jul 20

IREN books $2.8bn of GPU cloud contracts

IREN signed $2.8bn of new multi-year cloud contracts at a weighted average term of about four years and raised its end-2026 AI cloud run-rate target from $3.7bn to over $4bn, evidence that GPU capacity is still being sold forward faster than it is built.

Insights