Seven providers are serving Kimi K3 on OpenRouter, and most charge $3 and $15 a million tokens. That is what Moonshot charges for the same model, whose weights it published as a free 1.56 terabyte download. The licence stays free until a model-as-a-service business passes $20 million in revenue, at which point Moonshot wants a separate agreement. Serving a 2.8 trillion parameter model costs what it costs, whoever owns the weights.
OpenAI moved on Thursday, and only at the bottom of its range. It cut GPT-5.6 Luna by 80%, to 20 cents and $1.20 a million tokens, took 20% off the mid tier, and left the flagship where it was. The discount lands only where a free model is already competing. The company said its strategy is to make each generation do more work at a lower cost.
Serving is where the money goes, and it is getting more expensive. Amazon raised its 2026 capital budget to $220 billion from $200 billion and named rising memory prices as the reason. Its cloud backlog, the work sold and not yet delivered, is $496 billion, and it says the higher budget still will not cover the demand it already has.
All seven models in the token basket and all four neocloud GPU ranges came back unchanged from the 27 July snapshot; on the spot marketplace the H100 ceiling fell 28% to $3.07 an hour, the same rate as the H200 floor.

