Three weeks. That’s all it took for OpenAI to slice GPT-5.6 Luna’s asking price by 80 percent. Input per million tokens fell from $1.00 to $0.20. Output fell from $6.00 to $1.20. Terra, the mid-tier, fell only 20 percent. Sol, the flagship, did not move. Asymmetric price cuts are always the first map of strategic pressure. The question is not whether OpenAI can compete. The question is where the attack is actually landing.
The GPT-5.6 family was never one model. It is a three-tier architecture: Sol at the top, Terra in the middle, Luna at the bottom. OpenAI officially defines Luna as delivering 85 percent of Sol’s quality. That is a marketing definition, not a public benchmark. Three weeks after launch, Luna’s price dropped from $1/$6 to $0.20/$1.20. That is not a cost improvement announcement. It is a response to a revenue grab. A CNBC survey found Chinese models now process 46 percent of US enterprise token volume on OpenRouter. DeepSeek V4 Pro charges $0.435 input and $0.87 output per million tokens. Luna’s new input price is below DeepSeek. Its output price is not. The asymmetry is deliberate.
Let’s read that pricing table the way I read token contracts during the 2017 ICO blitz: line by line, looking for what the operator doesn’t say. OpenAI held Sol at $5/$30. It cut Terra by only 20 percent. It cut Luna by 80 percent. Why? Because small-model inference is a commodity. The workloads are classification, extraction, summarization, and formatting. They are price-sensitive, low-switching-cost, and high-volume. If a competitor can deliver 85 percent quality for 50 percent less, enterprises will move within a billing cycle. OpenAI is not defending Sol’s turf. It is defending the intake valve.
Read the price asymmetry against DeepSeek again. Input: $0.20 versus $0.435. Output: $1.20 versus $0.87. OpenAI wins the input side and loses the output side. In API economics, input tokens are the entry point: system prompts, batch documents, and retrieval context. Once a developer hard-wires a pipeline to your endpoint, output pricing becomes a tax, not a barrier. That is the same mechanism I flagged in DeFi yield farming during 2020. The yield was always a subsidy with a lock-in target. The deeply discounted input rate is a subsidy with a retention target. Sol’s price stays high because high-intelligence work still has pricing power. Luna’s price drops because low-intelligence work has no moat.
Then there is API Fast. OpenAI charges 2x the standard API rate for up to 2.5x speed. That is a separate profit center for latency-sensitive operations. In a price war, the seller must find pricing power somewhere. Speed is the most obvious moat. Agents, real-time workflows, and high-frequency decision loops need deterministic latency. The 80 percent cut on Luna is the blunt instrument. API Fast is the scalpel. A mature pricing strategy needs both. If OpenAI was truly panicking, the speed premium would be free. It is not. The company is exactly where intelligent defensive pricing sits.
Terra’s small cut is the quiet signal. Anthropic launched Sonnet 5 at a promotional $2/$10, scheduled to rise to $3/$15 after August 31. Terra sits at $2/$12. OpenAI’s output is more expensive than the promo and cheaper than the post-promo price. That is a defensive bracket against a direct competitor, not a structural cost breakthrough. If OpenAI had discovered an order-of-magnitude cost advantage on the core model, every tier would drop. Instead, the cuts are layered exactly where the competitive pressure is highest.
Now let’s talk about the blind spot. The market reads 46 percent Chinese-model penetration and calls it a geopolitical setback. I read it as an unsegmented number. Not all token volume is equal. Text classification on an inexpensive Chinese model is not equivalent to legal reasoning on Sol. The 46 percent is almost certainly dominated by low-value, high-volume tasks. Those tasks have the lowest switching costs and the highest price elasticity. They are also the easiest to win back with a price cut. If OpenAI wanted to fight for the enterprise crown, it would not cut Luna by 80 percent. It would cut Terra’s output to under $10 and undercut Anthropic’s promotion. It didn’t. Terra still costs $12 output, while Sonnet 5 runs a promo at $10. That gap tells you the fight is industry-wide, but selective. OpenAI is giving up margin where it must, holding it where it can.
Based on my audit experience across 500 token contracts and four market cycles, I have learned to trust pricing sheets over press releases. This is a transactional industry. The press release says OpenAI is democratizing intelligence. The pricing sheet says OpenAI is under attack in the lower-middle segment and is paying an 80 percent discount to keep the pipeline full. Luna’s 85 percent quality rating is controlled by OpenAI. The evaluation methodology is not public. Actual capability variance could be smaller or larger than advertised. But the strategic direction is clear: a single training run is being distilled into multiple price bands through quantization, sparse routing, or smaller derivative models. In crypto terms, that is the difference between building one settlement chain and launching ten rollups that inherit its security. It is not scale. It is fragmentation.
The model card is static. The pricing table is not. Static pricing dies first in a commodity market. Static analysis never explains a fivefold volume threshold. For the 80 percent cut to be economically neutral, Luna’s volume must grow roughly fivefold. If it does, OpenAI’s cost model is structurally sound and the family architecture is working. If it does not, this is straight revenue destruction. I will be watching API throughput, gross margin disclosures, and enterprise migration data, not demo videos. The terminal has already printed the thesis. The next chart to break will tell you whether this was a strategic retreat or a temporary subsidy.

