The subsidy is in the subscription
· ai, economics, opinion, llm
There is an argument going around that current AI pricing is propped up by investor money, and that one day the real number arrives and everyone has to decide whether the work was ever worth it. I see some version of it every week, usually with the implication that anyone building on these tools is standing on a trapdoor.
I went looking for the numbers behind it because I wanted to know how exposed I actually am. The argument turns out to be half right, and the half it gets right is not the half people are worried about.
Metered inference looks like a functioning business
Reporting by Ed Zitron on leaked fiscal statements, which the Financial Times subsequently verified, put OpenAI's implied gross margins at 28% in 2024 and 43% in 2025. That is a company whose unit economics are improving, not collapsing, and the 43% figure includes subscription revenue, which as we will get to is the part dragging the number down.
The price trend says the same thing. GPT-4 class capability sold for roughly $30 per million tokens in early 2023. Equivalent capability in 2026 runs under a dollar. Prices do not fall by a factor of thirty across three years in a market where every unit sold loses money.
There is also a useful sanity check from the open-weight side. The analyst scaling01 has argued that if GLM-5.2 can be sold profitably at $4.40 per million output tokens, and closed providers charge several times that, then the closed labs could be running margins north of 90% on inference. That is an inference about someone else's cost structure rather than a disclosure, so hold it loosely. It points the same direction as everything else.
If you pay per token, the provider is most likely making money on you, and the bill you are dreading has been falling for three years.
Flat rate is a different business
The distinction that reframed this for me is stated plainly in the piece linked above. Subscription economics and API economics are not the same business, and only one of them is subsidized.
A subscription loses money because of how intensely a small number of people use it, rather than because tokens are expensive to serve. One number makes it concrete. A fully utilized $200 per month Pro plan would cost somewhere between $3,500 and $14,000 a month at published API rates. That is a gap of roughly 17 to 70 times, depending on assumptions about true serving cost.
Nobody is paying $14,000 of value into a $200 plan and being quietly absorbed forever. That is where the subsidy lives. It belongs to a narrow band of users, the ones running agents continuously against a flat fee.
Where the correction lands
I will put a prediction here so it can be checked later.
The repricing arrives as caps, tiers, throttles, fair-use enforcement, and paid upgrades on flat-rate plans, rather than as a tenfold jump in token costs. Providers will keep metered API pricing competitive because that is a market with open-weight alternatives and visible per-token comparison, and they will tighten the all-you-can-eat tier because that is where the losses live and where switching is hardest.
Some of that has already started. Epoch projects hyperscaler capital expenditure crossing above operating cash flow around the third quarter of this year, which is exactly the condition that makes finance departments interested in metered billing.
If you want to know whether this is happening, watch subscription terms rather than API price pages.
I am the boring customer
Two weeks ago I wrote that I run a handful of subscriptions in the twenty to one hundred dollar range, have never once hit a usage limit, and add metered API spend with a cap when a project needs it. I said it then to argue that token spend is a meaningless flex.
Running the economics changes what that fact means. My provider is making money on me, and on almost everybody, because the losses concentrate at the extreme end of the distribution.
The person genuinely exposed here runs four agents in parallel from breakfast onward on a flat plan. That person is currently getting a remarkable deal and is going to notice when it ends. They are also, in my experience, not the person posting about their enormous token spend, because that person is on metered API and is paying every cent of it.
Portability is the hedge
The useful preparation is making sure a pricing change lands as a line item rather than an existential event. Spending less is beside the point.
Concretely, that means knowing whether your workflow can move. If everything you do is welded to one provider's flat-rate plan and its specific tooling, a terms change lands on you with full force. If your work runs through interfaces you could repoint at a metered API, a different vendor, or a local open-weight model with some degradation you could live with, then the same change is an afternoon of annoyance.
I would test that rather than assume it. Take a workflow you rely on, run it against a cheaper model or a different provider, and find out what actually breaks. Doing that once teaches you more about your exposure than any amount of reading about capital expenditure.
What it costs you is downstream of what you can check
The last piece connects to something I wrote about the returns on all this. Production got roughly ten times cheaper for me and verification got no cheaper, so my throughput is bounded by how fast I can genuinely understand output rather than by how many tokens I can buy.
That bound is also a budget. If your consumption is limited by your capacity to check the results, you were never going to be the customer whose economics do not work. The people at risk from repricing are the ones generating far more than they verify, which was already the expensive habit for entirely different reasons.
The bill arriving might be the thing that finally makes that visible.