The Economics of AI Labs | Fintech Inside #114
How AI labs like OpenAI and Anthropic actually make money - capex, GPU depreciation, inference margins, and why the business model looks more like a bakery than a SaaS company.
Hi Insiders, I’m Osborne, an investor in early-stage startups.
Welcome to another edition of Fintech Inside. Fintech Inside provides nuance and insight on the big trends shaping financial services.
Every week we get another frontier model launch. A smarter Claude, a faster GPT, a cheaper open-weight release that hits 95% of the state of the art at a fraction of the cost. Most of the discussion around these launches revolves around benchmarks: who’s smarter, who’s cheaper, who’s winning.
I got curious about something underneath all of it: how does an AI lab’s business model actually work? How is this being financed, who’s carrying the risk, and how does demand and supply actually play out for something this capital-intensive? Is this a genuinely new kind of business, or the same capital-intensive playbook other industries have run before, just with GPUs instead of factories?
This year gave me a good real-world case to dig into it with: Anthropic and OpenAI running opposite capital strategies at the same time, with real dollar figures attached, while a genuine accounting fight over GPU depreciation broke into the mainstream.
In this edition, I try to work through what that economics actually looks like, and what it means if you’re building or investing in fintech and AI-native businesses in India and Southeast Asia, several layers removed from the labs themselves.
Thank you for supporting me and sticking around. Enjoy another satisfying week in fintech.
Considering angel investing or investing in funds? I get a bunch of founders reaching out to me for investors. I’d be happy to put you in touch. Send me a DM here.
🤔 One Big Thought
The Economics of AI Labs
The Bakery That Has to Buy a New, Expensive Oven Every Year
As an outsider to the AI labs/frontier model space, I’ve been curious about the financial economics of these businesses. I used to think of them as a SaaS business, but, boy! was I far from the truth! Let’s try and unpack the AI lab business model, with an analogy. Let’s say you’re the owner of Frontier Flour Co., a bakery in Bandra, Mumbai.
As the owner of Frontier Flour Co., you have to buy a brand new industrial oven every year, and each oven costs more than the bakery made in revenue the year before, because the bakery two doors down just installed an oven that bakes twice as many loaves per hour for half the gas bill.
You can’t run last year’s oven for its full useful life and stay in business, everyone else selling bread is now doing it faster and cheaper. So you raise VC funds or take out a loan for the new oven, and the oven only pays for itself if it’s running near capacity around the clock, every idle hour is money lost on a very expensive appliance you’re still paying off.
Swap “oven” for “GPU training cluster” and “loaves per hour” for “tokens per second” and you have an AI lab.
A frontier training run costs billions, takes months to prepare and run, and the resulting model is meaningfully behind the frontier within a year or two, the same way last year’s oven gets outpaced by the one next door. The cluster has to keep running, training the next model, serving inference, anything, because idle compute is exactly as expensive as an oven sitting cold in the back of the kitchen.
Then, as a new neighbourhood bakery, you have to keep reinventing the menu to ensure these instagram-friendly GenZ kids keep coming through the door. So you have to keep experimenting with new flavours, new baking styles, new products and more.
Most of what your bakery tries in the kitchen never makes it onto the shelf, recipes that don’t rise right, flavours nobody likes, batches that work in a home kitchen but fall apart at scale. The bakery still pays for all those ingredients and test runs regardless of whether anyone ever eats the result.
AI labs run the same experiment. Most training runs, architecture tweaks, proprietary data purchases and research bets don’t ship, or don’t move the needle, but the lab pays for the compute either way. The recipe that does work has to earn back the cost of every failed batch in the test kitchen, and it has to do it fast, because once a rival bakery tastes it, they can often get close enough to sell it cheaper within months, not years.
Finally there’s how the bakery actually gets paid back. It doesn’t recoup an oven’s cost by selling one giant batch of bread on day one. It recoups it slowly, loaf by loaf, subscription box by subscription box, for as long as customers keep ordering, hopefully for years, before the next oven upgrade is due.
That’s the model AI labs are racing toward too: turning one extremely expensive training run into years of recurring API and subscription revenue, loaf by loaf, before the next training cycle demands a new oven.
Put those three together and you get the actual business: a bakery that has to buy a new oven every twelve to eighteen months, where every oven purchase comes with a kitchen full of recipes that might not work, and where the only way to ever pay any of it off is loaf by loaf, subscription by subscription, before the next oven is due.
If you think about it, an AI lab is three whole businesses in one: a semiconductor business buying the oven, a pharma company betting on the recipe, and a cloud company hoping you subscribe - all running out of the same Frontier Flour co. bakery.
Same industry, opposite playbooks
We don’t have to speculate about capital strategy anymore, because two labs are running a live A/B test on it. OpenAI went aggressive early: raise everything, commit everything, build ahead of demand. Its Stargate joint venture with SoftBank and Oracle targets $500bn in infrastructure by 2029, and Amazon alone has put roughly $63bn behind the company this year, $15bn invested directly into Series C preferred stock plus a commitment to buy up to $35bn more, layered on top of a separate cloud services and product collaboration agreement.
Anthropic ran conservative for longer, then moved fast once demand outran its own forecasts. In about twelve months it assembled a compute portfolio spanning a $200bn, five-year deal with Google and Broadcom for 5GW of capacity, up to another 5GW from Amazon under a $100bn server-rental arrangement, Azure capacity through a Microsoft/Nvidia tie-up, and, tellingly, a deal to lease spare capacity from a competitor, paying xAI roughly $1.25bn a month for compute it couldn’t build fast enough on its own.
The payoff so far: Anthropic’s run-rate revenue crossed $74bn in July 2026, up from about $9bn at the end of 2025 - 8x growth in 7 months! Generational company! This growth is largely attributed to compounding model-capability leaps landing inside a single quarter, and by mid-2026 that run rate had climbed further still to roughly $47bn.
Net dollar retention across its enterprise base is running above 500% annualised, and the number of customers spending $1M+ annually passed 1,000, doubling in under two months. Two labs, two capital strategies, both scrambling to keep pace with demand they underestimated. That’s the story benchmarks don’t show you. The model roadmap is downstream of the financing roadmap.
Anthropic’s CFO, Krishna Rao, described the logic behind these bets in terms that make the earlier analogy almost literal. Anthropic buys compute across three different chip families, Amazon’s Trainium, Google’s TPUs and Nvidia’s GPUs, and treats the whole GPU pool as fungible, shifting capacity between training, internal tooling and customer-facing inference depending on where the return is highest that week.
He’s talked about planning against what he calls a “cone of uncertainty,” a range of demand scenarios stretched over one to two years, and deliberately buying toward the top of that range, because underbuying compute knocks you off the frontier just as surely as overbuying it can sink the business.
On the $75bn+ raised since he joined two years ago, with another $50bn still to land from the Google and Amazon deals, his own framing is worth sitting with: that capital exists to cover the width of the uncertainty range, not to plug losses in a business he describes as already running efficiently.
Whether that framing survives a slower demand environment than the one it was built for is exactly the open question I keep circling back to.
The treadmill problem, priced
Once one model launches, the next one has already begun. There’s almost no pause, and understanding why means separating two costs that most coverage of this industry lumps together.
Training a frontier model is, in substance, a capital expenditure: a lab spends a few hundred million dollars, sometimes low billions, on compute, data and researcher time before the model exists at all, and that spend has to be recovered over the model’s useful life the same way a factory’s construction cost gets recovered over years of output. Once the model actually ships, a second, ongoing cost kicks in: inference, the compute burned every single time a customer uses it. That one doesn’t get paid once. It recurs on every request, for as long as the model stays in service.
This is where the margin story gets more interesting than the headline numbers suggest. It’s common to hear that mature AI products can run 50-70%, even up to 90%, gross margins once inference is optimised, and on a per-million token basis for a well-tuned, mature model that’s roughly right, providers have reported inference margins in the 70-80% range at the unit level.
But that’s not what shows up in the labs’ actual numbers. Anthropic’s blended gross margin ran closer to 40% in 2025, about ten points below its own internal forecast, driven by inference costs that came in 23% over budget. OpenAI’s reported gross margin has been closer to a third.
The gap between “70-80% on a mature, optimised token” and “40% company-wide” is the treadmill, expressed in numbers: the strong unit economics of an individual model get diluted the moment you’re discounting for enterprise deals, serving a mix of old and new models running at different efficiency levels, and burning cash training the next generation before the current one has even fully matured.
Which is exactly why, once a model is released, the priority shifts hard toward maximising how much of it gets used. A shorter payback period on that few-hundred-million-to-billion-dollar training bet means proving the capex was worth it fast, and the surest way to do that is to get as many tokens as possible flowing through the model, across as many customers and use cases as possible, before the next generation needs funding.
That’s the real reason behind the pace of product launches: a new agent, coding tool or workflow feature seemingly every other week isn’t primarily a product strategy decision, it’s about giving the oven more reasons to keep baking. Agents in particular matter here because they’re structurally heavier token consumers than a single chat reply: one agentic task can call the model dozens or hundreds of times over minutes or hours, chaining tool calls and reasoning steps together, which is a far more reliable way to keep utilisation high than waiting on one-off queries.
That’s a large part of why the entire industry has pivoted so hard toward agentic workflows in the last year, it isn’t just that agents are more useful, it’s that they burn tokens at a rate chat never could.
The risk sitting underneath all of this is on the demand side. If there’s slower demand of those loaves of bread, prepared with the new fancy, expensive oven, then the loaves will be sitting on the shelf, when that bread has a shelf life.
The treadmill only works if usage keeps climbing fast enough to justify the next round of capex before it’s spent. The scale involved stopped being abstract this year: the five largest US cloud and AI infrastructure companies, Microsoft, Alphabet, Amazon, Meta and Oracle, are projected to spend $660-690bn on capex in 2026 alone, nearly double 2025 levels, committed largely on the expectation that demand keeps compounding.
If usage growth merely levels off instead of continuing to compound, the payback math on the last training run gets uncomfortable fast, and the case for the next one gets harder to make.
The open-source/open-weights variable
The “cheaper open-weight models start to matter” line above undersells what’s actually happened. On OpenRouter (read Edition #113), the marketplace that routes API traffic across providers, Chinese open-weight models have gone from a rounding error to somewhere between 46% and 60% of tokens routed by US companies over the past year, depending on whose July 2026 count you use, up from roughly 30% US-hosted-model share just twelve months ago. Alibaba’s Qwen overtook Meta’s Llama as the most-downloaded model family on Hugging Face back in September 2025. Z.ai’s GLM-5.2, released in late June, saw its daily token volume grow roughly 27x and its customer count roughly 80x in its first week alone.
The pricing gap explains why. DeepSeek’s V4 Flash tier lists at roughly $0.17 per million output tokens, Kimi K3 is $14 per million output tokens, against $30 for GPT-5.5 and $50 for Claude Opus 5. That’s not “95% of the performance for 5% of the cost” as an abstract framing anymore, it’s a live, working price list. The latest open-weight model by US-based Thinking Machines, founded by Mira Murati, is $1.2 per million output tokens.
And it’s landing squarely on the highest-value workload in this whole piece: coding and agentic tasks made up roughly 11% of OpenRouter usage at the start of 2025 and over half by mid-2026, and Chinese models are disproportionately strong and cheap at exactly that work.
The Commoditisation Paradox
Open models commoditise inference. They don’t necessarily commoditise distribution, enterprise integrations, trust, compliance, workflow software or customer relationships. I’ve written about this in Edition #110
The model may become interchangeable. The platform may not, which is exactly why every frontier lab is moving beyond “the model” toward an integrated operating system for work, not because they suddenly became software companies, but because software has better economics than selling tokens.
Anthropic’s own pricing history is a useful real-world data point here. When it cut prices on its top-tier Opus model at the 4.5 release, usage didn’t just backfill the lost revenue, it multiplied it, the kind of Jevons paradox effect where cheaper access unlocks so much more consumption that total revenue rises even as the per-token price falls.
That’s a very different dynamic from the “price gets commoditised toward zero” story open models are supposed to trigger, and it’s a big part of why pricing across the Haiku, Sonnet and Opus tiers has stayed unusually stable even as the underlying capability keeps climbing.
However, a chunk of the enterprise conversation isn’t really about capability at all, it’s about data residency: a hosted Chinese model processes prompts under Chinese jurisdiction, which rules it out for a lot of regulated or sensitive workloads regardless of price. But none of that changes the demand math. A frontier lab doesn’t need to lose the benchmark war to feel this. It just needs enough price-sensitive, non-regulated volume, to peel off and slow the usage curve the whole treadmill is underwritten against.
Circular financing: who’s actually taking the risk?
This is the part of the story that doesn’t fit neatly into “revenue vs. cost,” and it deserves attention. Increasingly the money doesn’t just flow to labs, it flows through them, in circles. A cloud provider invests equity in a lab. The lab commits to spend that money, and more, buying compute back from the same provider, or from chipmakers who are themselves extending credit to fund the purchase.
Amazon’s arrangement with OpenAI is a clean example of this: an equity investment, a cloud services agreement, and a joint product collaboration, all signed in the same quarter. Compute vendors are doing something similar. One recent deal between OpenAI and an inference-chip provider included the vendor funding a roughly $1bn working-capital loan back to its own customer.
None of this is illegal, or even unusual, vendor financing is a well-worn tool in capital-intensive industries. But it does mean a good chunk of the revenue growth being celebrated on earnings calls is, at least partly, capital that started on the same balance sheet it ends up on. I keep wondering how much of the growth is real external demand versus financing engineering.
The depreciation debate hiding inside AI earnings calls
If you want the single most underrated number in this industry, it isn’t a valuation. It’s a depreciation schedule. When a company buys a GPU, the cash leaves immediately, but the accounting expense gets spread out, typically over four to six years. That gap between cash out the door and expense on the income statement is where a lot of “AI profitability” quietly lives.
Investor Michael Burry has argued hyperscalers are depreciating hardware far too slowly relative to how fast Nvidia’s chip generations actually turn over, closer to two or three years in practice given Nvidia now ships a new architecture annually. His estimate: this understates industry-wide depreciation, and overstates profits, by roughly $176bn between 2026 and 2028.
What makes this credible rather than just a hot take is that the industry can’t agree on the answer itself. In the same year, Amazon shortened the useful life of a chunk of its servers from six years to five and took a $920M accelerated depreciation charge for it, citing the pace of AI hardware change.
Meta moved the opposite direction, extending its schedule to 5.5 years and cutting reported depreciation expense by $2.9bn. Same underlying Nvidia hardware, opposite accounting conclusions, in the same reporting year. That divergence is the treadmill problem translated into a line item. If chips really do become obsolete every two to three years, the reported profitability of the entire sector is partly an assumption every one of these companies is choosing for itself, not a fact.
What happens to yesterday’s bread
There’s a question that follows naturally from all this: when a new model launches, does the old one just get thrown out? I don’t know for sure but I think mostly, no, and this is where the bakery analogy earns its keep again. A new flagship model is almost always trained from scratch rather than incrementally patched, the way a new aircraft is a new aircraft even though it’s built on everything engineers learned from the last one.
But the previous generation doesn’t just go in the bin. It usually gets sold at the day-old rack: pushed down-market as the free or budget tier, since its cost is already sunk and every additional user is close to pure margin. It gets compressed into smaller, cheaper “mini” or “flash” versions that inherit most of the capability at a fraction of the serving cost, the way a bakery might turn yesterday’s loaves into croutons or breadcrumbs rather than wasting them outright.
Sometimes it just keeps selling because it’s now the cheapest way to serve a customer who doesn’t need the newest capability. Only once a model is worse, slower and more expensive to run than what replaced it does it actually get retired. The practical upshot is that a frontier lab isn’t really selling one hero product, it’s running a whole bakery counter of models at once, premium loaf, everyday loaf, day-old rack, priced and positioned differently, all funded by the same original oven.
The question I keep coming back to
I started this piece out of curiosity. I wanted to understand what kind of business an AI lab actually is: how the capex gets financed, why margins look the way they do, why a new agent or product shows up every other week, and whether any of this is genuinely new or just an old capital-intensive playbook wearing a GPU costume. Working through it, my honest answer is: a bit of both, and I’m not fully sure yet.
What I do think is worth carrying forward, especially if you’re building or investing in fintech and AI-native businesses, is that none of us are as far removed from this as it feels. Every founder pricing an API call, every investor underwriting a business built on top of a frontier model, is downstream of a capital structure that’s still being figured out in real time, several layers removed from the labs themselves.
So the thing I’ll actually be watching from here isn’t a benchmark score or a funding round headline. It’s the demand curve: whether usage keeps compounding fast enough to justify the capex being committed today, and what that means for the labs’ business models, and for everyone building on top of them, if it doesn’t. That’s less a conclusion than a question I plan to keep coming back to in future editions.
1-min Feedback: Your feedback helps me improve this newsletter. Click UPVOTE 👍🏽 or DOWNVOTE 👎🏽
🎵 Song on Loop
Bored of all this tech talk? here’s a song for you: This week, it’s System of a Down’s “Chop Suey” (Youtube / Spotify). I think SOAD is performing across EU at the moment and people are making videos of them head-banging on their way to the show and transitioning to head-baning in the concert. Chills.
✨ Call Outs
[Video] China quietly saved the world last month: (this is about the drop on China’s oil imports, how that saved the world from collapsing, why China did it and more)
[Video] How Bridgewater Built an AI Analyst That Does Hours of Expert Research in Minutes (this video is a presentation by Bridgewater’s analyst and tech team showing off their agentic analyst platform)
[Article] Human Capital in Venture Capital: Evidence From 100,000 Venture Capitalists (fascinating study of 100,000 VC professionals, showing that investment success is extremely concentrated).
👋🏾 That’s All Folks
If you’ve made it this far - thanks! As always, you can always reach me via DM at osborne.vc/dm. I’d genuinely appreciate any and all feedback. If you liked what you read, please consider sharing or subscribing.
See you in the next edition.












