Here’s a number that should reframe how you think about AI costs: OpenAI says the price of intelligence has dropped roughly 99% over the last few years. Not 40%. Not 60%. Ninety-nine percent. That’s the central claim in a new piece from OpenAI CFO Sarah Friar, who lays out what the company calls its full-stack AI strategy — a compounding chain of improvements across silicon, infrastructure, models, and products that, taken together, are supposed to make AI both dramatically more capable and dramatically cheaper to run. The post, titled “The Full Stack Behind Abundant Intelligence,” is part strategy memo, part investor pitch, and part technical roadmap. It’s worth reading carefully, because the argument Friar is making has real consequences for how enterprises budget for AI, how competitors position themselves, and whether the economics of this industry actually work long-term.
Why OpenAI Is Talking About the Stack Now
For most of its public life, OpenAI has communicated through model releases and product launches. A new GPT version drops, the benchmarks look good, the coverage follows. What’s different here is that Friar — a finance executive, not a researcher — is the one making the case. That’s intentional. This is a message aimed squarely at CFOs, procurement teams, and enterprise buyers who are tired of being told AI is transformative and want to know what it actually costs at scale.
The timing matters too. OpenAI is reportedly on track for revenues well north of $10 billion annually, but it’s also burning through compute spend at a rate that makes profitability a genuine question. Laying out a coherent story about cost curves isn’t just good PR — it’s necessary to justify continued investment from partners like Microsoft and to maintain confidence as the company navigates its ongoing restructuring into a for-profit entity.
There’s also competitive pressure. Google has been aggressive about publishing its own infrastructure efficiency numbers. Anthropic has made reliability and safety its cost-of-entry narrative. Meta’s open-weight Llama models have pushed inference costs toward the floor for teams willing to self-host. OpenAI needs a story about why its closed, full-stack approach wins anyway. This post is that story.
What the Full-Stack Argument Actually Says
Friar’s core claim is that OpenAI doesn’t just build models — it optimizes across every layer of the AI delivery chain simultaneously. When improvements compound across all those layers, the result is better than any single breakthrough could deliver alone. Let’s break that down layer by layer.
The Chip and Compute Layer
OpenAI has been investing heavily in custom silicon through its partnership on the Stargate Project, the massive infrastructure joint venture with SoftBank, Oracle, and others. The goal is to reduce dependence on third-party GPU supply — primarily Nvidia — and build compute capacity that’s optimized specifically for OpenAI’s training and inference workloads.
This matters because off-the-shelf Nvidia H100s and B200s are expensive and constrained. A chip designed around the specific mathematical operations that transformer-based models use most can, in theory, deliver far better performance per dollar. Apple did this with the Neural Engine in its M-series chips. Google has done it with TPUs. OpenAI wants the same advantage at data center scale.
The Infrastructure and Training Layer
Friar points to significant improvements in how OpenAI trains models — better parallelism, more efficient data pipelines, smarter use of mixed-precision arithmetic. These aren’t glamorous, but they’re multiplicative. A 20% improvement in training efficiency stacked on a 20% improvement in chip utilization stacked on a 20% reduction in data center overhead doesn’t add up to 60% savings. It compounds closer to 73%. That’s the math behind the 99% cost drop claim, and it’s how you get from GPT-4’s eye-watering inference costs to whatever GPT-5’s successors will eventually run at.
The Model Layer
OpenAI has leaned hard into a tiered model strategy — smaller, faster, cheaper models for routine tasks; larger reasoning models for complex ones. The o-series reasoning models like o3 and o4-mini sit at one end; lighter models like GPT-4o mini sit at the other. The idea is to route queries intelligently so you’re not burning expensive compute on tasks that don’t need it.
This is actually where I think OpenAI has been most clever. Competitors benchmark on capability. OpenAI is increasingly benchmarking on cost-per-useful-output, which is a different and arguably more enterprise-friendly metric. If a cheaper model gets 90% of the job done at 10% of the cost, most businesses will take that trade.
The Product Layer
At the top of the stack sits ChatGPT, the API, and increasingly, vertical products like ChatGPT Work and Codex. Friar frames these not just as revenue drivers but as feedback loops — real-world usage at scale surfaces the failure modes that lab benchmarks miss, which drives model improvements, which feeds back into infrastructure choices. It’s a closed loop that pure API players and open-source efforts struggle to replicate.
- Custom chips: Stargate infrastructure optimized for OpenAI workloads, reducing Nvidia dependency
- Training efficiency: Compounding improvements in parallelism, data pipelines, and precision
- Tiered models: Routing tasks to appropriately sized models to minimize cost per output
- Product feedback loops: Billions of real queries shaping model and infrastructure decisions
- Cost trajectory: ~99% reduction in AI inference costs claimed over recent years
Who This Actually Helps — and Who Should Be Skeptical
For enterprise buyers, the message is clear: the cost curve is going down, keep buying. And honestly, that tracks with what we’ve seen. API pricing has dropped substantially. Models that would have cost dollars per query two years ago now cost fractions of a cent. Companies like Asana, which used Codex to clear five years of engineering backlog in two weeks, couldn’t have done that economically at 2023 pricing.
For developers building on the API, the tiered model strategy is actually good news. It means more granular pricing options and the ability to right-size your model choice to your use case. We covered how that plays out practically in our piece on GPT-5.6 coming to Kiro and what developers actually get — the short version is that model diversity gives builders real options, not just a one-size-fits-all offering.
But here’s the thing: the full-stack argument is also a moat argument. Friar is essentially saying that OpenAI’s vertical integration makes it harder to displace, because you’d have to beat them at chips, infrastructure, training, and products simultaneously. That’s a legitimate claim — but it’s also the same claim IBM made about mainframes, and Oracle made about enterprise databases, and it has historically been vulnerable to abstraction layers that let customers ignore the stack entirely.
The Llama-shaped question hanging over all of this: if Meta continues releasing capable open-weight models, the market for hosted AI could bifurcate sharply. Enterprises with the engineering resources to self-host will face a very different cost curve than those who need the managed, full-stack experience. OpenAI’s bet is that most of the market, most of the time, wants the managed experience. That’s probably right today. Whether it’s right in three years is genuinely uncertain.
What This Means for Different Audiences
If you’re an enterprise buyer, use this as negotiating context. OpenAI’s costs are falling, which means contract pricing negotiated 18 months ago may not reflect current economics. Push for consumption-based pricing tied to actual inference costs, not fixed per-seat rates.
If you’re a developer, the tiered model strategy is your friend — but only if you actually instrument your application to route queries appropriately. Defaulting to the most powerful model for everything is leaving money on the table. Tools like OpenAI’s model routing capabilities are worth actually configuring, not just enabling by default. It’s also worth keeping an eye on OpenAI’s data retention policies as you scale, since those have real compliance implications.
If you’re a competitor, Friar’s post is a signal that OpenAI is going to compete on economics as much as capability. Pure capability benchmarking is a game they can win on multiple dimensions now. Differentiation on safety, specialization, or open-source flexibility becomes more important, not less.
What is OpenAI’s full-stack strategy?
It’s OpenAI’s approach to controlling and optimizing every layer of AI delivery — from custom chips and data center infrastructure through model training to end-user products. The argument is that compounding improvements across all layers deliver better economics than optimizing any single layer alone.
How much have AI costs actually dropped?
OpenAI claims roughly 99% cost reduction in inference over the past few years. Independent analysis of published API pricing broadly supports a massive cost decline, though the exact figure varies by model and task type. The trend is real even if the specific number is imprecise.
Does this affect how businesses should buy AI?
Yes, significantly. Falling cost curves mean pricing locked in early may be unfavorable now. Businesses should review their AI contracts and consider whether tiered model access or consumption-based pricing better reflects actual usage patterns.
How does OpenAI’s approach compare to Google or Anthropic?
Google has similar vertical integration through TPUs and Gemini, making it the closest structural competitor. Anthropic focuses on safety and reliability as differentiators but lacks the chip-level integration. Meta’s open-weight strategy is philosophically opposite — pushing costs to zero for self-hosters rather than optimizing a managed stack.
The 99% cost drop is a headline number, but the more interesting story is what happens when that curve keeps going. If AI inference gets cheap enough, the constraint shifts entirely from cost to imagination — what problems are actually worth throwing intelligence at. OpenAI is clearly betting that when that moment arrives, being the full-stack provider means you capture the value, not just deliver the commodity. Whether the market agrees will define the next five years of this industry.