How Enterprises Are Moving From AI Chat to AI That Acts

How Enterprises Are Moving From AI Chat to AI That Acts

Most companies are still using AI like a fancy search engine. Ask a question, get an answer, move on. But a small group of enterprises — the ones OpenAI is now calling “frontier firms” — have figured out something different. They’re not just chatting with AI. They’re letting it do things. And according to OpenAI’s new enterprise research published August 12, the gap between those firms and everyone else is already widening fast.

The Shift From Assistance to Execution in Enterprise AI

Here’s the thing: there’s always been a version of this story. Every major technology wave produces an early-adopter class that gets ahead, then a long tail of laggards who eventually catch up. What’s different about enterprise agentic AI adoption is how steep that curve looks right now.

OpenAI’s research draws a sharp line between companies using AI for assistance — drafting emails, summarizing documents, answering internal questions — and companies using AI for execution. Execution means AI that takes actions. It browses, codes, calls APIs, coordinates across systems, and completes multi-step tasks without a human holding its hand through every click.

That second category is still a minority. But it’s growing fast, and the companies in it are pulling away from competitors in measurable ways: faster software shipping cycles, leaner operations, more output per employee. OpenAI isn’t being subtle about what they think this means. The subtext of this research is a direct pitch: get on the agentic train now, or spend the next two years catching up.

To understand why this moment feels different from previous AI hype cycles, it helps to remember where enterprise AI actually was 18 months ago. ChatGPT Enterprise launched in August 2023, and the early use cases were almost embarrassingly basic — search augmentation, meeting notes, first-draft generation. Useful, sure. But not transformative. The missing piece was always agency: the ability for AI to act in the world, not just respond to prompts.

What Frontier Firms Are Actually Doing With AI

OpenAI’s research identifies a cluster of behaviors that separate high-adoption enterprises from the pack. It’s not about which tools they’re using — most large companies have access to the same models. It’s about how they’ve integrated them.

A few patterns stand out:

  • Agentic workflows over one-shot prompts: Frontier firms have moved beyond single-prompt interactions. They’re running multi-step agent pipelines where AI plans, executes, checks its own work, and loops back when something fails. This requires real infrastructure investment, not just an API key.
  • Codex for software development at scale: OpenAI Codex shows up heavily in the research. Engineering teams at frontier firms aren’t just using it for autocomplete. They’re deploying it for autonomous code review, test generation, bug triage, and in some cases, full feature development on well-scoped tasks. This mirrors what we’ve seen in fintech, where AI agents are now owning entire workflow categories rather than just assisting humans within them.
  • Custom GPTs and internal tooling built on the API: Rather than relying on off-the-shelf ChatGPT, these companies have built purpose-specific AI tools for their teams — trained on internal knowledge bases, connected to proprietary data sources, and embedded directly into existing workflows.
  • Measurement culture: This one’s underrated. Frontier firms are actually tracking what AI does. They have metrics for AI-assisted output, time saved per workflow, error rates. Companies without that measurement infrastructure can’t iterate, and they can’t make the business case for deeper investment.
  • Executive buy-in at the deployment level: Not just “we support AI” statements from the CEO, but actual leadership involvement in deciding which workflows to automate and which to leave alone. OpenAI’s own CFO has written about how this looks from the inside — it’s a management problem as much as a technology problem.

The Codex Factor

Software development is where the execution gap shows up most clearly. Engineering teams that have fully integrated Codex into their CI/CD pipelines aren’t just writing code faster — they’re changing what it’s feasible to build at all. When AI can handle the 60% of engineering work that’s routine and repetitive, human engineers can focus on architecture decisions, product judgment, and the genuinely hard problems.

That’s not a minor efficiency gain. That’s a different kind of engineering org.

OpenAI’s research suggests frontier firms are already operating this way. Their engineering velocity metrics look different from companies still treating AI as an optional co-pilot that individual developers can choose to use or ignore.

Who’s Using ChatGPT Work and Operator-Level Features

The research also highlights adoption of ChatGPT Work — the product formerly known as ChatGPT for Teams — and deeper operator-level API integrations. Zapier’s use of ChatGPT Work to fix its lead funnel is a good example of what this looks like in practice: not a moonshot, but a specific, measurable business problem solved with targeted AI deployment. That’s exactly the pattern OpenAI’s research describes as characteristic of high-adoption firms.

Operator-level integrations — where companies configure system prompts, control model behavior, and connect ChatGPT to internal data — are apparently much more common among frontier firms than OpenAI expected at this stage. The implication is that companies willing to do the integration work are getting significantly better results than those using AI in its default consumer-facing form.

The Gap Is Real, and It’s Growing

Let me be direct about what this research is really saying: there’s a stratification happening in enterprise AI adoption that’s going to be very hard to close later. The firms that figured out agentic workflows in 2024 and early 2025 have 12-18 months of institutional knowledge about what works. They’ve made the mistakes, trained their teams, built the measurement systems. That’s not something a competitor can replicate just by buying enterprise licenses.

This is actually a common pattern in enterprise software — the same thing happened with cloud adoption, with data infrastructure, with DevOps practices. The leading companies compound their advantage because the knowledge is organizational, not just technical.

For OpenAI, publishing this research is clearly strategic. They want laggard enterprises to feel the urgency. The message is: your competitors are doing this. Here’s proof. Call us.

What About the Competition?

It’s impossible to read this research without thinking about where Google and Anthropic fit. Gemini’s scale numbers are impressive, but consumer reach and enterprise depth are different games. Google has the enterprise relationships through Workspace and GCP, but OpenAI has the mindshare in the developer and technical communities that matter most for agentic deployment.

Anthropic is the interesting wildcard. Claude’s reputation for reliability and instruction-following has made it a legitimate enterprise contender, and some of the workflow automation use cases that OpenAI describes in this research are areas where Claude competes directly. I wouldn’t be surprised if Anthropic publishes something similar in the next quarter — this kind of enterprise credibility research is too valuable to cede to OpenAI without a response.

The model capabilities are close enough across the frontier labs that deployment infrastructure and enterprise relationships are becoming the real differentiators. OpenAI knows this, which is why this research exists.

What This Means for Enterprise Leaders Right Now

If you’re running an engineering, operations, or product team at a mid-to-large company, here’s the practical read on OpenAI’s findings:

  • Pilot programs aren’t enough anymore. Companies with isolated AI experiments aren’t seeing the gains. You need workflows, not experiments.
  • The agentic transition requires infrastructure investment. Multi-step agent pipelines don’t just happen — they need engineering time, data access, and security review.
  • Measurement is the unlock. If you can’t measure what AI is doing, you can’t improve it and you can’t justify deeper investment to finance.
  • Hire or develop people who can bridge AI and domain expertise. The scarcest resource right now isn’t model access — it’s people who understand both the technology and the business domain well enough to design effective workflows.
  • The gap closes slowly. If your competitors are already operating agentic workflows and you’re still in pilot mode, expect 18+ months to close that gap even with aggressive investment.

How Does This Apply to Smaller Teams?

OpenAI’s research focuses on enterprise-scale deployments, but the principles aren’t exclusively for Fortune 500 companies. Smaller technical teams can move faster on agentic workflows precisely because they have less organizational friction. The barrier isn’t scale — it’s willingness to invest engineering time in proper integration rather than defaulting to the chat interface.

The companies that treat AI as infrastructure rather than a feature are the ones seeing compound returns. That mindset is available to teams of any size.

OpenAI’s research lands at a moment when the enterprise AI market is genuinely sorting itself out — early movers are cementing advantages, and the cost of waiting is rising with each quarter. Whether this research changes behavior at laggard firms or just confirms what frontier firms already know, it’s a useful map of where enterprise AI actually is in August 2025, not where the press releases say it should be. The next version of this research, probably 12 months from now, will tell us whether the gap closed or widened further — and my bet is on further.

Frequently Asked Questions

What does OpenAI mean by “frontier firms” in enterprise AI?

OpenAI uses the term to describe companies that have moved beyond basic AI assistance into agentic, multi-step AI deployments with real workflow integration, measurement systems, and executive alignment. It’s less about company size and more about depth of adoption — how thoroughly AI is embedded into actual work processes rather than used as an optional add-on.

What is enterprise agentic AI, and how is it different from regular ChatGPT?

Agentic AI refers to AI systems that can take sequences of actions autonomously — browsing the web, writing and running code, calling external APIs, coordinating across tools — rather than just generating a single text response. Regular ChatGPT usage is typically one prompt, one response. Agentic deployment means the AI is doing work over multiple steps with minimal human intervention at each stage.

Is Codex still relevant for enterprise software development in 2025?

According to OpenAI’s research, yes — significantly so. Frontier firms are using Codex not just for code completion but for autonomous test generation, bug triage, and scoped feature development. The key is integration into existing CI/CD pipelines rather than using it as a standalone tool. Teams that have done that integration work report meaningful changes in engineering velocity.

How do I know if my company is behind on enterprise AI adoption?

A few honest signals: if your AI usage is primarily the chat interface without custom system prompts or data integrations, if you don’t have metrics tracking AI-assisted output, or if AI decisions are still made project-by-project rather than as part of a deliberate strategy, you’re probably in the laggard category by OpenAI’s framework. That’s fixable, but it requires treating AI adoption as an operational priority rather than an IT experiment.