NextFin News - Frontier firms — the top 10% of enterprise AI users — now generate 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January. The gap has more than tripled in six months, and it is not being driven by better models. The firms pulling ahead are the ones that have stopped treating AI as an assistant and started treating workflows as operating capability.
OpenAI's September 1 analysis of AI-native companies, paired with its Enterprise Signals data released in August, makes the mechanism explicit: the competitive distance in enterprise AI is no longer measured in model access, but in how much work an organization has converted into teachable, repeatable, measurable agent workflows. Three companies — accounting-agent builder Basis, revenue-platform Clay, and search-infrastructure Exa Labs — illustrate the pattern, and the underlying data shows the gap widening across industries faster than most leaders expect.
The Frontier Gap Is a Workflow Gap, Not a Model Gap
The headline number is the 8.3x multiple. But the mechanism behind it matters more. OpenAI ranks enterprise customers monthly by output tokens per active user, defining frontier firms as the top 10% and typical firms as those in the 45th to 55th percentiles. As of June 2026, that frontier cohort produced 8.3 times the output of the middle of the pack — a threefold increase over the 2.6x spread recorded in January. The gap appears across industries and company sizes, which means intensive AI use is not limited to technology companies.
What changed between January and June was not the availability of frontier models. All enterprise customers have access to the same model tier. What changed is how the work is structured. Agentic AI — defined in OpenAI's data as Codex tokens — accounted for 64% of combined Codex and ChatGPT output tokens among enterprise customers as of June. Agentic workflows generate more output because they carry out longer, multi-step tasks, so the figure reflects both how often Codex is used and how much output those tasks produce. The shift from asking to delegating explains the token math: a chatbot conversation produces one answer, while an agent workflow produces code, documents, pull requests, test runs, and follow-up actions — each of them output tokens. The widening gap is therefore a proxy for something specific: frontier firms are handing agents longer, multi-step tasks with tools attached, while typical firms remain in question-and-answer mode.
The spread is also uneven across the organization, and that unevenness is the story. Since February, the number of weekly active enterprise Codex users has grown 108 times in legal, 41 times in sales, 41 times in recruiting, and 26 times in marketing, compared with 5 times among engineers. The fastest growth is in functions that had almost no agentic usage in February. Legal, sales, and recruiting are catching up to engineering — the future of agentic work is as much a contracts review as a code review.
Frontier firms also use the capabilities that make agents effective. Among weekly active users, 21% at frontier firms use plugins and 19% use skills, compared with 9% and 3% at typical firms. OpenAI's own internal usage points to the ceiling: within OpenAI, every department now uses Codex as its primary AI tool, and Codex accounts for 99.8% of weekly output tokens generated inside the company. The implication is that typical firms are not just behind today; they are behind the part of the stack where the learning compounds.
The gap is also uneven across industries. The largest token-usage gap is in information and technology at 11.7 times, while the smallest is in manufacturing at 5.3 times. Token usage among typical firms is similar across industries and has grown only modestly over the past year, which suggests many organizations are still using simple chat assistants and have meaningful room to deepen adoption.
Three Patterns: Teachable Skills, Persistent Context, Tested Execution
OpenAI's case studies of Basis, Clay, and Exa Labs translate the aggregate data into operational terms. Each company started with a specific, stable job and converted it into a workflow with a clear trigger, known steps, access to the right tools, and a definition of "done." The progression, in OpenAI's framing, runs from teaching an agent a stable process, to giving it persistent context as work changes, to letting it carry opportunities into tested action.
Teach an agent a stable process, give it persistent context as work changes, then let it carry opportunities into tested action.
Basis turns onboarding into a reusable skill. At Basis, which builds AI agents for accounting firms, first-day onboarding now takes 30 minutes instead of two hours. On day one, employees receive access to Codex and a company-specific onboarding skill — a reusable set of instructions and resources for that workflow. Codex welcomes them, introduces key company concepts, and uses their computer to complete integration setup in the background. When recurring questions or exceptions appear, HR updates the skill before the next cohort. The process no longer depends on one person's availability, yet the team can step in for exceptions. Onboarding becomes consistent, repeatable, and easier to improve — and new employees gain an immediate model for working with AI.
Clay gives scattered work a persistent home base. Clay, which builds a self-learning revenue engine for go-to-market teams, faced the familiar sales problem: deal context scattered across CRM records, email, Slack, calls, presentations, text messages, and conversations with internal teams and customer champions. A go-to-market engineer built a persistent workspace and a dedicated subagent for every account. Each subagent reviews primary sources and updates its deal folder overnight; every morning, a coordinating agent turns those updates across all her accounts into a short list of priority moves — answer a lingering customer question, fill a gap in the buying committee, or give a prospect a reason to re-engage. The workflow saves her roughly an hour of inbox triage each night, according to Clay, and the supporting evidence stays close to each recommendation so sellers can inspect primary sources before acting.
Exa carries an opportunity into tested action. Exa Labs, which builds web search infrastructure for AI agents, wants its search API available wherever developers could use it — a goal the team calls "Exa everywhere." Pursuing it once required developer relations and account teams to monitor repositories, identify promising integrations, gather context, and coordinate across systems. The team turned that sequence into a defined workflow for Codex, with clear priorities, access to the necessary sources, and human review before anything ships. Codex now monitors for high-priority integration opportunities, gathers context, creates pull requests, runs tests, and prepares weekly updates using sources such as Slack and Notion. When appropriate, it can also draft the next step, including an initial announcement, for the team to review. People still decide which opportunities matter, which commitments Exa should make, and how external relationships should be managed — but tests and review points make the agent's work visible before it ships.
The common thread is that each company made improvement part of the workflow. Onboarding exceptions reveal where a skill needs refinement; new account activity and seller validation keep deal context current; tests and human review sharpen the boundaries for future execution. The division of labor becomes clearer through use, and responsibility expands as the workflow proves itself.
Why This Is Structural, Not Cyclical
The central judgment for enterprise leaders is whether the frontier gap is a cyclical adoption lag that will revert as laggards catch up, or a structural regime shift that will persist. The evidence points to structural.
A cyclical gap would narrow on its own as tools diffuse and typical firms simply adopt what frontier firms already do. But the mechanism here is compounding, not diffusing. A workflow encoded as a skill improves with each run: exceptions feed back into the instructions, evidence accumulates in persistent folders, and test results redefine the boundaries of safe execution. That feedback loop is owned by the organization that built it, not by the model vendor. Two firms can license the same model and end up with different operating capabilities because one has converted its work into repeatable workflows and the other has not.
The adoption data supports the regime-shift read. A 2026 survey of enterprise AI agent adoption found that three in four enterprises are now actively using or testing agents, and nearly half — 47% — use them for data management, the single most common use case. Human-in-the-loop approval gates are the most common management approach at 38%, and 30% of leaders say they see the most potential for agents in automating routine workflows. Over four in five enterprise leaders — 84% — say it is likely or certain their organization will increase AI agent investments over the next 12 months. These are operating disciplines, not tool settings. Once an organization has built them, they do not revert when the model cycle turns.
There is also a scale effect visible in the usage data. The frontier advantage is beginning to compound: OpenAI's B2B Signals report found frontier firms use 3.5 times as much intelligence per worker as typical firms, up from 2 times a year earlier, and message volume explains only 36% of that gap — the majority comes from deeper usage. Microsoft's 2026 Work Trend Index identifies a group it calls Frontier Professionals — the most advanced AI users, representing 16% of surveyed AI users — of whom 80% use agents for multi-step workflows and routinely rethink where agents can augment or automate. That 16% is small, but it is the same shape as OpenAI's top-10% frontier firms: a minority doing structurally different work, not just more of the same work.
The Counter-Thesis: Is the Frontier Gap Just Token Inflation?
The strongest challenge to the structural read is that output tokens are an imperfect measure of value. OpenAI itself warns that a short response can be highly valuable, while a long one may add little.
A short response can be highly valuable, while a long one may add little.
On this view, the 8.3x gap could reflect frontier firms generating more low-quality output, burning more compute without producing more economic value — a token-inflation bubble that will deflate once leaders demand outcome metrics over activity metrics.
That objection has force, and it is the right discipline to impose. But it does not overturn the structural conclusion for two reasons. First, the case studies are anchored in time and outcome, not token volume: Basis cut onboarding from two hours to 30 minutes; Clay recovered roughly an hour per night per seller; Exa reduced handoffs across research, engineering, and communications and ships tested artifacts. Tokens are the proxy OpenAI uses to measure depth of use, but the operating gains are reported in hours and cycle time.
Second, the gap is widest where the work is most verifiable. The largest industry token gap is in information and technology at 11.7 times; the smallest is in manufacturing at 5.3 times. Code, tests, and integrations produce immediate, binary feedback — a test passes or it does not. If frontier firms were merely inflating output, the gap would be largest in the least verifiable functions, not the most. Instead, the functions with the tightest feedback loops show the widest spreads, which is consistent with learning compounding where results can be measured.
The falsifying signal is specific: if, over the next two quarters, frontier firms' token multiples continue to rise while their reported cycle-time and outcome metrics stagnate or deteriorate, the workflow-capability thesis is wrong and the gap is token inflation. Leaders should watch for that divergence — rising activity with flat results — before scaling agent budgets.
What This Means for Enterprise Leaders
The practical translation is that AI strategy should be workflow strategy. The guidance OpenAI lays out for leaders — give employees room to test consequential workflows, measure success, and turn the strongest experiments into repeatable practice — reduces to a single operating principle: convert work into teachable, persistent, tested processes, then let responsibility expand as each workflow proves itself.
Short term, the winners will be organizations that identify stable, high-frequency workflows with clear definitions of done and encode them as skills with human-in-the-loop checkpoints. The losers will be organizations that buy model access and leave work unstructured, assuming adoption will follow. Medium term, the advantage shifts to firms that give agents persistent context — the deal folder, the onboarding cohort, the integration backlog — because context is where institutional knowledge accumulates. Long term, the regime belongs to firms whose workflows improve themselves: exceptions refine the skill, tests redefine the boundary, and the division of labor between human and agent becomes an asset rather than a cost center.
The base case is that the frontier gap continues to widen through 2027 as agentic work spreads beyond engineering into legal, sales, recruiting, and finance. The upside case is that workflow-encoded firms begin licensing their internal capabilities as products, turning operating capability into revenue. The downside case is that outcome metrics fail to keep pace with token growth, triggering a budget correction that resets spending toward narrower, verified use cases.
The signal to watch is the ratio of outcome metrics to token volume inside an organization's own deployments: if cycle time, error rates, and revenue-per-workflow improve faster than token spend, the workflow thesis is confirmed; if tokens grow while outcomes flatten, the correction is underway.
The market is not paying for AI access anymore. It is paying for the work an organization has taught its agents to do.
Explore more exclusive insights at nextfin.ai.
