Home / Insights / Making AI pay

Making AI pay

Your AI Programme Is Not Compute Constrained

Olivier GomezOlivier Gomez (OG), 11 min read

I ask two questions in every AI conversation I walk into.

Which model are you using?

What number goes up if this works?

The first question always gets an answer. Immediate, confident, detailed. Vendor, version, context window, the reasoning behind the selection, the bake-off that preceded it, sometimes a slide.

The second question gets a pause.

Then it gets a word that is not a number. Efficiency. Productivity. Innovation. Employee experience. Occasionally it gets something that looks like a metric until you hold it up to the light. Documents processed. Tickets summarized. Prompts run per seat per month.

That gap between the two answers is the whole story of enterprise AI right now. Not model quality. Not compute scarcity. Not talent. The gap.

I have spent twenty-five years in enterprise IT operations and automation delivery, across three global service providers and more than twenty countries. I have watched this exact pattern run through client-server, offshoring, virtualization, cloud, and RPA. The technology changes. The gap does not.

So let me be direct about what I think the actual constraint is, because almost every remediation plan I see is aimed at the wrong one.

Compute is a procurement problem

Start by clearing the thing that gets blamed most often.

You can rent compute by the hour, from at least four credible suppliers, at prices that have fallen every year I have been watching. Capacity is contended at the frontier and there are real queues for the newest silicon. But for the workloads that ninety percent of enterprises are actually running, inference against a hosted model on internal documents and internal processes, compute is a line item with alternatives.

If compute is genuinely your bottleneck, you have a procurement problem. It is a real problem, it is annoying, and it is the easiest one on this page. It has a supplier, a contract, a price, and a resolution date.

Nothing else on this page has those.

The reason compute gets blamed is that it is legible. It appears on an invoice. Executives can act on it. A constraint you can put in a purchase order feels more solvable than a constraint that requires you to admit nobody defined success before spending the money.

What is actually constraining you

You are benchmark constrained.

Not in the academic sense. In the plainest possible sense. There is no number, agreed before the work started, that tells you whether the thing is working.

Without that number, four things happen in sequence, and I have watched all four so many times I now treat them as a single mechanism. I call it the Four Gates.

Promised. Something gets approved on the strength of a narrative. A vendor demo, a peer company’s press release, a board member who read something. The business case contains a range rather than a target, and the range is wide enough to be unfalsifiable.

Piloted. The pilot is built. It works. Of course it works, because “works” was never defined tightly enough to fail. Screenshots circulate. Someone presents it internally. This is the stage where the most political capital gets spent and the least economic value gets created.

Production. A smaller number survive to production, and here the first real cost appears: integration, access management, exception handling, the change management nobody scoped. Costs become concrete while benefits remain adjectival. This is where most programmes quietly plateau. Not cancelled. Plateaued.

Permanent. Fewer still become permanent, meaning the organization would notice and object if you switched it off, and the value shows up somewhere a finance director recognizes.

Value leaks at every gate. The leak is rarely caused by the model being bad. In my experience, it is rarely the model. It is caused by nobody having defined, before the pilot started, the single number that would justify passing through the next gate.

Here is the uncomfortable consequence. If you cannot name the number, the pilot cannot fail. That sounds like safety. It is the opposite. A pilot that cannot fail also cannot succeed. It can only continue. Indefinitely, at a monthly cost, generating screenshots.

That is not a technology outcome. That is a subscription.

Output, outcome, impact

The counter-argument I get is that people are measuring. They have dashboards.

Most of what sits on those dashboards is output. Output is what the system did. Documents processed, calls transcribed, drafts generated, hours of usage. Output metrics have one useful property, which is that they always go up, and one fatal property, which is that they always go up.

Outcome is what changed in the process. Cycle time removed. Rework eliminated. Escalation rate down. Backlog cleared and staying cleared. Outcome metrics can go the wrong way, which is precisely why they are worth measuring.

Impact is what changed in the business. Margin. Retained revenue. Headcount growth avoided against a rising volume. Penalty clauses not triggered. Contract renewals won on service performance.

A dashboard that stops at output is not a benchmark. It is a comfort object.

At IAC, we work on gain-share. We get paid on outcome. That is not a positioning choice; it is a forcing function. You cannot sign a gain-share contract without agreeing on the number first, with a measured baseline, before anyone builds anything. It removes the option of vagueness from both sides. I recommend the discipline, whether or not you ever buy anything from me.

Why nobody names the number

If this were only an analytical error, it would have been fixed years ago. It persists because naming the number is politically expensive and staying vague is politically free.

Consider what a specific benchmark actually does inside an organization. It creates a date on which someone is publicly right or publicly wrong. It converts a sponsor’s enthusiasm into a personal exposure. It gives finance a reason to ask a follow-up question next quarter. And it removes the most valuable feature of an AI initiative as currently practiced, which is that it makes the sponsor look forward-leaning at zero measured risk.

Vagueness is not laziness. It is a rational response to an incentive structure. The people avoiding the number are usually the sharpest in the room, and they understand exactly what they are avoiding.

Two more forces make it worse.

Vendors have no interest in fixing it. A supplier paid per seat, per token, or per license is optimized for usage, and usage is an output metric. Every commercial arrangement in the market rewards the metric that always goes up. When your supplier’s dashboard is the dashboard your board sees, you have outsourced the definition of success to the party being paid.

And the technical team is usually not empowered to set the number anyway. Ask an engineering lead what number should go up, and you will often get a good answer, followed by the observation that they do not own the process, the P&L line, or the baseline data. They can build anything. They cannot decide what counts.

So the number goes unnamed, and everyone involved is behaving reasonably.

This is why I treat benchmark definition as a governance act rather than an analytics exercise. It has to be owned above the delivery team, agreed before the build, and written where it cannot be quietly revised after the fact. Otherwise it will be revised after the fact. Every time.

This is not a new lesson, and the evidence is stronger than you think

Alexander Wissner-Gross made an argument on the TEDxBoston stage recently that I want to borrow, because it is the historical version of everything above.

He claims that the great unlocks in AI were dataset constrained. Not algorithm constrained. Not compute constrained. His list is speech recognition, statistical machine translation, chess, Jeopardy, Go, and general-purpose conversation. In each case, the field did not stall for want of a clever idea. It stalled until somebody built the dataset and defined the benchmark that could tell you whether an idea was working. Once the scoreboard existed and a community formed around beating it, the unlock followed.

He goes further and argues that the field’s early damage came from optimizing for algorithmic elegance instead of a measurable target. His phrase for what the field did instead is religious warfare over whose approach was more elegant or more neuro-inspired. Anyone who has sat through an enterprise architecture review knows that sound. We still make it. We changed the nouns.

I disagree with a great deal of what else he says. He is an avowed accelerationist who expects alien contact, mind uploading, and cures for the top five thousand diseases inside a decade. I have no way to test any of that, and neither does he.

But the dataset argument is not a prediction. It is a description of fifty years that already happened, and it is being demonstrated again right now in the one domain that just got a proper scoreboard.

Mathematics got formal verification. Erdős problems have started falling to AI systems, with the proofs machine-checked in Lean rather than eyeballed. Problem 728 in January of this year, then 397 and 281 within days, then 1196 verified and marked proved.

Take Terence Tao’s caveats with it, because they matter and they are his own. Tao has warned that Erdős problems vary in difficulty by orders of magnitude, that many of the ones falling are long-tail problems nobody seriously attempted, and that a problem open for fifty years is not evidence it resisted fifty years of effort. He has flagged cases where an AI solution turned out to already exist in the literature, and pointed out that unreported failures make success rates uninterpretable.

All of that is fair, and it does not touch the point I am making. The point is what changed to make progress measurable. Lean verification turned “is this proof correct” from an opinion into a check. The moment the scoreboard became machine-readable, progress became visible and fast.

That is the pattern. Not superintelligence. A scoreboard.

Your organization does not have one.

The dataset you cannot buy

If the benchmark is the constraint, the dataset is the asset, and here is where enterprises consistently look in the wrong place.

Everyone wants proprietary data, and almost everyone reaches for the wrong pile. The customer records, the product catalogue, the document repository. That material is valuable, but it is not differentiating, because your competitors have structurally identical versions of it.

The dataset nobody else has is your process exhaust. Ticket histories with resolution paths. Exception logs. Rejected quotes and the reason codes. Escalation trails. Change failures and the post-incident notes. Every decision your organization has made under time pressure, and what happened next.

That is the asset. It is usually sitting in systems nobody has touched in years, in a schema nobody documented, owned by a team that would rather not discuss it.

The model is not the asset. The model is rented, and everyone is renting the same three.

The other half nobody prices

Suppose you do all of this and it works. There is a second exposure, and it is the one I would put on a risk register today.

Wissner-Gross was asked about another AI winter. He said we are overdue but that it will not resemble the historic ones, and his candidate mechanism is capital expenditure. Data centre buildout at a scale that requires frontier lab revenue to justify it. He compared it to the telecom overbuild of the early 2000s and then moved on in about thirty seconds.

I would sit with that comparison much longer, because I lived through the enterprise side of it and it is the most useful thing in the talk.

What the telecom overbuild actually did to enterprises was not remove the technology. Bandwidth got cheaper, not scarcer. What it did was reprice contracts violently, take out suppliers, force consolidations, and expose a set of companies who had built critical operations on a commercial arrangement they did not control and could not exit. The ones who came through were not the ones with the best network. They were the ones who could switch.

Apply that. If a capex correction lands, your model does not stop working. What moves is your unit economics, your contract terms, your rate limits, and your supplier’s willingness to keep serving your workload at last year’s price. The question stops being whose model is best and becomes: what do you own when the rent goes up?

Five locks. Model. Data. Talent. Cost. Exit.

Rent all five, and you are a passenger in your own P&L. You do not have to own all five, and anyone telling you to is selling something. But you need to know which ones you have rented, deliberately, with the switching cost priced and written down. Most organisations have not done that inventory. They will do it under time pressure, which is the most expensive way to do it.

Notice this is the same argument as the benchmark argument in different clothes. A benchmark tells you whether the thing works. An ownership audit tells you whether you keep the value when conditions change. Neither is a technology question. Both get delegated to technologists by default, which is why both get answered badly.

Five moves

None of these require a view on superintelligence.

Name the number. One metric per use case, at outcome level or above, agreed before the build, with a baseline actually measured. If you cannot measure the baseline, that is your first project, and it is more valuable than the pilot.

Set a kill criterion. Below this number by this date, we stop. Write it down and circulate it. A programme with no kill criterion is not a programme.

Mine the exhaust. Find the process data nobody has looked at. That is your differentiated dataset, and it is already paid for.

Price the exit. For every AI dependency, answer in writing what happens if the price triples or your tier is withdrawn, and how many days migration takes. If nobody can answer, that is a risk register entry today.

Separate the sizzle from the sequence. Some of the ten-year predictions may land. Your capital cycle is twelve to thirty-six months. Plan against the cycle you control and treat the rest as upside.

What would change my mind?

I will hold myself to the standard I just applied to everyone else.

If capability keeps improving at the current rate and organizations that never defined a benchmark still capture large, durable margin gains, then I am wrong, model choice mattered more than the scoreboard, and the discipline I am arguing for was overhead.

I have not seen it. I will say so plainly if I do.

Until then, the answer to the second question is the only thing that predicts anything.

Which model are you using is a purchasing decision.

What number goes up if this works is a strategy.

Pick the scoreboard first. The budget comes after.

OG Approved, no BS.

First published in the OG Approved newsletter on 13/08/2026. Read it on Substack or subscribe to get the next one.