All articles
Outcome-Led Consulting

You cannot price what you cannot evidence

AI is removing the visible effort that consulting fees have always been anchored to. Selling outcomes instead does not remove the risk. It moves it onto you. What determines whether you can carry that risk is a third kind of observability, and almost nobody is equipping consultancies to produce it.

Time was never a good measure, only a visible one

For decades consulting has had a convenient proxy for value: the hour multiplied by the rate.

A senior consultant costs more than a junior one. Six months costs more than six weeks. Ten people cost more than three. None of that described the outcome the client actually got, but it was understandable and easy to buy. And clients had no alternative route to the knowledge the consultancy offered.

AI is breaking the relationship. Some work that took four consultants three weeks can increasingly be done by one consultant in three days, and a consultant can get answers from an LLM in minutes (but so can their client). Good news for the cost and speed of delivering consulting, a serious problem for the economics of selling it. Anchor the fee to visible effort and the firm has used AI to cut its own price.

The obvious answer, it seems to much of the market, is to stop selling effort and start selling outcomes. That answer is correct at least in part, but it is considerably harder of course than it sounds.

The map firms are about to price against

A working paper published last month by Erik Brynjolfsson and Georgios Petropoulos, Pricing Consulting Services in the Age of Agentic AI, gives the clearest framing of the problem I have seen. It positions five commercial models - time-based, project-based, access-based, outcome-based and hybrid - against two dimensions: input cost observability, how clearly the client can see what the work costs you to deliver, and outcome observability, how clearly they can verify the result and your contribution to it. Their conclusion is that AI pushes firms towards hybrid and asset-based structures.

The pricing strategy mapA plane with outcome observability increasing along the horizontal axis and input cost observability increasing along the vertical axis. Time-based pricing, hours times rate, occupies the top left. Project-based pricing, a fixed fee on scope, occupies the top right. Access-based pricing, a flat retainer, occupies the bottom left. Outcome-based pricing, value-linked fees, occupies the bottom right. Hybrid pricing, a base plus a variable component, occupies the interior. After Brynjolfsson and Petropoulos, 2026, Figure 1.The pricing strategy mapFive regimes, each optimal in a different information environment.INPUT COST OBSERVABILITY →OUTCOME OBSERVABILITY →Time-basedHours × rateProject-basedFixed fee on scopeHybridBase + variableAccess-basedFlat retainerOutcome-basedValue-linked feesRedrawn after Brynjolfsson and Petropoulos (2026), Figure 1. Hybrid pricing occupies the interior of the map.

It is a map worth having on the wall, and I think it is right. My question is a different one. Before a firm can choose a position on that map, it has to be able to tell which situation it is actually in - and for consulting specifically, I do not think the two dimensions are enough to tell you.

Outcome pricing sounds easier than it is

Take a consultancy that agrees to lift a software company's win rate in return for a share of the additional revenue. Everyone wins: the firm has skin in the game, the client pays in proportion to the result.

Then the win rate rises four points and the CFO asks the only question that matters. Improved compared with what? The pricing change that launched in the same quarter, the competitor who pulled out of two of their largest deals, the three senior hires who joined the sales team in March, or the consultancy's work on qualification discipline? That is not a pricing problem. It is an evidence problem, and it cannot be reconstructed convincingly if it is left as a task at the end, by anyone, however senior.

The capability the map assumes

The Stanford map gives you two things a client can observe: what the work cost, and what the result was. For consulting, I would argue the decisive information sits between them.

We could call it intervention observability. That is, can anyone see what actually happened between the starting position and the eventual result?

The full chain runs:

Baselinediagnosishypothesisrecommendationinterventioncapability movementbusiness outcome

The middle of that chain is not administrative detail. Capability is the only lever a consultancy actually has, whether it builds that capability in the client or supplies it itself for the life of the engagement. You cannot move a client's revenue directly; you change what the organisation is able to do, and the outcome moves as a consequence.

The distinction between building and supplying matters more than it looks, because supplied capability is capability the client is renting. It is one of the reasons outcomes so often revert after a firm leaves, and one of the reasons a firm should know which of the two it is selling. Either way, if you cannot measure capability, you cannot be specific about what has to change in order to produce the outcome, and you are left recommending activity and hoping.

Most firms can evidence the first link weakly and the last link circumstantially. Everything in the middle is typically undocumented, which is why the strongest defendable claim is so often that revenue improved while you were there.

But if the middle of the chain becomes measured and managed, connecting the baseline to the outcome, the conversation changes shape.

  • This was the starting position.
  • These were the capabilities we believed were constraining the outcome.
  • These were the interventions actually undertaken.
  • This is how those capabilities moved.
  • And this is what happened to the agreed business measures over the same period.

That still does not prove causality. Consulting does not happen in a laboratory: markets move, leadership changes, competitors act, clients make their own decisions. But it is a different order of evidence from comparing a KPI at the start and end of a project, and it changes what a firm can responsibly put behind a fee.

There is a difference between this axis and the other two. Input cost observability and outcome observability describe the engagement you happen to be in - you read them off the situation.

Intervention observability describes the firm you have chosen to build. It is not a condition you find yourself in but a capability you either have or do not, and it travels into every engagement regardless of the conditions.

Which is also why the map cannot be used until you have it.

Choosing the right commercial model requires knowing which situation you are in, and knowing that requires a fast, repeatable read of how you would drive the outcome and what is still unknown. The paper tells you which model fits the conditions. It does not tell you which condition applies to a specific client, what your risks are with them, and what you can therefore propose.

The axis you have to build for yourselfThe same plane as the pricing strategy map, with input cost observability increasing up the vertical axis, but with intervention observability increasing along the horizontal axis in place of outcome observability. Top left, only the effort is visible, and the commercial option is time or fixed project. AI augmentation moves firms downward into the bottom left, where neither is visible and there is no defensible basis for price. Evidence moves a firm rightward into the bottom right, where only the change is visible and hybrid or outcome-linked pricing becomes possible. Top right, both are visible and every commercial model remains open.The axis you have to build for yourselfThe same vertical axis. A different question along the bottom.HIGHLOWLOWHIGHINPUT COST OBSERVABILITYHow much of your effort the client can still seeINTERVENTION OBSERVABILITYCan you show what changed, and why?Only the effort is visibleThe client sees the team, the weeksand the deliverables. Nothing showswhat actually moved.COMMERCIAL OPTIONTime or fixed projectBoth are visibleEffort is legible and thechange is proven. Everymodel is open to you.COMMERCIAL OPTIONAll models openWhat AI augmentation does.The effort stops being visible, so the map movesyou down and the fee loses its anchor.Neither is visibleAI removed the effort the feewas anchored to, and nothingwas ever instrumented toreplace it.COMMERCIAL OPTIONNo defensible price basisOnly the change is visibleYou cannot show the hours,but you can show the baseline,the intervention and themovement it produced.COMMERCIAL OPTIONHybrid or outcome-linkedWhat evidence does.Moves you right.Input cost observability is falling for every firm, whether or not itchooses to move. Intervention observability is the only one you can act on.Moving right does not remove the risk. It moves it from the client to you.Evidence is what lets you price the risk you are taking.

Where the paper stops

It would be easy to assume the answer to all this is that AI will fix measurement. As agents embed in client operations and generate continuous telemetry, outcomes become easier to see, so the argument for outcome pricing gets stronger on its own.

The paper is well ahead of that. It separates outcome noise into two parts: how precisely a result can be measured, and how cleanly it can be attributed to the consultant's work rather than to market conditions, other initiatives or the client's own execution. Telemetry compresses the first. It does nothing to the second. Attribution risk therefore sets a floor on how informative an outcome signal can be, a floor no amount of instrumentation lowers, and that floor caps how far any of this can push a firm towards outcome-based pricing.

So I am not disagreeing with them. I am agreeing, and asking the next question.

Their advice is to invest in measurement infrastructure before shifting to outcome-based pricing, and to work out whether measurement or attribution is the binding constraint. Where attribution binds, they prescribe credible causal identification: randomised holdouts, controlled rollouts, clean counterfactuals. And they note, correctly, that many engagements cannot support any of that.

Which is to say: most consulting engagements. You cannot run a randomised holdout on a transformation programme. There is one client, one leadership team, one attempt, and no control group.

So the question the paper leaves open is the one that matters commercially. What does attribution technology look like for the engagements that can never be randomised?

Part of the answer is that attribution is not only a sensing problem. Capability is not a thing in the world waiting to be instrumented. It exists only relative to a model that says which capabilities matter for this outcome. Point an agent at a client's systems without that model and you get telemetry - activity, throughput, cycle times - which is more measurement but not attribution, because nothing in it tells you which of the observed changes were the ones that mattered.

Better instrumentation tells you the number on a KPI moved. It does not tell you it moved because of you, or even why. The second part is the one that matters.

If you cannot randomise, the alternative is to make the causal chain explicit and then measure along it. That needs three things: a model of which capabilities drive the outcome, an instrument that reads that model consistently against a client's own evidence, and a way of carrying the reading across the engagement so that what was recommended, what was done and what moved are all on the same record.

Those are the three things we build at TheAX: the Atomic Model, AXAT and Outcome Pathways. But the argument holds whether or not you use ours. Some firm has to build this, because the pricing map cannot be used by anyone who cannot locate themselves on it.

Where this goes next

That is the diagnosis, and it leaves two questions open. Before reading on, they are worth putting to your own firm.

Before your next outcome-linked proposal

Can you state the client's starting position as a number today? If not, anything you claim at the end is an assertion, and they are entitled to treat it as one.

Can you show which of your recommendations were actually implemented, and how that changed the outcome? If not, you cannot separate your contribution from everything else that happened that year.

Can you run the same instrument again in six months, regardless of which senior consultant and which client contact were involved the first time? If not, you have a snapshot rather than evidence of movement.

The first open question is practical. Establishing a baseline sounds like precisely the work clients refuse to pay for: weeks of data collection and analysis before anything useful happens. If the evidence has to be designed in at the start, it has to be quick, cheap and repeatable - and repeatable in a stricter sense than the word usually carries, because the reassessment at the end has to run against the same capability definitions and the same scoring logic as the baseline.

Without that reassessment against a consistent model, what you are measuring is the consultant. That constraint has consequences for how a firm has to be built, and it is where part two of this series goes.

The second open question is about what happens when the same model is used repeatedly. One engagement tests what your firm believes about this client. A hundred engagements can test the model itself: which hypotheses held, which failed, which capabilities moved as predicted, and where the firm's reasoning had to change. That is part three.

If you would rather see it as a working example than read about it, our Sunshine Power worked example runs the whole sequence on a fictional client: a first reading with its uncertainty made explicit, a refusal to price the main programme until that uncertainty is settled, and the same measurement logic used to track what changes through delivery.

AI is removing the hours from consulting. It is not removing the value. The firms that can evidence what their expertise changes may finally have a better basis on which to charge for it than the time it happened to take.

The shift

This is one piece of a longer argument

Depending on a few senior people has always capped how fast a firm can grow. Clients moving to outcome-based work is about to make that considerably more expensive. The full argument sets out why the constraint has held for seventy years, and what changes now.