The part clients will not pay for
In part one I argued that making the Stanford pricing map commercially practical requires a capability most firms do not have. You can observe what the work cost and how long it took, and you can observe the result. But the thing that determines what a firm can safely promise sits between them: whether anyone can see what actually happened, and whether the capabilities that drive the outcome moved.
Which means the evidence you collect must be designed into the engagement at the start rather than assembled at the end. And that runs straight into the reason most firms never do it.
Establishing a baseline sounds like precisely the work clients refuse to buy. Weeks of interviews, data collection, diagnostics and analysis, producing a document that describes the problem they already told you they had. Very few clients want to pay for that today, and they are right not to. So the baseline either gets skipped, or it gets buried inside the programme fee where it becomes something the firm absorbs rather than something the client values.
For evidence to survive a procurement conversation it has to be quick, cheap and repeatable. Not accurate instead of those things. Accurate and those things, because an accurate baseline nobody will fund is not a baseline.
Repeatable means something stricter than it sounds
Quick and cheap while maintaining quality are the obvious constraints that any professional service business battles with. Repeatable, however, does not mean what most firms mean by it.
Defensible evidence of movement requires the reassessment at the end to run against the same capability definitions, the same scoring logic and the same evidence standards as the baseline. Not a similar framework. The same one. And not just run by the same consultant, but run in exactly the same way.
If two consultants would score the same client differently, or the same consultant would score it differently in March and October (which they always will, having gained more knowledge of the client and the project in between), then the movement you report is not the client's movement. It is a mixture of the client changing and the reader changing, and there is no way to separate them afterwards.
Measure with a different instrument and the movement you report is not the client's movement. It is a measurement of the consultant.
This is why evidence of movement cannot be added at the end as a reporting exercise. By the time you are writing the closing report, the property that would have made it defensible was either built into the first reading or it was not.
Comparability is not attribution, and it is worth being exact about that. It tells you the client moved. It does not tell you that you moved them.
Attribution comes from the fuller chain: a baseline, a hypothesis stated in advance, a recorded intervention, measured capability movement, outcome movement and the context around all of it. Even then you are reducing the uncertainty rather than proving causation. But comparable measurement is the part without which none of the rest is worth anything.
Which is why it cannot live in someone's head
The usual argument for systemising consulting expertise is cost and scale: you want your best people's thinking available without your best people having to perform every piece of work. That argument is true and valuable, and I have made it myself. It is not the argument that matters most here.
The constraint that forces systemisation is comparability. A model that has to produce equivalent readings six months apart, run by whoever is available, applied to a client the original assessor never met, cannot be a way of thinking that a senior consultant carries around with them. It must be explicit: the capabilities named, the questions defined, the scoring rules written down, the evidence standards fixed in advance.
Firms tend to resist this because it sounds like turning judgement into a checklist. It is the opposite of a static checklist applied as a box-ticking exercise. The judgement is what gets encoded: which capabilities matter for which outcome, what good looks like, what evidence is sufficient. What gets removed is the variability in how consistently that judgement is applied, which was never the valuable part, and which has always been the thing that stops outcomes being delivered repeatably.
What that looks like on a reading
Our Sunshine Power worked example was built to show how this runs, and what it changes commercially. Sunshine Power is a fictional mid-market B2B solar infrastructure business whose revenue is growing but not predictably, and whose CRO calls in BIG Consultants Inc for advice.
In the example, the first assessment produces a corroborated position of 2.4 out of 5, a capability floor of 1.7 and 84% confidence in the reading, alongside 42 unresolved questions - of which 11 could move a capability rating on their own. All of it inside about 30 minutes of a consultant's time on the platform.
That last number is the one that makes doing the assessment commercially feasible in the first place. A baseline that takes half an hour is something a client will buy, or something the firm can carry as a cost of sale. A baseline that takes three weeks is something a client has to be talked into.
The conventional move would be to use that first diagnosis to scope a transformation programme and price it. The assessment dossier declines to do that, on the grounds that a fixed-scope, fixed-price programme built on this reading would be priced against a picture that may not survive contact with the evidence.
So instead BIG Consultants proposes a nominal £1,725 of fixed-price work, a day and a half, to settle the reading first, then generates the corrected roadmap and prices the larger engagement from the plan rather than from the guess. That is under 2% of the programme fee, spent only where a playback has already landed, and spent by a client gaining confidence in the firm's data and in its evidence of how the outcome will be reached.
The client is fictitious, but the numbers are not invented. They are real outputs from TheAX's algorithms, run against the test data and the underlying model built for the example.
The detail worth copying is what else that time is used for by a consultancy. The same sessions agree the outcomes the work is meant to move and baseline the KPIs that evidence them, taken from the client's own reporting where the number already exists and named as a gap where it does not. The measures get fixed while both parties still have an open mind about them, rather than argued over at the end when one side has an invoice to defend.
The principle the example draws out is the important one: a commitment made against an unchecked reading is not a commitment. It is a guess with a signature on it.
Risk does not disappear. It changes hands.
To be precise about which risk is moving, there are two of them.
The first is where the hourly effort is visible and charged for, and the change is not. Here the client carries the outcome risk and your exposure is commercial rather than operational.
This is where your work becomes devalued, or stops being buyable. AI can make the work look quick and easy to deliver, and it can turn access to the knowledge you are selling into a commodity, which is then valued and priced accordingly.
If you cannot show the effort, because AI has made your consultants faster and the hours have disappeared, and you cannot show the change, because you do not have the repeatable systems that would evidence it, commoditisation accelerates. However good a consultant you are, there is nothing left to compete on but price, which is the argument part one makes at length.
The reflex consultancies are making today is to believe they can offer outcome-based terms to win the work back. But that creates the second risk. Offering outcome pricing does not create value. It moves risk off the client's balance sheet and onto yours. To manage that risk you have to have the systems that let you carry it.
So the dilemma is that staying on hourly rates is a risk; moving to outcome pricing without the systems to carry it is risk transfer in the wrong direction.
The opposite is also true. A firm with the systems to evidence the chain can price it. It knows what it believes should move and why, what evidence stands behind that belief, and how it will test that belief on this engagement. A firm that cannot is writing the same terms without knowing the odds, and only finds out at the end of the project whether the guess was right, which is the worst available moment to discover it.
The axis that matters is not really about establishing proof for the client's benefit, but for the consultancy. Evidence is not only the thing you show the client. It is the thing that tells you what you can afford to promise them.
None of which argues that every engagement should be outcome-priced. Sometimes the evidence is weak, the result takes years, or you control very little of what happens after the recommendation lands. The advantage is not being forced into a single model. It is having enough evidence to choose, to show the client why so they understand the situation in data, and to know before you sign which risk you are carrying.
That is what a good baseline buys you on one engagement. Part three of this series is about what a hundred of them buy you, and why a firm with a tested model will be answering a question its competitors cannot.


