All articles
Consultancy-as-a-System

The hundred-and-first engagement

The fashionable answer is to gather everything your firm has written and point AI at it. That gives you a knowledge layer, not an expertise layer. The stronger move is to run every engagement against an explicit model of what your firm believes, so that each one becomes a test of it.

Part one · What you can promisePart two · What it takesPart three · Testing what you believe

Most of the AI conversation in consulting is still about speed

How can consultants research faster. Analyse faster. Build the deck faster. Write the report faster. Use fewer junior people to do it. Get a smaller, leaner machine.

Those are real productivity benefits. But if that is all a firm does with AI, it has bought a productivity gain for the same business model. Worse, as I argued in part one, it has made its own visible effort disappear faster, which is the thing its fee was anchored to. That problem gets hardest quickest for the firms adopting AI most enthusiastically.

The larger opportunity is not speed. It is what the firm knows at the end of an engagement that it did not know at the start, and still knows once the people who ran it have gone. That is the part that builds a firm's IP, and in an AI world it builds business value beyond what sits in consultants' heads. But most importantly, an expertise layer is what makes a consultancy scalable, what supports pricing strategies, and what opens the door to new Consultancy-as-a-System business models.

The firm has not learned. Its consultants have.

A consultancy that has delivered a hundred broadly similar engagements might call that knowledge its greatest asset. But look at where the learning sits. A consultant learns from the engagements they personally run, so a firm of fifty consultants has probably fifty private learning curves and a limited institutional one. Its experience is really the sum of what its people happen to remember.

What is retained in writing is decks, final reports and workshop notes. The records that might have bridged the gap were never created, because no client pays for the hours it takes to write down why a conclusion was reached rather than what it was. You can ask a person instead, and a good senior consultant will give you a useful answer. But only while they still work for you.

Pointing AI at the case studies will not fix it

There is a fashionable answer to this. Gather every deck, report, proposal and case study the firm has produced, point a model at the corpus, and call the result the firm's expertise layer.

The problem is that the category is wrong before the technology is even considered.

Knowledge management gathers what a firm has produced into a corpus, and an LLM pointed at that corpus makes an excellent knowledge layer for search and retrieval. An expertise layer encodes something different: how the firm interprets evidence, decides what to do, and what would change its mind. One retrieves the products of expertise. The other makes expertise executable.

Two problems follow.

First, the corpus contains mostly outputs rather than the reasoning behind them. A deck records the conclusion, not why one diagnosis beat the alternatives or what changed the consultant's mind. Some firms retain working papers and decision logs, but much of the judgement was never written down. AI cannot reconstruct reasoning that is not there.

Second, the evidence that remains shows correlation, not attribution, and there is too little of it. If intervention X precedes movement Y in sixty cases, you know they travelled together, not that X caused Y. And a hundred engagements is a lot of consulting but very little data once divided across sectors, contexts, starting positions and outcomes.

What you get is a powerful retrieval system, and that is useful. But it contains assertions made by different people at different times, not a common claim that successive engagements can test. Nothing is systematically contradicted, so nothing is systematically corrected.

A repository remembers. A model can learn.

And that is the real limitation: the firm has bought a better search engine, while remaining dependent on the same handful of senior people.

Run every engagement against a model instead

The alternative is not a bigger corpus. It is to change what the engagements are being run against.

In part two I argued that comparable measurement forces a firm to make its expertise explicit: which capabilities matter, what good looks like and what evidence is sufficient. That was necessary for measurement. Its second consequence is more powerful.

An explicit model of expertise makes a claim. It says that for a given outcome, certain capabilities matter, and that moving one should influence another. Which means the model can be wrong. And anything that can be wrong can be tested.

So the sequence changes. The model identifies what should matter. The baseline establishes the organisation's position against it. The hypothesis states what the firm expects to change and why. The intervention applies that reasoning. Reassessment shows whether the capability moved, and outcome measures show what followed. The evidence then strengthens, weakens or changes the model.

That sequence is what we built TheAX to run: the Atomic Model holding what the firm believes, AXAT taking the reading, and Outcome Pathways carrying the hypothesis and the result across the engagement. Our Sunshine Power worked example shows the first two steps of it on a single client. But the argument holds whoever builds it.

Three things a hundred engagements can leave behindThree columns compared. The first, documents and memory: slide decks, final reports, workshop notes and the memory of whoever ran the work. Asked a question it returns roughly what was done, if the file can be found. The second, a corpus with AI pointed at it: every deck, report, proposal and written case study, holding conclusions but not the reasoning behind them and nothing about what was tried and failed. Asked a question it returns what the firm has said before in cases that look similar, but it holds no belief and so cannot be wrong. The third, a tested model: what the model predicted, the hypothesis stated in advance, what was done, whether the predicted capability moved, whether the outcome followed, and where the model was wrong. Asked a question it returns what the firm believes, how far that belief has held, and where reality has disagreed with it.Three things a hundred engagements can leave behindOnly the third one can be wrong in a way that teaches you something.WHAT MOST FIRMS KEEPDocumentsand memorySlide decksFinal reportsWorkshop notesThe memory ofwhoever ran itASK IT A QUESTIONRoughly what we did,if you can find the fileand the person isstill here.THE FASHIONABLE FIXA corpus, withAI pointed at itEvery deck and reportProposals and pitchesWritten case studiesConclusions, but notthe reasoningNothing about whatwas tried and failedASK IT A QUESTIONWhat we have saidbefore, in cases thatlook like this one. Itcannot be wrong.WHAT WE ARGUE FORA model thathas been testedWhat it predictedThe hypothesis, statedbefore the workWhether that capabilityactually movedWhere it was wrongASK IT A QUESTIONWhat we believe, howfar it has held, andwhere reality hasdisagreed with us.A repository holds what the firm has said. A model holds what the firm believes,which is the only one of the two that reality can correct.

The hundred-and-first engagement should not start with a hundred case studies. It should start with a model that has been tested a hundred times.

The distinction is between correction after the fact and prediction before it. A record of how a senior consultant fixed the output tells you what your best people would have done. It says nothing about whether the client's forecast accuracy improved because of it. Only a hypothesis stated beforehand can be refused by reality.

The failed hypothesis is the valuable one

Suppose the model says that moving pipeline governance from level two to level three should improve forecast accuracy. You diagnose it, intervene, governance moves as intended - and forecast accuracy does not.

Under the old arrangement, that is an awkward result. Under this one, it may be the most informative thing that happened all year.

Perhaps the causal link is weaker than the firm believed. Perhaps another capability is a prerequisite. Perhaps the KPI was wrong. Each is a specific correction to what the firm believes, and none would have surfaced from records mined for patterns because those records contain no statement of what anyone expected.

This is not statistical proof. A hundred engagements across a hundred organisations are still confounded by different contexts and conditions. But a stated hypothesis makes each result interpretable. You know what you expected, so you can distinguish evidence that supports the model from evidence that challenges it.

That is the difference between collecting experience and learning from it.

You do not need a hundred engagements to start

The obvious objection to anything that compounds is the cold start. If the asset only becomes valuable at engagement one hundred, what happens for the first ninety-nine?

This is where the model matters, and it is why the Atomic Model is the first thing we build with a firm rather than the last. It does not begin empty. It begins with the encoded judgement of the firm's best consultants: the diagnostic logic they already use, the capabilities they know matter and the thresholds they recognise. Decades of experience, made explicit rather than carried in a few people's heads.

So engagement one already starts with the firm's best current understanding and produces the first test of it. Engagement two starts with that understanding plus what the first revealed. By engagement one hundred, the model has been challenged repeatedly and corrected where reality disagreed.

That is very different from starting with a corpus and waiting for enough data to discover the expertise. The model starts with what your experts already know, then uses every engagement to test and improve it.

The question clients will start asking

Consulting buyers have always asked some version of: how many people will you put on this, and how senior are they?

AI is weakening that proxy. As teams and timelines compress, buyers will care more about a different question: how confident are you that you can produce this result, and what is that confidence based on?

A firm with a corpus and a chatbot over it still answers that question the way consultancies always have: reputation, credentials and the calibre of the people in the room.

A firm with a tested model can answer with what it believes, how often that belief has held and where it has failed. Over time, that becomes much harder to compete with.

You can hire a senior consultant next month. You cannot hire a model that has survived a hundred engagements.

The test worth running on your own firm

Write down what your firm believes about one outcome you sell. Which capabilities constrain it, and which one you would move first. If two of your partners write different answers, you do not yet have a model.

Now take your last engagement of that kind. Can you say what you expected to move, and whether it did?

And if it did not move, is that written down anywhere? If it is not, the firm has kept the result and thrown away the learning.

An engagement should leave the firm better equipped for the next one than it was for the last. Most do not, because the thing that would have made that possible - a statement of what the firm expected to happen - was never made. Solving that is worth considerably more than doing the same work faster.

Which is where the three pieces end up, and it is a larger argument than the one they start with. Part one looks like it is about pricing. Part two turns out to be about how a consultancy operates. This one is about what a consultancy becomes: a firm whose expertise is an asset it owns and improves, rather than an attribute of the people it currently employs.

The series

Part one · You cannot price what you cannot evidence
Why AI weakens effort as a basis for price, and what determines whether a firm can carry outcome risk.

Part two · The baseline nobody will pay for
Why comparable evidence has to be designed in at the start, and why that forces expertise out of individual heads.

Part three · The hundred-and-first engagement
This article. Why a corpus gives you a knowledge layer, and what an expertise layer has to encode instead.

The shift

This is one piece of a longer argument

Depending on a few senior people has always capped how fast a firm can grow. Clients moving to outcome-based work is about to make that considerably more expensive. The full argument sets out why the constraint has held for seventy years, and what changes now.