AI Consulting

How we choose which job to fix first.

The method, in enough detail that you could run it yourself this afternoon without us. Two axes, five tests, and one rule about which corner to start in.

Why we are giving the method away

The obvious commercial move is to keep this vague, so that choosing the first project looks like something only we can do. We are writing it out instead, for two reasons.

The first is that a method you can check is worth more to you than a method you have to trust. The second is more practical: if you can run this yourself and the answer comes out “nothing here is worth doing”, we both save several weeks. That outcome is a success, not a lost sale.

The two axes

Every candidate job gets placed on a grid with two axes, and neither is about technology.

  • How much would this actually help.In money or hours, from your own figures. Not “strategic value”, which is a phrase that exists to avoid this question.
  • How hard is it to do. Which includes the technical build, but in a small business is usually dominated by something else: how many people have to change what they do on a Tuesday morning.

Four corners fall out of that, and they are not equally interesting. High help and low difficulty is where you start. Low help and high difficulty is where a great deal of what gets sold as AI actually sits, and we will tell you to skip it even though we could charge you to build it.

Start herehelps a lot, easyPlan for ithelps a lot, hardOnly if freehelps little, easySkip ithelps little, hardHow hard it is to doHow much it helps
A great deal of what gets sold as AI sits in the bottom right. We will tell you to skip it even though we could charge you to build it.
The grid is not the clever part. Almost everyone draws this grid. The clever part is being disciplined about the vertical axis, which is the one everybody inflates.

The five tests

A job that passes all five is a good first project. A job that fails two or more is not, however appealing it sounds. Score each one yes or no, not out of ten — scores out of ten are how a no becomes a “six”.

1. Can you count it today?

Is there a number, or could there be one inside two weeks by hand? If not, you will never know afterwards whether it worked, and you will decide on a feeling. That is not a small risk: expert developers in the METR trial felt 20% faster while being 19% slower, a 39-point gap between the feeling and the fact.

2. Does it happen often?

Volume is what turns a small per-instance gain into money. A task done four hundred times a month with a modest improvement beats a task done twice a month with a dramatic one, and the second is far more tempting to build.

3. Is the person doing it inexperienced at it?

This is the least obvious test and it has the best evidence behind it. Across 5,179 customer support agents, an AI assistant produced 14% more resolutions per hour on average, 34% more among the least experienced, and close to nothing for the most experienced (Brynjolfsson, Li and Raymond). The gain concentrates where the skill gap is. Pointing a tool at your best person is the most common way to get a null result.

4. Can somebody check the output?

Not proofread everything forever, but catch a wrong one before it reaches a customer or an account. If nobody can tell a good output from a bad one at a glance, the failure mode is silent, and silent failures are the expensive kind.

5. Is there one named person who will switch the old way off?

Named. A person, not a department. If the old route stays open, the new one is optional, and optional loses to habit every time. This test kills more candidate projects than the other four together, and it is free to apply.

Which corner to start in

Start with the highest-help job that passes all five, not the most impressive one. First projects have a job beyond their own payback: they establish whether this business can absorb a change at all. A modest thing that works and gets used teaches you more than an ambitious thing that half-works.

And if the highest-scoring job in your whole list still does not clear the cost of doing it, that is the answer. Not a smaller version of the project. The answer.

The honest limit of this

We should be clear about what these five tests are and are not.

They are a structured way of asking questions that usually go unasked, drawn from the evidence above and from our own work. They are not a validated instrument. We have not scored a population of past projects against their eventual outcomes and computed whether the score predicts success. Until we have done that, calling this “proven” would be exactly the kind of claim this guide spent a whole chapter taking apart in other people’s marketing.

What we can say is narrower and true: applying them takes an afternoon, costs nothing, and reliably removes the candidates that would have failed test five. That is worth doing whether or not we are ever involved.

The next chapter is what happens when we do it with you rather than you doing it alone.

Where these numbers came from

Every figure on this page, with what it is and where it is from. If a number is illustrative rather than measured, it says so here and it says so in the text.

  1. 39 pointsThe gap between what expert developers believed about their own productivity with AI and what was measured. They predicted 24% faster, felt 20% faster, and were 19% slower.METR randomised controlled trial, 16 experienced developers across 246 real issues on mature codebases, July 2025.From a named study
  2. plus 34%The productivity gain for the least experienced customer support agents given an AI assistant, against 14% on average and close to nothing for the most experienced.Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025, pages 889 to 942. Staggered rollout across 5,179 support agents; first circulated as NBER working paper 31161 in April 2023.From a named study
  3. not validatedWhether the five tests below actually predict which projects succeed. We have not scored past projects against their outcomes and computed a correlation, so this is a structured judgement rather than a measured instrument.Stated as a limit rather than claimed. The same admission appears in our research paper "Visible before it starts".We counted it