AI Consulting

Thirty days, one spreadsheet, no analyst.

Most “this delivered a forty per cent improvement” claims are explained by something duller than the software: it was measured after a bad month, and bad months are followed by ordinary ones. Here is a protocol that survives that, and it fits on one page.

Measure anything just after an unusually bad stretch and it improves, whether or not you did anything.

the mean

MeasuredRegression to the mean, a property of any measure with noise in it. Not a criticism of vendors; a criticism of measuring after the month that made you angry.

A protocol that survives scrutiny needs one number, defined before you start, and two equal windows.

1 number

AssumedThe protocol is set out in full below. It is a design, not a finding, and its value is that it is fixed before you can be tempted by the result.

How far vendor-claimed improvement sits above measured improvement.

not measured

Not measuredWe have not run this protocol against a set of vendor claims and published the gap. When we do, our own underperforming automations go in the same table.

Here is how nearly every impressive improvement figure gets made, and almost nobody involved is lying.

A bad month happens. Complaints, or missed deliveries, or a run of no-shows. The owner has had enough and buys something. It goes in. The next month is better. Everyone agrees the thing worked, and a number gets written down.

The difficulty is that the next month was going to be better anyway. Anything that goes up and down will, after an unusually bad stretch, tend to come back towards its ordinary level, for no reason at all beyond that being what unusual means.1 A bad month is not the start of a trend. It is a bad month. And we are extremely bad at accepting that, because a cause is more satisfying than a distribution.

Which means the single worst moment to measure the effect of anything is immediately after the thing that made you buy it - and that is exactly when everybody measures.

Buying something after a bad month and measuring the month after is a procedure that produces a good result whether or not you bought anything.

what it is normallyyou bought it herethe “improvement”that was coming anyway
The same noisy measure either way. If you draw the arrow from the worst point, you will find an improvement every time, and it will be real, and it will not be yours.

There is a second thing, and it is smaller than people say

People also behave differently when they know something is being watched. This gets called the Hawthorne effect, after factory studies from the 1920s and 30s whose original numbers have been picked apart repeatedly since, so I would not lean any weight on the studies themselves.

The uncontested part is enough for our purposes: in the first weeks after you install something, everyone knows it is being judged, so everyone tries harder. That effect is real and it fades. It is another reason a thirty-day window immediately after go-live flatters whatever you installed.

The protocol

One page. No analyst. The whole value of it is that every decision gets made before you can see the result.

  1. Pick one number, and write down its definition.Not three. One. “Bookings that did not turn up, counted at close each day.” Write the definition down before you start, because the most common way this goes wrong is the definition quietly improving along with the number.
  2. Get thirty days of it from before, and not the worst thirty. This is the step everybody skips and it is the one that matters. If the last month was unusual, go back further and take a normal one, or take ninety days and use the average.
  3. Write down what you expect, as a number, with a date. Before switching anything on. Sealed, in the sense that you do not get to revise it afterwards. This one step is what separates measuring from storytelling, and it is uncomfortable precisely because it can catch you out.2
  4. Switch it on, then wait two weeks before you start counting. The first fortnight is everybody trying harder because they are being watched. Throw it away deliberately.
  5. Count thirty days. Same definition, same person counting, same time of day.
  6. Compare the two numbers, and compare both against your written prediction. Three numbers, one line. If the result beat the before but missed your prediction badly, that is worth knowing too, and it is the finding nobody ever records.

The one question that spoils a good result

When it improves - and it often genuinely does - ask this before you celebrate:

What else changed in those thirty days? A festival. A new person on the counter. A competitor shutting. Weather. A price change. If anything else changed, you have a result you cannot attribute, and the honest thing to write down is “improved, confounded,” not a percentage.

You will not always be able to isolate it, and a small business cannot run a controlled trial. That is fine. What is not fine is reporting a number as though you could.

What to do with the answer

If it worked, you now have a number you can defend, which is worth more than a bigger number you cannot. Keep the page. Do the next thing the same way.

If it did not, you have learned something for the price of thirty days instead of the price of two years, and you have learned it before you rolled it out to four branches.

And if the vendor pushes back on the protocol before it runs - the window is wrong, the number is the wrong number, it needs longer - listen carefully, because they may be right. Then ask them to say which number and which window, in writing, before the start. Anybody happy to do that is worth working with.

How this paper was made

This paper contains no measurements. It is a protocol and an argument for why the protocol is shaped the way it is.

The regression-to-the-mean argument is a statistical property, not an empirical claim about vendors, and it is cited below. We mention the Hawthorne effect once and flag it carefully: the original factory studies it is named after have been reanalysed and disputed for decades, so we lean on the uncontested half - people behave differently when they know they are being watched - and not on the original numbers.

We have not yet run this protocol across a set of live automations and published the claimed-against-measured gap. Pillar three says so. When we publish it, the automations of ours that underperform their claim will be in the same table as everybody else’s, or the exercise is not worth publishing.

On the date at the top of this page. This paper is dated 14 September 2026 because that is its slot in the series. The writing and the working were done on 26 August 2026, when the series was compiled ahead of its slot. We would rather say that here than have you find it in the page history.

References

  1. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. Chapter on regression to the mean, including why we invent causes for it.
  2. Tetlock, P. E. and Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown Publishers. Scored, dated, falsifiable predictions beat confident narrative. Almost nobody keeps score, because keeping score is how you find out you were wrong.