Nobody can prove it worked, and that was decided at the start
A striking number of transformation programmes cannot demonstrate what they returned. BCG puts it at roughly three quarters of organisations unable to prove the ROI of their transformation programmes. Morgan Stanley, surveying S&P 500 companies, found only about a fifth could cite a measurable AI benefit at all. MIT research on AI pilots is starker still, finding the overwhelming majority delivered no measurable profit and loss impact.
The usual reading is that measurement was neglected. That is not what happened. Measurement was made impossible in the first ninety days, and everything after that was consequence.
A benefit statement is not a hypothesis
Most business cases contain a benefit statement. Efficiency will improve. Cycle time will fall. Decisions will be faster. Each is directionally true and none is testable, because none of them names what would have to change in behaviour for the benefit to appear, or what measure would move if it did.
A value hypothesis is a different object. It says: this specific group will work this specific way, which will move this specific measure, and here is the baseline it moves from. It can be wrong, and being capable of being wrong is what makes it worth writing.
The chain has to be traceable, not asserted
The second failure sits between the hypothesis and the measure. A programme can hold a perfectly good hypothesis and still be unable to prove anything, because nobody built the path from the work to the number, and nobody was made accountable for that number landing.
This is what SA-2 measures. Not whether KPIs exist, which they always do, but whether there is a traceable chain from the work being done to the measure that would prove it, with a named owner at the end of it.
Go-live is not the finish line. It is the starting line for value realization.
Why it is invisible until late
A programme with an untestable value case reports green for a long time. Delivery milestones are real and they are being met. Nothing in the reporting structure is designed to notice that the benefit chain was never built, because reporting measures activity and activity is genuinely happening.
The bill arrives at benefits review, typically a year or more later, when someone asks what it returned and discovers that the question has no answer available. At that point the cost of building the measurement retrospectively exceeds what anyone will authorise, so it is quietly not built.
What to ask in the first ninety days
Three questions, and they take an afternoon rather than a workstream. What specifically will people do differently. What measure moves if they do. Who owns that measure landing, by name, and what is its baseline today.
If the third question has no answer, the programme has a benefit statement rather than a value hypothesis, and it is already on the path to being unprovable. That is worth knowing in month one, when it is cheap, rather than in month eighteen when it is not.
The absorption ceiling nobody scores
Concurrent initiatives draw on the same finite capacity, and no single programme can see the total draw. The ceiling is real, it is measurable, and almost nobody measures it.
Shadow workflows are quietly holding your adoption numbers up
Go-live does not eliminate the parallel process. It drives it underground, where it keeps the numbers looking healthy while the transformation quietly fails to land.
Your governance was accurate once, on the day it was written
Governance does not fail loudly. It is correct at publication and degrades from that moment, while the document keeps looking authoritative because documents always do.
Start with the free diagnostic
Ten questions, an indicative readiness band, and one prescribed next step. Returned on screen, no call required.
See what it measures