A project gets a start date, a budget and usually a presentation with an upward curve. What a project almost never gets is a moment when someone establishes that it has done what it was supposed to do. Not: delivered. Done. The work relocated, the behaviour changed, the old way of working no longer needed.
That distinction has always been difficult. It becomes sharper now that part of the improvement consists of AI taking over tasks. With a new system or a new procedure, you can still assume that people will use it because it is mandatory. With AI work, that assumption is not self-evident: someone must keep checking the result, someone must dare to trust an outcome they did not write themselves, and someone must have the authority to approve or reject that outcome. If those three things are not organised, the system runs, but nothing lands.
There are measurements that are reliable because they are factual. How many hours of work shift from performing a task to overseeing that task. How many of the outcomes are approved by a human without changes, how many are rejected, and for what reason. Whether the time between delivery and decision becomes shorter. Whether the number of escalations to a senior employee decreases because the first assessment holds up.
This is capacity, not a promise. Freed-up hours are a fact as soon as you count them; what an organisation does with that freed-up capacity is a separate question, with its own legal frameworks as soon as it affects staffing. Keeping those two questions apart is a precondition for a credible measurement, not a side issue.
A usage percentage says that people press a button. It says nothing about whether they also believe the result. A team can approve an AI outcome for form's sake and do nothing with it in practice, or conversely adopt everything without checking — both behaviours look identical in a dashboard as "adoption 80%". That is the reason a single percentage without explanation is not usable: it describes an action, not a change in how the work gets done.
There is also a timing problem. An improvement that measures well in the first month can decline in the fourth month because exceptions pile up and no one planned for them. Conversely, an improvement that starts slowly can become stable after six months once the people who need to approve it have gained trust in it. A measurement at a single point in time is therefore incomplete by definition; what is needed is a series of measurements that shows whether the curve repeats itself or flattens out.
One company sees AI work land within a quarter, another is still testing after a year. That difference rarely lies in the technology. It lies in who has been trained to assess an AI outcome, who has the decision authority to let that outcome proceed, and whether that decision authority has been formally established or still rests with a single person who "does it on the side". More on who should actually fulfil that role is described in who is needed to make work done by AI succeed.
Another common cause: the project remains technically successful but organisationally isolated. The AI works, the pilot runs, but no one outside the IT department uses the outcome structurally. Why that happens and how you recognise that pattern is described in how do you prevent an AI project from staying stuck in the IT department.
And sometimes the delay is not an organisational problem but a legal or compliance issue that was not factored into the planning; that situation is described separately in what to do when the ambition is possible but compliance is not.
The reason a single KPI is never sufficient is that readiness itself does not fit into one number. Something only lands once it holds up on multiple dimensions at the same time: the people who need to assess it are trained, the data is usable, the process recognises the exception, and the decision authority has been explicitly assigned. If one fails, you can measure progress on the other seven and still have no landing. That is worked out in why readiness is not a figure but a series of questions.
The underlying question — which work in this company can truly be taken over by AI, and which part remains human work — is answered per task with the work scan of FTE TO AI.
No measurement is complete without a time dimension, a cause check and a separation between usage and trust. Concretely, that means: do not just measure whether something is used, but whether the outcome holds up without correction. Do not measure once, but at fixed moments after the start. And ask for the reason with every rejection, because that reason tells you more than the score itself.
If the ambition is too large to test in one go, a phased approach is often more realistic than a one-year plan; how that phasing works is described in how you phase an ambition that is too large for a year. And for the question of how much pace an organisation can actually bear, without the measurement itself causing the delay, there is how fast can an organisation truly change.
A reliable measurement does not start with a dashboard, but with the question of where your organisation truly stands at this moment. The free readiness check asks eight short questions, one per dimension, and gives a picture of where you are furthest along and where you are not. The full ambition test, with the four layers and five confidence gates, is under construction.
Vertel wat u wilt bereiken, dan kijken we samen wat daarvoor moet staan.
Answers come from this site’s knowledge base. Not tailored advice, and not a scan of your company.