Metrics for AI projects - what your team feels versus what the board wants to see

Metrics for AI projects - what your team feels versus what the board wants to see

Most AI programmes I’ve worked on reach the same point about six months after go-live: the technology is stable, people are using it, the feedback is good, and attention moves from whether the tool works to whether the business is better off. That second question is harder to answer than it looks, and it tends to catch programmes out because nobody planned for it at the start.

A typical update at that stage looks something like this. 

  • A platform live across more than 1,000 users

  • A dozen workflows running every day

  • Weekly active usage above 70%

  • A satisfaction score of 4.2 out of 5

  • An average saving of an hour per person per day

Then the Finance Director asks where the hour went.

At 1,000 people saving an hour a day, that is roughly 125 full-time equivalents handed back to the business. So, which part of the business is using them? In my experience nobody can firmly answer the question, and the measurement is not usually the problem.

The programme measured whether the tool saves time, but the question on the table was whether the business changed. Those are different questions, and most programmes only ever answer the first one. The reason the second one is difficult has less to do with measurement than with how organisations behave once capacity is freed up.

A returned hour per person has three possible outcomes. 

  1. The same people produce more. 

  2. The work gets done with fewer people - less overtime or less contractor spend. 

  3. A permanent backlog finally clears.

There is no fourth outcome, and if none of those three numbers has moved once the programme has settled, the hour was not saved but absorbed.

Absorption is not a failure of the tool, and it is not a failure of the people using it either. It is what happens in any organisation with more work available than capacity to do it, which describes most of the functions I have worked in. Saved time does not wait in a holding account for a management decision. It goes on the slightly longer review, on the meeting that had been compressed out of necessity, on the low-priority requests that used to get declined. Some of that is valuable and much of it is work expanding to fill the space.

It is also invisible to the people it happens to. Nobody notices their saved hour disappearing. They notice a less pressured week, which is an improvement in working life and a poor substitute for a number a board can act on.

None of this is an argument against measuring time saved. It is an argument for agreeing what else you will measure across the whole process before anything is built. I learned that on a programme I led about a year and a half ago.

A worked example from an invoice pipeline replacement

The programme replaced a legacy optical character recognition pipeline, which read invoices arriving in a shared mailbox and posted them to the enterprise resource planning platform, or ERP. The replacement used a large language model to read the invoice visually and extract the fields, rather than matching characters against per supplier templates. We agreed on three measures before anything was deployed and took a baseline for each. That decision is the only reason I can tell you what actually happened.

Measure Baseline After a year in production
Invoices posted automatically 32% 94%
Invoices needing manual initial posting 68% 6%
Corrections needed in the ERP 7% of volume 9% of volume
Time to post 48 hours 24 hours

The automation rate is the number everyone wanted to talk about. The old system handled 32% of invoices on its own and left 68% to be keyed in by hand. A year in, 94% were posting automatically and manual keying had dropped to 6%.

The second measure is the one that made the case credible, because it moved the wrong way. Invoices posted automatically were sometimes posted wrongly, and someone had to find and fix them in the ERP. Those corrections went from 7% of total volume to 9%. The team was now fixing more invoices than before, and we reported it that way rather than leaving it out, because we had accepted that trade when we chose the design. What we got in exchange was the removal of manual keying across most of the invoice book, which was not a close contest.

The two figures also said something about the extraction quality that neither said on its own. Corrections at 7% of total volume against 32% automated meant the legacy system got roughly one invoice in five wrong. Corrections at 9% against 94% automated meant the new pipeline got roughly one in ten wrong. Accuracy per automated invoice improved considerably, and the absolute correction volume only rose because far more invoices were being automated in the first place.

Reporting the automation rate alone would have described a clean win and hidden the extra correction work, which is the sort of omission that costs a programme its credibility when somebody finds it six months later. Reporting the correction rate alone would have suggested quality had degraded, which was not true.

The third measure is where the value showed up outside finance. Invoices that took 48 hours to reach a posted state were getting there in 24, which moved a meaningful volume of payments inside early settlement windows. Those discounts were cash, rather than notional productivity.

Setting up the right measurements before starting an AI programme

Three things from that programme carry straight into the way I would set up measurement on an AI deployment today. 

The first is to measure the process from end to end rather than the tasks inside it. Task level savings disappear into queues and handovers later in the process. A document drafted in 20 minutes instead of 90, changes nothing for the customer if it then sits in a review queue for four days, and the same arithmetic is now catching out delivery teams using AI coding assistants. When development took five or six weeks, another three or four weeks of change boards, firewall approvals and security reviews felt proportionate. Build the same thing in a fortnight and the bottleneck has not gone away, it has moved to the approval cycle.

The second is to take the baseline before the tool arrives. It cannot be reconstructed afterwards, and without it every later conversation about value comes down to opinion.

The third, and arguably the most important to convince senior stakeholders, is to agree where the freed capacity is going at the point of funding, and to give that decision to an owner in the business rather than in the programme. Either it absorbs a planned hire, or it goes to a named backlog, or it comes out as cost. Settling that early is the difference between a benefit you can report and one you have to argue for.

The hour in your business case is probably real. Being able to say where it went, and who is using it now, is what turns it into a result a board will accept, and that comes from a few decisions taken before the build starts.

Alvaro Escudero Angulo

Alvaro Escudero Angulo is a Senior Consultant at WeBuild-AI. With over 4 years of experience in cloud and AI, focussed on the manufacturing sector, he has led programmes centred around replacing backend processes with AI-powered pipelines, to maximise enterprise efficiency and spend, to support teams to focus their time on where they are most needed. He now works with clients in both financial services and energy sectors, drawing on his background in both computer science and business management, to drive success in our AI programmes.

Previous
Previous

The AI-Native SDLC eBook

Next
Next

The Knowledge Graph Tool Inside Our AI Accelerator