Skip to content

Business

How do you measure whether automation worked?

Automation worked when a step of the work is gone from the process and the number you singled out beforehand has moved. Four measurement points settle the question, and the three numbers most often reported are not among them.

By Sveinung Totland6 min read
An older man in a knitted sweater marks a sticky note on a wall-sized process map of printed sheets and colored notes, while a younger woman beside him holds an open laptop.

Automation worked when a step of the work is gone from the process and the number you singled out beforehand has moved. Four measurement points settle the question: the share of cases that ran the whole way without a person, the share of cases that came back, the cost per case, and the time from the customer making contact to the case being closed.

Apps AS has been building its own products since 2008 and now runs an agent that writes this blog. The figures in this article were counted in Apps AS's own systems.

What does it mean that automation worked?

Automation worked when a step that used to need a person no longer needs a person, and the process still produces a result at least as correct as before. Both conditions have to hold at once. A step moved from a paper form to an agent, where someone still reads every single answer before it goes out, has been digitized rather than automated.

The distinction between digitized and automated is worth keeping because the two states look alike in a status meeting and different in a set of accounts. A digitized step moves the work onto a screen. An automated step removes the work. The operations side of the same distinction is in Personalization is an operations task, not a campaign.

Which numbers measure something other than what you think?

An activity number tells you how much the system has produced. An outcome number tells you how much work is gone. Most automation reports open with an activity number, because the activity number is the only one that is easy to pull straight out of the system itself.

The number reportedWhat the number measuresWhat you need alongside it
Cases the agent handledVolume into the processThe share of cases finished without a person
Hours savedA conversion made with an assumed hourly rateHours actually billed, or spent on other work
First response timeHow quickly something happensThe time from the customer asking to the case being closed
Share of staff who adopted the toolInternal distributionThe share still using the tool in month three
Accuracy measured on a test setHow well the system does on known examplesThe error rate in production, on cases nobody has seen before

The bottom row is the most expensive one to skip. A system that does not give the same answer twice cannot be approved once and counted as approved from then on. How risk moves when software acts on its own is covered in What changes when software acts on its own.

5 of 12
article drafts created in the Norwegian article archive on apps.no between 21 and 25 September 2026 have been published. Four are still drafts and three were archived without being published. The production figure for that week is twelve. The outcome figure is five. Source: Apps AS's own CMS, counted 28 September 2026.

Which four measurement points show whether automation worked?

Four measurement points are enough for most automated processes. All four share the same denominator: the number of cases that entered the process in the period.

  1. Share of cases finished without a person. The share of cases that went from start to a finished outcome without a person doing anything. This is the only number that says directly whether a step of the work has been removed.
  2. Return rate. The share of finished cases that come back as a complaint, a correction or a fresh enquiry about the same thing, within a window you set yourself. The return rate is the protection against making mistakes faster than before.
  3. Cost per case. Model usage, integration, operations and the human review that still exists, divided by the number of cases. The review belongs in the calculation, and why the review does not disappear is covered in Manual review of AI agents is not a cost you can remove.
  4. Time to a closed case. The time from the customer making contact to the case being finished, measured across the whole distribution rather than on the average alone. The median and the ninetieth percentile say more than the mean, because the tail is what the customer remembers.

Measure only the first point and the share finished without a person can rise while the return rate rises just as much. The process has then become faster at producing cases it has to do over.

Why does the baseline have to be measured before automation goes live?

The baseline has to be measured before automation goes live because there is no way to reconstruct the baseline afterwards. After launch you have one number and nothing to compare it with, and any movement in the number then gets read as an effect of the automation.

Three things make a baseline hard to recover after the fact. Seasonality: a support service in January and the same service in July are two different services. Case mix: automation usually takes the easiest cases first, so the remaining cases are heavier and the average looks worse even when each individual case goes better than before. Other changes made at the same time: a new form, a new price list or a new integration in the same quarter makes it impossible to attribute the movement to the agent.

Measure the four points for at least four weeks before you start, and write down which other changes were made in the same period.

What do you do when the number does not move?

When the number does not move after a quarter, the reason is usually that the process was not mapped well enough to be automated. Go through the three checks below before you build more.

Three checks before you build further

Who owns the process now? Without a named owner with the authority to change the rules, the process stands still however good the agent is.

How many exceptions are there? A process with 20 exceptions is 20 processes. Count the exceptions before you automate more of them.

Which step has actually been removed? If nobody can point to a step that is gone, the process has been digitized. The measure is not the thing that is wrong.

A project that has not moved the number after two quarters has reached a stopping point and does not need more development. The figures on where agentic AI actually pays off are in How Will Agentic AI Improve My Business?. What a supplier can guarantee when the outcome is what is being sold is in What a consultant delivers when agents do the work.

Count the four numbers for one process this week

Pick one process you have automated this year. Count the four measurement points for that process: share of cases finished without a person, return rate, cost per case, and time to a closed case. Send the four numbers to hello@apps.no. We will reply with which of the four we would make the steering measure, and which one we would stop reporting.

Sources

  • Apps AS, article archive on apps.no, count of Norwegian article drafts created 21–25 September 2026, counted 28 September 2026. apps.no
  • Apps AS, Manual review of AI agents is not a cost you can remove, read 28 September 2026. apps.no
  • Apps AS, Personalization is an operations task, not a campaign, read 28 September 2026. apps.no

About this article. Written by Sveinung Totland, chief executive of Apps. A satellite to the pillar From digitization to automation. The count of article drafts was made in Apps AS's own CMS on 28 September 2026 and covers Norwegian drafts created between 21 and 25 September 2026. Who created each individual draft is not separated out in the count. The article uses no customer figures, and no figures from third-party surveys, because none of the surveys read this week could be checked against a primary source.

FAQ

How long do we have to measure before the numbers mean anything?
Allow a quarter for a process with a steady caseload, and two for a process with seasonal swings. What decides it is not the calendar but the number of cases: below a few hundred cases in the period, movement in the number is as likely to be noise as a real effect. If your volume is low, measure over a longer stretch rather than concluding early.
What do we do if we never measured the baseline before we started?
Use a comparable process that has not been automated as the reference, or the share of cases that still runs manually. Both are weaker than a real baseline and should be described that way when the number is presented. The alternative is to switch the automation off for a share of cases for a few weeks, which gives you a sound comparison but costs time.
Should we count salary costs we did not actually cut?
No. Hours freed up and not used for anything else are an opportunity rather than a saving. Write down what the freed hours were actually spent on instead. If the answer is that they went into reviewing what the agent produced, that review is a cost in the calculation and not a gain.
Who should own the numbers, the IT department or the department that owns the process?
The department that owns the process should own the numbers, because that department is the one that can change the process when the number stands still. The IT department should own the underlying data and make sure the numbers can be pulled out the same way every month. Split any other way, reporting tends to land with whoever has the easiest access to the data and the least authority to act on it.
How do we measure a process where the agent only proposes and a person approves?
Measure how many proposals were approved unchanged, how many were edited, and how many were discarded. The share approved unchanged is the closest thing to a completion rate in a process of that shape. If the edited share climbs over time, the usual cause is a shift in the mix of cases rather than the system getting worse.