Short answer
Measure four things the same way before and after: hours the work takes, how long it waits, how often output needs correcting, and whether the team still uses it in week four. Take the before-numbers first. Without a baseline you are comparing today against a memory, and memory flatters whichever side you prefer.
Key takeaways
- Take the baseline before anything is built. It takes an afternoon and it is the only chance you get.
- Adoption predicts everything else. An agent that saves four hours and gets quietly abandoned saves nothing.
- Correction rate near zero is not automatically good news. It can mean nobody is checking.
- Review monthly, decide at 90 days: expand, fix, or stop. Agents nobody decides about are the ones that rot.
The most common way this goes wrong is not a bad agent. It is that nobody wrote down what the work cost before, so three months later the conversation is two people trading impressions. Half an afternoon of measurement at the start prevents that entirely.
The four numbers
Hours: how long the work takes, counted the same way before and after. Cycle time: how long it sits waiting between steps. Correction rate: how often output needed fixing, measured for the human version too. Adoption: whether the team still uses it in week four without being reminded.
| Number | How to take it | What good looks like | How it misleads |
|---|---|---|---|
| Hours | Time the task for a week before, and the same week after | A drop you can feel in the schedule, not a rounding error | Counting only the agent’s run time and ignoring the review time |
| Cycle time | From request to finished, including the waiting | Usually improves before hours do | Improving cycle time while pushing work to somebody else |
| Correction rate | Count corrections on both the human and agent versions | Comparable to the human rate within a few weeks | A rate near zero because the approver stopped reading |
| Adoption | Is it still in use in week four, unprompted | Nobody asks whether they have to use it | Counting logins instead of completed work |
Compare against reality, not against the ideal
The baseline has to be how the work is really done, including the days it does not get done at all, the rework, and the waiting. Teams often benchmark an agent against an imagined version of their own process that never existed, then conclude the agent underperformed a standard nobody was meeting.
The reverse trap is just as common. If the human version had a 6% correction rate and the agent has 5%, that is a win, and it will still feel like a loss because agent mistakes are strange and human ones are familiar. Write both numbers down and look at them together.
Why adoption is the one that predicts the rest
An agent that saves four hours a week and is abandoned in week five saves nothing. Abandonment almost always traces to one of three things: nobody owns it, it produces work somebody else has to clean up, or the process changed and the agent did not. All three are fixable, and none is fixable if you do not notice.
The monthly review, and the 90-day decision
- Once a month, the named owner brings the four numbers and one sentence on what changed.
- Look at corrections by cause, not by count. A missing rule, a missing example and a changed process need different fixes.
- At 90 days, make an explicit call: expand it to more of the work, fix a specific thing, or turn it off.
- Write the decision down. Agents that nobody ever decides about are the ones still half-running a year later.
What not to measure
Volume of output, number of runs, tokens consumed, and anything a dashboard offers because it is easy to count. They describe activity, not whether the business is better off. If a number would not change your decision to expand, fix or stop, it does not belong in the review.
Questions people ask
What if we never took a baseline?
Take it now on the part of the work still done manually, and be explicit that the comparison is partial. Reconstructing a baseline from memory produces whichever answer the person reconstructing it expected, which is worse than admitting you do not have one.
How long before we judge it?
Four weeks for a first read on hours and adoption, 90 days for the decision to expand, fix or stop. Judging in week one measures novelty. Waiting six months means paying for something nobody decided about.
Our correction rate is zero. Is that good?
Check before celebrating. A rate at zero often means the approver has stopped reading, which is the failure mode that looks exactly like success. Sample a few outputs yourself and compare them against the standard the procedure defines.
How do we measure hours honestly?
Count the whole loop: the agent’s run, the review time, and any rework. Comparing the agent’s run time against the human’s whole loop is the most common way a modest gain gets reported as a dramatic one.
What is a reasonable first result?
A visible drop in cycle time, hours moving in the right direction, corrections comparable to the human rate, and a team still using it unprompted at week four. Anything dramatic in month one is usually a measurement artefact.
Sources
- delegAIte Founder Dashboard and Operating Cadence deliverables, delegaite.co (read 20 September 2026): The measure, monitor and make right cadence, and revenue-per-agent style metrics as part of the install.
Published . Last reviewed .
Want this in your business?
Book a call and we will scope exactly which part of your operation an agent should take first.
Book a call with the team