Measurement
What a Good Baseline Looks Like
You can't improve a number you never wrote down. The four measurements we take before any AI touches a workflow — and the one everyone forgets.
Priya Nair · Jun 24, 2026 · 6 min read
Every disappointing AI project we've been called in to rescue shares one trait: nobody wrote anything down before it started. The team feels like things got better, the invoice says things got expensive, and there is no third column. A baseline is the third column. Here's how we build one in the two weeks before a pilot.
The four numbers we always take
- Volume — how many of this task arrive per week, and how spiky the arrival is
- Touch time — minutes of human attention per item, measured, not estimated
- Quality — whatever the team already trusts: CSAT, error rate, reopen rate
- Latency — how long the requester waits, wall-clock, including the queue
None of these are exotic. The discipline is in taking them before the pilot, over a long enough window to smooth out a bad week, and agreeing in writing which one the pilot exists to move.
The one everyone forgets
Exception rate: what fraction of items are weird — missing context, angry customer, three intertwined questions, policy edge case. This number decides your ceiling. If 30% of tickets are exceptions, no assistant will safely touch more than the other 70%, and your business case has to survive that arithmetic from day one.
The baseline isn't there to prove the AI worked. It's there to make it impossible to lie to yourself about whether it did.
Measured, not remembered
When we ask a support lead how long a ticket type takes, the answer is usually half the measured figure. Not dishonesty — memory compresses the tab-switching, the lookup in the second system, the interruption in the middle. So we time real work: a week of it, sampled across the team and across the day. The gap between remembered and measured time is often the single biggest correction to the business case, in either direction.
Keep it boring, keep it small
A baseline that takes a quarter to assemble is a project of its own, and it will die of scope. Ours fit on one page: four numbers plus the exception rate, each with a date range and a source. If a stakeholder can't glance at that page and say "yes, that's our workflow," it isn't done. When the pilot ends, the same page gains a second column — and the decision to scale or stop mostly makes itself.
Lumen helps support and operations teams put AI to work — one measured pilot at a time. If any of this sounds like your team, book a call.