Blog
Share on
Reimagined ITOps is a weekly series about what IT operations becomes when AI stops assisting and starts operating. This is article 3 of Season 1. Stay with it.
The market’s automation numbers have stopped agreeing with each other. One large platform vendor reports that its AI now resolves 90% of the tickets on its own internal help desk. Independent analyses published this year put realistic deflection in customer deployments at 20 to 40%.
Contracts are being signed on both numbers.
Here is what those two numbers have in common. Neither one was measured on your estate.
My claim: an automation potential figure that was not derived from your own ticket data is an estimate, whatever the slide calls it. We stopped estimating automation potential and started measuring it against real ticket estates, ticket by ticket. The measured number came back lower than the estimates it replaced. It is also the first automation number I would defend in a contract.

Fig 1: Estimated vs measured automation potential
Estimates start from the category table. Take the ticket volume report, find the biggest categories, assume some share of each is automatable, and multiply. It feels rigorous because a spreadsheet is involved.
But a category tells you what a ticket was about. It tells you nothing about what should happen to it. A password reset and a license request have nothing in common topically and almost everything in common operationally. Category volume is a map of effort, and effort is the wrong thing to count.
The test that separates effort from automation is convergence. Take 50 resolved tickets from one category and ask a single question: Did the resolutions converge? If 50 tickets were resolved 3 ways, you have found automation. If they were resolved 50 different ways, that category is not automatable, no matter how large it is, and the variety usually means the category itself is a fiction that covers several distinct faults.
Volume tells you where the effort is. Convergence tells you where the automation is. They are not the same list. Mistaking one for the other is the most common error I see in automation business cases, including some of our own older ones.
The second source of inflation is politeness toward data. Estimates trust labels. A populated resolution field is counted as a resolution. A filled owner field is counted as an owner. A knowledge article linked to a fault is counted as coverage of that fault.
None of those things follow. A resolution field that says “fixed, closing ticket” contains no resolution. An owner field pointing at a team disbanded 2 reorganizations ago contains no owner. A knowledge article can exist for a fault and still not resolve it.
We now state this as a rule: attribute presence is not the property it stands for. Content beats labels, every time. Any assessment that counts fields will overstate what the estate can support, because it is scoring the paperwork, not the knowledge.
Measurement means reading the contents. Classify each ticket from what was actually written and actually done, not from what the form claims about it. It is slower. It is also the difference between a number and a guess.
The third correction changed our results the most. When you assess a ticket, the question is not what the human did with it. The question is what should have happened to it.
Score history and you get history. You rebuild the current operation in software, ceiling included, because the current operation is exactly what generated the data. The 8 escalations a ticket went through are not evidence that it needed 8 escalations.
So, the assessment assigns an ideal disposition to each ticket. Should this have been self-served, resolved automatically, resolved with assistance, prevented from existing at all, or does it need a human decision the enterprise has not delegated to any system? The gap between what was done and what should have been done is the size of the opportunity, and it is measurable at the estate, tower, and queue levels.
So, the assessment assigns an ideal disposition to each ticket. Should this have been self-served, resolved automatically, resolved with assistance, prevented from existing at all, or does it need a human decision the enterprise has not delegated to any system? The gap between what was done and what should have been done is the size of the opportunity, and it is measurable per estate, per tower, per queue.
That last part matters. A measured assessment does not return one number. It returns a map: which queues can carry autonomy now, which can carry assistance, and which should not be touched yet. The single headline number is the estimate’s habit, and it is the habit to break.
Across the estates we have assessed, the measured potential came back lower than the modeled estimates it replaced. Lower than the market’s claims, and lower than our own earlier models. I regard that as evidence that the method works, not against it.
A concession because measurement has failure modes of its own. An evaluation that saw information production will not have produces a flattering, false number, and we have caught exactly that in our own work. A measured figure also moves as the estate and the models change, so it carries a date and a version on its face, or it is not a measured figure.
But a lower, dated, defensible number does something no borrowed benchmark can. It survives contact with your CFO, your risk team, and the third quarter of the program. Programs sized on borrowed numbers stall when reality arrives, and the automation team inherits the credibility damage the estimate earned.
If you are sizing an agentic program right now, the discipline is one sentence. No coverage number enters the business case unless it was derived from your own resolved tickets. When a vendor quotes a rate, the question that matters is: measured on which tickets, whose estate, and when?
Open your current automation business case and sort its numbers into two piles: derived from our own ticket data and arrived on a slide. Most cases I review are one pile.
So, the question: has any vendor ever shown you a potential figure measured on your own tickets before asking you to sign? If yes, I would like to hear how it held up. If no, ask yourself what that says about the numbers you were shown instead.
Next week: signal is the floor, not a rung. Why nearly every operations maturity model in the market is drawn wrong.
Sanjesh Rao leads product strategy and innovation for Hexaware’s Agentic ITOps platform.
Estimates category volume, trusts field labels, and scores what humans historically did with tickets. Each choice inflates the result: categories describe what a ticket is about rather than what should happen to it; a populated field does not prove the property it names; and scoring history rebuilds the current operation, including the ceiling.
Take 50 resolved tickets from one category and ask whether the resolutions converged. If 50 tickets were resolved 3 ways, that is automation. If they were resolved 50 different ways, the category is not automatable no matter how large it is. Volume tells you where the effort is. Convergence tells you where the automation is.
A measured figure is derived from the customer’s own resolved tickets, carries a date and a version on its face, and returns a per-queue map of where autonomy and assistance can be applied rather than a single headline number. The article’s discipline: no coverage number enters a business case unless it was derived from your own resolved tickets.