Why do most AI ROI numbers fall apart?
Three recurring failures. Company level averages hide which workflows pay and which leak, because nobody can attribute money to a single system. Hours saved get booked as money saved, although cutting a task from twenty minutes to five usually redirects the remaining work rather than removing anyone’s salary. Costs get undercounted too: licences and usage fees are visible, while building, reviewing outputs, and maintenance hide in team time. Hours saved measure activity; money saved requires the workflow itself to change.
What does honest measurement look like?
Pick one workflow with a countable unit: tickets resolved, documents processed, invoices matched. Record its cost and cycle time before launch, then compare after, counting every input: development spend, run costs, review time, and the oversight work. Judge the result per unit, not in totals: an automation that resolves each ticket for a fraction of the handled cost is working even if monthly spend rose. Review on a fixed cadence with agreed keep or stop rules, because value decays as models and usage shift. The discipline pairs naturally with checking AI readiness before building and weighing build vs buy for AI systems before scoping; our notes on token cost attribution show what the cost side looks like in practice.