The KPIs I Give to Leaders Who Ask: ‘How Do We Know If AI’s Working?’
Frequently Asked Questions
Q: What are the most common mistakes in how organisations measure AI performance?
The most common mistake is measuring AI capability rather than AI impact — tracking how accurate the model is, how many queries it handles, or how fast it responds, rather than whether the decisions, products, or experiences it contributes to are actually better. Capability metrics are important for engineering teams; they are the wrong primary metric for executive oversight. The second common mistake is measuring only the benefits while ignoring the costs — the hidden costs of AI deployment include oversight labour, error correction, technical debt, and the organisational capability that atrophies when AI substitutes for human judgment.
Q: What KPIs actually tell you whether AI is working?
Decision quality improvement — measured by comparing outcomes in decisions made with AI assistance versus comparable decisions made without it, over sufficient time to capture tail-risk outcomes as well as typical ones. Human capability development — whether the people working alongside AI are developing judgment, or whether they are becoming dependent and less capable over time. Trust calibration — whether the people using the AI system have appropriately calibrated confidence in it (not over-trusting in domains where it fails, not under-using in domains where it performs well). And exception rate — how frequently the AI requires human intervention, which is the most direct measure of where the system is operating outside its reliable range.
Q: How does the HUMAND framework inform AI performance measurement?
HUMAND provides the prior question: what is this AI system supposed to be doing that humans, machines, or combinations could not do as well? The answer to that question defines the right measurement framework. If the AI is supposed to be freeing human time for higher-judgment work, the KPI is the quality of the higher-judgment work — not the volume of tasks the AI processed. If it is supposed to be improving decision accuracy in a specific domain, the KPI is decision outcome quality in that domain — not the model’s internal accuracy metrics.
Q: Can Morris Misel run a workshop on AI governance, performance measurement, and the HUMAND framework for our executive team?
Yes. AI governance and performance measurement are core workshop topics for executive teams, boards, and technology functions. Book at morrismisel.com.