An agent in your company quotes a price, issues a refund, sends an email or approves a request, and it is wrong. The first question in the room is usually whose fault it is, and that is the wrong question. Accountability is not assigned after the incident. It is fixed by five things that were either in place before the agent ran or were not: an identity, a supervisor, a spend limit, a management system, and an insurance policy. A Canadian tribunal answered the legal half of it in 2024 with ordinary negligence law and a small claims file worth CAD 812. The other half is work you control, and this post covers what each of the five actually buys you.
A tribunal answered this in 2024, and Air Canada paid
On 14 February 2024 the British Columbia Civil Resolution Tribunal decided Moffatt v. Air Canada, 2024 BCCRT 149. A passenger asked Air Canada’s support chatbot about bereavement fares, the chatbot described a retroactive refund process that did not match the airline’s own policy page, and the passenger booked on the strength of it.
Air Canada’s defense is the part worth reading. Tribunal member Christopher C. Rivers summarized it like this: “In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission.” He then disposed of it in two sentences: the airline is responsible for every piece of information on its website, and it makes no difference whether that information came from a static page or a chatbot.
The finding was ordinary negligent misrepresentation, not novel AI law: “I find Air Canada did not take reasonable care to ensure its chatbot was accurate.” The airline was ordered to pay CAD 812.02, made up of CAD 650.88 in damages, CAD 36.14 in pre-judgment interest and CAD 125 in tribunal fees.
The award is trivial and the principle is not. The decision binds Air Canada and nobody else, which is what makes it worth reading: no new AI doctrine was needed to reach it. Your company published the agent, and reasonable care is measured by what you put around it.
The evidence test: courts will keep reaching for existing doctrine, because it fits. The question worth your time is what evidence they could produce that reasonable care was taken, and that evidence has to exist before the incident.
An agent needs a name before it needs permissions
Before anyone can say who did something, the system has to record what did it. An agent running on a shared service account, a developer’s personal API key, or a long-lived token inherited from a pilot produces actions with no owner. That is a non-human identity problem, and it is now the majority case rather than the edge case.
Palo Alto Networks’ 2026 Identity Security Landscape, a survey of 2,930 global cybersecurity leaders, states it directly: “Machine identities, including AI agents, now outnumber human identities 109:1.” The same research reports that 96% of respondents see human identities operating with access far beyond what their roles require, and that the same report expects the population of AI agents to grow by 85% this year.
Read those three findings in order and the shape of the problem is obvious. The population that already carries too much access is the smaller one, it is outnumbered a hundred to one by machine identities, and the part of that machine population which takes actions rather than moving files is the part expected to grow 85% this year. An agent with its own identity, its own scoped credentials and a named human owner turns “something charged the customer twice” into a log line with a subject. Without it, your incident review starts with archaeology.
The practical bar is low and most companies still fail it: every agent has its own credential, that credential is short-lived, its permissions are scoped to the task rather than to the team, and a named person in your org chart owns it.
Something has to be watching while the agent works
Oversight after the fact is an audit. Oversight during the run is a control, and in the EU it is about to carry a date. Article 14 of the EU AI Act requires high risk systems to be designed so they “can be effectively overseen by natural persons during the period in which they are in use”. It is not in force yet: the AI Act Explorer records Article 14 coming into force on 2 December 2027 for high risk systems under Annex III and 2 August 2028 for those under Annex I, under Article 113(c). Note the wording of the obligation. Not reviewable afterward, overseen while in use.
In practice nobody watches every step by hand, so the supervision is itself software: a guardian agent, a cheaper and narrower model or rule set that inspects what the working agent is about to do and stops it. The pattern is now a first class object in the major SDKs. OpenAI’s Agents SDK documents it as guardrails: checks and validations run on the user’s input and on the agent’s output.
The same documentation is honest about the trade-off, which is the part that matters: “Blocking execution guarantees that the expensive model does not start; with parallel execution, the expensive model may already have started before the guardrail completes.” Blocking costs latency on every request. Parallel keeps the experience fast and accepts that the thing you are trying to prevent may already be underway. That is a business decision about which failures you can live with, and it should be made by the person who owns the process, not by whoever wired the agent up.
A guardian agent does not replace the person. It makes their oversight tractable at volume, by escalating the small number of cases that need a human instead of asking for approval on everything.
Most agent incidents arrive as an invoice
The failure a mid-market company is most likely to meet this year shows up on a statement rather than in a hearing: an agent that loops, retries and spends. A runaway agent is a cost event before it is anything else, and it is the one failure mode that arrives with no incident report attached.
AI FinOps is the discipline that catches it: attributing token and inference spend to the agent and the process that caused it, then setting budgets and alerts against that attribution. The FinOps Foundation’s State of FinOps 2026, drawn from 1,192 respondents representing more than $83bn in annual cloud spend, records how fast this became standard practice: 98% of respondents now manage AI spend, up from 31% two years earlier.
Two years ago fewer than a third of those teams tracked AI spend at all. A jump that steep usually means a lot of people opened a bill they had not modelled. A per-agent spend cap, an anomaly alert on cost rather than on errors, and cost attribution granular enough to name the offending workflow are cheap to add before launch and awkward to retrofit after a bill arrives.
The paperwork your customer will eventually ask for
The three controls above are yours to design. The fourth is the one you hand to somebody else, because at some point a customer’s procurement team, an insurer or a regulator will ask what your AI governance actually is, and “we are careful” is not an answer that survives a questionnaire.
ISO 42001 is the certifiable answer. Microsoft’s compliance documentation describes it as the international standard for setting up, running and improving an AI management system inside an organization, and records what certification means in practice: Microsoft’s own AI systems go through regular independent third-party audits against it.
Certification is not the point for most mid-sized companies, and chasing it early is a good way to spend a year on documentation instead of delivery. What is useful earlier is the shape it asks for. The same page defines an AI management system as the set of organizational elements that establish policies and objectives, and the processes to achieve them, for the responsible development, provision or use of AI systems. Strip the formality and it asks one question: are your AI decisions made by a process somebody owns, or by whoever happened to be in the room? Answer that with documents rather than opinions and you are most of the way to the evidence a tribunal called reasonable care, whether or not you ever pay for an audit.
What insurance carries, and what it does not
The last control is financial, and it is newer than the others. AI liability insurance has moved from a theoretical product to one you can buy. Munich Re’s aiSure line covers both sides of the relationship, offering AI providers performance warranties that let them indemnify their clients for financial losses or legal liabilities directly caused by AI errors, and separately covering corporations deploying AI at scale against losses arising from AI errors.
Two things follow for a buyer. First, if you are the one selling something AI-assisted, a backed warranty is now a competitive instrument and not just a risk transfer. Second, if you are the one deploying, the existence of a dedicated product is itself the signal: read your general liability and professional indemnity wording before assuming AI error already sits inside it, and ask your broker where the line falls.
Insurance pays for outcomes. It does not produce the evidence of reasonable care, and an underwriter will ask for exactly the four things above before quoting. It is the backstop, not the plan.
The five questions we ask before an agent goes live
We run the same short list on every agent we put into production for a client, whichever framework it is built on:
- What identity does it act under, and who in the org chart owns that identity? If the answer names a team rather than a person, it is not finished.
- What stops it mid-task, and what is the cost of that check? Blocking or parallel is a business call, so make it explicitly and write down which one you chose.
- What does one unit of its work cost, and what happens at ten times that? A cap that nobody has tested is a hope.
- What would we hand a customer who asks how this is governed? If the honest answer is a conversation, write the document.
- Who pays if it is wrong, and have we read that policy this year?
None of these is exotic and none of them requires a governance program. They are the difference between an incident with an owner and an incident with an argument. If you are deciding whether to build an agent on a platform or from parts, the same questions apply on either side, and we wrote about where that line falls in Copilot Studio or a custom agent.
Reasonable care is a design decision
Air Canada did not lose because a chatbot made a mistake. It lost because it could not show it had taken reasonable care that the chatbot was accurate, and then argued the software was somebody else. Every control in this post exists to make sure that argument never has to be made: an identity so the action has a subject, a supervisor so the action can be stopped, a spend limit so the damage has a ceiling, a management system so the care is documented, and a policy so somebody pays. Decide all five before the agent runs, because afterwards you are not designing anything, you are explaining.
Sources
- Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal, 14 February 2024
- 2026 Identity Security Landscape, Palo Alto Networks, survey of 2,930 cybersecurity leaders
- Article 14: Human Oversight, EU AI Act Explorer
- Guardrails, OpenAI Agents SDK documentation
- State of FinOps 2026, FinOps Foundation
- ISO/IEC 42001:2023 Artificial intelligence management system, Microsoft Learn
- aiSure, Munich Re