Guardian agent

BusinessSafety and governancePublished By Simon Budziak

Guardian agent is the category name for software that supervises other AI agents. It sits beside a working agent, inspects each action the agent wants to take, and then allows it, blocks it, redirects it toward something safer, or holds it for a person to decide.

What does a guardian agent actually check?

The useful ones combine fixed rules with judgment. Deterministic policy catches what can be written down, and a model reviews what cannot, which is the same split that makes guardrails work at the edges of a system. The analyst taxonomy that named the category splits it into reviewers, monitors and protectors, and places it inside AI TRiSM as the runtime half of governance. The distinction that matters when a vendor pitches you one is whether it can block an action or only report on it afterward.

Is it a control or a product category?

Both, and the buying question is which you are being sold. A guardian agent that gates actions is doing the work of tool approval with more context; one that only writes dashboards is observability with a new name. Ask what it does when it disagrees with the agent it watches, then check that consequential actions still reach a person through human in the loop rather than being auto-approved by another model.

Frequently asked questions

Is a guardian agent the same as a supervisor agent?

No. A supervisor agent is an orchestration pattern: it plans work and delegates it to workers to get a job done. A guardian agent is a control: it does not own the task, it judges the actions another agent proposes.

Do guardian agents replace human approval?

Not for consequential actions. They reduce how often a person is interrupted by handling the clear cases, which makes the remaining escalations worth reading. Anything touching money, external messages, or production data still needs a named human decision.

Summarize this page with

See this working in a system we built