← All articles

Agentic AI

Agents in regulated work need a place to stop

In tax, finance or healthcare, an agent that is right most of the time is not enough. Someone is accountable for every output.

Most agent demos show an AI completing a task from start to finish with nobody watching. In regulated work that is the one thing you cannot ship. When the output is a tax filing or a compliance decision, a named person is accountable for it. They need to see what happened and why, and they need to be able to defend it months later.

I work on AI for tax compliance. The design problem is rarely "can the agent do it". It is "can a professional trust, check and defend what the agent did". These are the four design choices I come back to.

1. Decide where the human sits

Reviewing everything removes the benefit. Reviewing nothing removes the trust. So I sort each step of the workflow into one of three groups.

Group What the agent does Good fit when
Automate Acts without asking Low risk, easily reversed, high volume
Propose Drafts, and a person approves Judgement is needed, or the action is hard to undo
Escalate Stops and hands over with context The case is outside what it was built for

Gathering documents usually belongs in the first group. Classification and exemption reasoning usually belong in the second. The third group is the one teams forget.

An agent that knows when to stop is worth more than one that always produces an answer. In practice this means designing the stop conditions explicitly: low confidence, conflicting evidence, missing data, or a value above a threshold. If you cannot list the conditions under which your agent hands over, it does not have any.

2. The review screen is the product

If a step is in the "propose" group, the person approving it spends their day in a review queue. That queue is where the product succeeds or fails, and it deserves more design attention than the chat window.

A good review screen shows, on one page:

  • What the agent recommends.
  • The evidence it used, next to the recommendation.
  • Which rule or policy it applied.
  • How confident it is, in words a reviewer understands.
  • One-click approve, edit or reject, with a reason.

The test is simple. If the reviewer has to redo the work to check it, the agent has saved nothing. If they can check it in a tenth of the time, you have a product.

3. Log everything as an audit trail

Every input, action and approval should be recorded in a form that makes sense a year later. In regulated industries the question "why did we do this?" arrives long after the decision, often from someone who was not there.

A useful audit record answers four things: what the agent saw, what it did, who approved it, and which version of the prompt and model was running. That last one is easy to skip and painful to reconstruct.

The audit trail also turns out to be the best debugging tool you have. When quality dips, it tells you whether the inputs changed or the agent did.

4. Treat privacy as a constraint on design

Decide early what data the agent may read, what it may send to a model, and what it may keep. Retrofitting privacy controls after a pilot is slow and painful. It is far cheaper to scope the agent's access narrowly from the first prototype, and widen it only when a use case needs it.

A useful exercise: for each tool the agent can call, write down the worst thing that could happen if it called that tool at the wrong moment. If the answer makes you uncomfortable, that tool belongs behind an approval.

Where to start: exception triage

If I had to pick one first use case for an agent in a regulated workflow, it would be exception triage: sorting the cases that need attention and attaching the context a person needs to resolve them.

It works as a first project for three reasons. The agent does not make the final call, so the risk is contained. The value shows up quickly, because exceptions are where experts lose their time. And the team learns how the agent behaves on real cases before giving it more responsibility.

Autonomy can grow from there, one step at a time, as the evidence supports it. Moving a step from "propose" to "automate" should be a decision someone makes with data, not something that happens because the approvals became routine.

Questions to ask your team this week

  • For each step in the workflow, is it automate, propose or escalate?
  • What are the agent's written stop conditions?
  • How long does a reviewer take to check one output?
  • Could we explain a decision from six months ago using only our logs?

I cover the underlying patterns in the agentic design patterns kit.

Found this useful? Share it, or get the next one through The AI Product Playbook, my LinkedIn newsletter.

The AI Product Playbook

Get new articles by email

Leave your email and I'll send you each new article on AI product management as it is published. You can also follow the newsletter on LinkedIn.