Agentic AI in the enterprise: pilots to production
The gap between a convincing demo and a system you can put in front of a regulator.

Agent demos are easy to build and hard to trust. A pilot shows that a model can complete a task. Production asks a harder question: can it complete the task reliably, explain what it did, and fail in a way the business can absorb.
The pilot-to-production gap is mostly not about the model
Teams that stall rarely stall on model quality. They stall on the surrounding machinery: no evaluation set, no way to reproduce a bad run, no boundary on what the agent is permitted to touch, and no answer when someone asks why it did what it did last Tuesday.
- Evaluation: a fixed set of real cases with known-good outcomes, run on every change.
- Observability: every tool call, input and output retained long enough to reconstruct a decision.
- Permissions: the agent holds its own narrow credentials, never a human's.
- Escalation: an explicit, tested path to a person, used by default when confidence is low.
An agent that cannot explain a decision has not automated the work. It has moved the work to whoever has to audit it.
Choose workflows where being wrong is cheap and visible
The best first candidates share a shape: high volume, structured inputs, a human already reviewing the output, and an error that surfaces quickly. Triage and classification fit. Drafting fits, because a person edits before anything leaves the building.
Workflows where errors are silent and compound are the worst possible starting point, however appealing the time savings look on a slide.
Write down what the agent may not do
Capability lists grow on their own. Constraint lists have to be written deliberately. Useful ones are specific: no irreversible action without confirmation, no spend above a stated threshold, no access to systems outside the named set, no outbound message to a customer without review.
Encoding those as enforced permissions rather than prompt instructions is the difference between a policy and a wish.
Keep a human in the loop on purpose, not as a fallback
Review that exists only because the team does not yet trust the system quietly decays. People approve by reflex once the output has looked fine for a few weeks. Review that samples deliberately, including cases the agent was confident about, keeps working.
What to measure
- Rate of escalation to a person, and whether it is falling for good reasons.
- Time from a reported bad output to a reproduced bad output.
- Share of actions a reviewer would have taken differently, sampled rather than self-reported.
- Cost per completed task, including the review time the business still pays for.
A system that scores well on these is boring to operate, which is the goal.
Working through this on a live estate?
Talk to an expert

