Agentic AI in Production (Evaluation & Governance)
Date Published
Deployed a LangGraph/CrewAI multi-agent system for internal insurance claims document-processing automation (routing, extraction, summarisation).
PoC just proved agents could chat. For Prod, I added strict timeouts (5 secs per agent), token-cost budgeting (hard stop if cost/query exceeded £0.02), and deterministic fallbacks (if agent confidence <90%, escalate to human, no hallucinated outputs allowed).
Ran a 4-week shadow-mode comparing agent decisions against historical human decisions. Tracked precision (94%), recall (89%), and human-escalation rate (target <15%).
Governance Framework:
Every tool-call and reasoning trace logged with a unique UUID, fed into a separate "Guardrail Agent" that flagged policy violations (e.g., PII leakage) in real-time.
Weekly review board with Compliance and Legal—they held veto power over any new tool permissions.
Automated rollback trigger: if human-override rate spiked >20% within an hour, the system auto-switched to read-only mode and alerted the on-call engineer.
Result: Processed 1,000+ documents autonomously in production over 2 months with zero compliance breaches.