We've been running multi-agent systems in production for about eighteen months now. Not proofs of concept — actual production systems that process real data, take real actions, and have real consequences when they go wrong. We've learned things that don't appear in the research papers because the research papers mostly evaluate agents on benchmarks rather than on what it's like to operate them on a Tuesday afternoon when something unexpected happens.
Here's what eighteen months of this has actually taught us.
Orchestration is harder than the individual agents
Building a capable individual agent for a well-defined task is achievable. The interesting engineering challenge is what happens when you wire multiple agents together. How does the output of agent A get passed to agent B in a format agent B can actually use? What happens when agent A is confident but wrong, and agent B makes a consequential decision based on that confidence? Who in your organisation is accountable for the output of a pipeline where four agents each made a reasonable decision but the combination of them produced a bad result?
We've settled on two orchestration patterns that we use in almost everything now. The first is a linear pipeline with validation gates: each agent's output passes through a validation check before the next agent uses it, and validation failures short-circuit to a human rather than propagating through the system. The second is a parallel-with-reconciliation pattern: two or more agents process the same input independently and their outputs are compared before any action is taken. Disagreements go to a human. Agreement at sufficient confidence proceeds autonomously.
The failure modes nobody tells you about
The most common failure mode in production multi-agent systems isn't individual agent errors — it's handoff failures. Agent A returns a result that's technically correct but formatted or structured in a way that causes agent B to misinterpret it. This sounds like an engineering problem that should be solved once and forgotten, but in practice the real-world messiness of input data means you keep encountering novel formats and structures. Robust handoff logging — recording exactly what was passed between agents and in what format — is how you debug these efficiently instead of spending four hours wondering why the output is wrong.
Practical agentic AI insights for UAE and GCC leaders — no spam, just what's actually working in production.
The second failure mode is confidence calibration drift. An agent that was well-calibrated at launch — meaning its stated confidence levels accurately predicted its accuracy — can drift over time as the underlying data distribution shifts. We've seen agents that were 92% confident and 90% accurate at launch drift to 92% confident and 78% accurate six months later, because the real-world inputs changed in ways the training data hadn't anticipated. Monitoring confidence calibration continuously, not just output accuracy, is something we now build into every deployment from day one.
The accountability question is organisational, not technical
After eighteen months, I'm convinced that the hardest problems in multi-agent AI aren't technical — they're organisational. When a four-agent pipeline produces an output that turns out to be wrong, someone needs to be accountable for that. In the organisations we work with that handle this well, there's a named human who owns the entire pipeline's output, who reviews the audit trail of agent decisions regularly, and who has the authority and the tools to intervene when something looks wrong. In the organisations that handle it badly, the pipeline is treated as a black box and accountability is diffuse — nobody really owns it, which means nobody catches the slow drift problems until they've become big problems.
What we'd do differently if starting again
Start the audit trail design before you write any agent code. Not as an afterthought — as the first design constraint. Every agent handoff, every tool call, every confidence score, every human escalation needs to be recorded in a structured format that someone outside the engineering team can actually query. We've retrofitted this onto early deployments and it's genuinely difficult. The systems that were designed around this from the start have been dramatically easier to operate, debug, and explain to the stakeholders who need to trust them.