AI Agents

Why AI Agents Fail in Production (And Why a Better Model Won't Save Them)

Why most AI agents stall between pilot and production, the five failure patterns behind it – and what it takes to get beyond 90% accuracy.

Dougal Watt CEO & Co-founder,
Graph Research Labs
Making AI Agents Work · Part 1 of 3
October 2026 · 5 min read
The Problem
In BriefAI agents usually fail in production because they do not understand the organisation they work in, not because the model lacks capability. The Five Agent Failure Patterns explain where agent accuracy breaks down – meaning drift, the fragmented view, invented relationships, rules in the prompt and no memory of why. The fix is a shared layer of meaning – a knowledge graph built on an ontology – combined with governance over what agents are allowed to do.

Almost every organisation I speak with is building AI agents. The pilots can look impressive, particularly to those less familiar with how agents work – the agent reads a request, calls a few tools and writes a clear, confident answer. That is understandable, as agents are new to most organisations. However many teams discover during the pilot that the agent is right only some of the time – often well short of what production needs.

The natural assumption is that the model is not good enough yet, and that the next release will fix it. In practice, that is rarely the issue.

An AI agent arrives in an organisation knowing an enormous amount about the world from public training data but almost nothing about that organisation. It does not know what the organisation's terms mean, how its systems connect, or which rules apply to which decisions. Every gap in that knowledge gets filled the same way – with the most plausible answer the model can produce.

Plausible is fine in a demo. It is not fine when the agent is approving a payment, answering a regulator or advising a clinician.

The Five Agent Failure Patterns

There are five common patterns that undermine agent accuracy – a framework I call the Five Agent Failure Patterns.

Failure Pattern 01

The first and most common pattern is meaning drift, which is when one word means different things across an organisation's systems, and an AI agent can't tell which meaning applies – so it guesses. Take "active customer". In the CRM, it is a customer contacted in the last twelve months. In billing, it is one with an open invoice. In the data warehouse, it is one with a transaction this quarter. People carry these differences in their heads. An agent cannot tell which meaning is relevant, so it answers by picking the meaning it thinks is most relevant – which might not be the correct one.

Failure Pattern 02

Second is the fragmented view, which is when an AI agent can't see how the data in different systems connects – so it only ever sees part of the picture. An organisation's CRM, billing and support systems might each hold the same customer under different names – duplication is very common. To the people who work there, it might be a single customer. To the agent, it is different customers.

Failure Pattern 03

Third, and perhaps most dangerous, is invented relationships. When the agent needs a connection it cannot find – which subsidiary owns which contract, which role approves which spend – it infers one from what is statistically likely. The result is a relationship that sounds right, reads well and does not exist. It is the most convincing kind of error, because nothing about it looks like an error.

Failure Pattern 04

Fourth is rules in the prompt. An agent is guided by a set of written instructions, known as a prompt. Each time the agent gets something wrong, someone adds another instruction to fix it: treat a customer as active unless their bill is more than 90 days overdue; ignore accounts with no region recorded, except older accounts. Six months later the instructions run to pages of exceptions that nobody can fully check. The organisation's business rules are now buried in text written for a machine – the hardest place to test them, and the hardest place to reuse them.

Failure Pattern 05

The fifth pattern is no memory of why. When an agent makes a decision, the organisation usually cannot reconstruct which facts it used, which rule it applied or which version of the data it saw. There are logs, but logs are not explanations. The first time an auditor or a regulator asks why the agent did something, the project stops.

The problem is not the model. It is that the agent has no shared meaning to work from.

So Will a Better Model Fix It?

No. A better model reasons more fluently over the same missing information. It closes some gaps and guesses more convincingly across the rest. Even an agent that is right 80% of the time, running ten thousand times a day, still gets two thousand things wrong every day – and most of those answers look exactly like a right one. And many organisations are not at 80% accuracy yet, which means there is a lot of work to do just to reach that point. At Graph Research Labs, our aim is to take agents beyond 90% accuracy. We get there by using the Five Agent Failure Patterns to find where accuracy is lost, then our ontology methodology and tools to fix it at the source.

The patterns are not edge cases – we see them in almost every agent that works across more than one or two systems. However they are not limitations of AI either. They are the result of asking an agent to act on a business that has only ever been described in people's heads, in documents and in the structure of individual applications.

The organisations getting agents into production have noticed this. They have stopped treating the agent as the hard part – models and frameworks are becoming interchangeable – and started investing in the layer the agent reads from: a single, shared description of the business, with agreed meaning and connected facts, that every agent uses.

When an agent works from that shared description, each of the five patterns is addressed. Meaning stops drifting because there is one definition. The fragments join up because the connections are explicit. Relationships are looked up rather than invented. Rules move out of the prompt and into something that can be tested. And every decision can be traced back to the facts behind it.

That layer is a knowledge graph, built on an ontology that defines what everything in it means. The rest of this series shows how to build it.

Part 2 – Why AI Agents Need a Knowledge Graph. What knowledge graphs and ontologies are, and why most teams start in the wrong place.

Part 3 – How to Get AI Agents into Production. Five steps from pilot to production, including the AI-Assisted Ontology Modelling Methodology, production testing with GRL Generators and an agent harness that governs every action.

Together, they turn an impressive pilot into an agent an organisation can trust.

Questions

Frequently Asked Questions

Why do AI agents fail in production?

Most AI agents fail because they lack an accurate description of the organisation they work in. The model can reason; it cannot understand a business it has never been told about. Most teams also lack the methods and tools needed to build that description, test it against real data and keep it current as the organisation changes.

Will a better AI model fix agent accuracy?

Not on its own. A better model reasons more fluently over the same missing information. Accuracy improves when agents work from a shared, explicit description of the organisation. Even that is not enough on its own. Agents also need their actions checked against the organisation's rules before they run, and a record of why each decision was made.

Can Graph Research Labs help with AI agent accuracy?

Yes. Graph Research Labs works with organisations in regulated industries to take AI agents from pilot to production, using the organisation's own team. The work combines GRL's Five Agent Failure Patterns methodology, the AI-Assisted Ontology Modelling Methodology and a suite of tools, with the aim of taking agents beyond 90% accuracy and giving organisations governance over their agents.