Most agentic AI programmes stall somewhere unexpected. Not on the model, not on the orchestration, not on the interface — on the question of whether the customer in this record and the customer in that one are the same person. It is the oldest problem in enterprise data, it is profoundly unfashionable, and it has quietly become the thing that decides whether an agent can be allowed to do anything consequential.

Why it matters more now than it did
For twenty years, bad entity resolution produced bad reports. A duplicate customer inflated a count, a mismatched record skewed a segment, somebody noticed at quarter end and applied a correction. The cost was wrong information, absorbed by humans who knew to be sceptical of the dashboard.
An agent doesn’t produce a report. It acts. It applies the credit to an account, sends the notice to an address, closes the case, releases the payment. When the resolution is wrong, the output isn’t a misleading chart — it’s a correct action performed on the wrong entity, executed confidently and at machine speed, and discovered later by the party it was done to.
That is a different class of failure, and it’s why a data problem that organisations tolerated for two decades suddenly gates everything.
The failure mode is plausibility
What makes this dangerous is that nothing looks broken. The agent behaves impeccably: it reasons well, cites its sources, follows policy, produces a defensible trace. Every step is right. The subject is wrong.
And because the reasoning is sound, review tends to pass it. A human checking the work sees a well-argued decision and approves it, because the thing that was wrong happened upstream of everything they’re looking at. Resolution errors are the hardest errors to catch in review, precisely because they corrupt the premise rather than the logic.
The model cannot fix this, and shouldn’t try
It is tempting to point a capable model at the problem — it’s good at fuzzy matching, and it will produce confident, plausible answers about which records refer to the same thing. This is exactly the wrong application of a model, by the principle I’d apply anywhere: put the model where being wrong is cheap and recoverable. A wrong resolution is neither. It is expensive, it propagates silently into every downstream action, and it is unauditable after the fact if nothing recorded why the match was made.
The model has a real role here — but it is to propose, never to conclude. It can surface candidate matches, rank them, explain the evidence. The decision that two records are the same entity should be made deterministically where the keys allow it, and by a named human where they don’t, with that confirmation stored.
What good looks like
Deterministic keys wherever they exist, used in preference to cleverness. A government identifier, an account number, a verified email — unglamorous, exact, and defensible in a way probability never is.
A resolution graph as a durable artifact, not a step in a pipeline. Which records are joined, on what evidence, confirmed by whom and when. This is the thing that accumulates value: every confirmed match is permanent, reusable, and compounding — and unlike a model’s output, it gets more valuable as it grows.
Confirmation with a name attached. If a human adjudicated an ambiguous match, that person’s identity belongs in the record. When a decision is challenged two years later, “the system matched them” is not an answer; “confirmed by this person on this date, on this evidence” is.
An explicit unresolved state. The most dangerous design choice is forcing a match. Agents must be able to encounter “I cannot establish that these are the same entity” and stop — routing to a person rather than guessing. Unresolved is a valid, safe outcome. Silently resolving to the most likely candidate is not.
The organisational problem underneath
The technical pattern isn’t controversial. The reason it doesn’t get built is that entity resolution has no natural owner. It isn’t a feature, so no product manager wants it. It isn’t a model, so it attracts no excitement. It sits between systems, so every team considers it someone else’s prerequisite. It appears in no roadmap and blocks every roadmap.
Which makes it a leadership problem rather than an engineering one. Somebody has to fund unglamorous work with no demo, before the agent programme that depends on it gets approved — and has to keep funding it when a more exciting proposal arrives.
The question worth asking
Before approving the next agent initiative, one question tells you more than any architecture review: if this agent acted on the wrong customer tomorrow, how would we find out — and how long would it take?
If the honest answer is “when they complain,” the model was never the risk.
A companion to the Agentic AI series.
Related: The Substrate War — Will Your Data Platform Become Your System of Record? · Put the Model Where Being Wrong Is Cheap.




Leave a Reply