Agents stopped answering and started acting
Not long ago, an AI system produced text for a person to read. Now it holds a tool, an API key, and a budget.
It files the ticket, changes the number, approves the document, routes the payment, or rebalances the position. Its output is no longer merely read by a person. It is consumed by another system and acted on immediately.
The shift is uneven, but the direction is clear. It is moving fastest where the consequences are financial.
An agent that drafts a bad email wastes a minute. An agent that misreads a figure and acts on it moves money.
Every action rests on a claim
Beneath every one of those actions is a factual claim.
This invoice totals 4,812.
This clause permits early termination.
This address appeared in three previous chargebacks.
This counterparty is investment grade.
The action is sound only if the claim beneath it is sound. Yet the claim often arrives without evidence, without a challenge process, and without anyone standing behind it.
By the time an error is discovered, the action has already been taken. The failure does not surface where it was made. It appears downstream as a reconciliation break, a bad payment, a complaint, or a loss, and often to someone who had no idea an AI system was involved.
The cost of being wrong is real, and it is already being paid.
It is simply paid by someone other than the system that was wrong.
Companies already pay people to catch this
This is not a hypothetical, and the industry has noticed.
Teams deploying agents already assign engineers and domain reviewers to check whether their outputs hold up. Their job is to read AI answers, compare them with the underlying information, and decide whether they are safe to use.
That approach works. But it has three structural limits.
Its cost scales with total volume rather than error volume: every correct answer still has to be reviewed, even though most answers will be fine.
It operates at human speed, so the check often arrives after the action it was meant to protect.
And the reviewer is appointed and paid by the same organization that deployed the system. There is no adversary in the process, no independent party that earns anything by finding a mistake.
A score cannot catch a live error
The other answer on offer is the score.
Scores answer a real question: should we trust this system in general?
That question matters. But it is asked before the fact, about the system as a whole, across many outputs, on a schedule.
Failures do not happen to systems as a whole. They happen one output at a time.
A system with an excellent evaluation grade can still produce the specific wrong claim that costs you money. Its score cannot tell you which claim failed, when it failed, or who is accountable for the result.
The score was never a verdict on your output.
A score and a verdict are different objects.
A score summarizes general performance. A verdict resolves a specific claim, between identifiable parties, with a consequence attached.
Production needs both. Most of today’s infrastructure serves the first. The second is still largely missing.
What the missing system looks like
What is missing is not another score for the system as a whole. It is a way to make each claim reviewable the moment it is published, before it travels through the system and triggers an action that cannot be undone.
The party publishing the claim must stand behind its accuracy. . If the claim is proven wrong, the resulting cost should remain with the publisher, not be passed on to the people or systems that relied on it.
But accountability on the publishing side is only half of the mechanism. Independent parties must also have a reason to scrutinize claims and identify errors. A challenge that proves correct should be rewarded; one that fails should carry a cost. Without both incentives, claims will either go unexamined or be challenged without sufficient grounds.
Once a claim is challenged, an independent resolution layer must assess the competing positions and issue a verdict that other systems can rely on.
This changes the life cycle of an AI output.
A claim is no longer simply generated, accepted, and passed forward. It is published with accountability attached, kept open to challenge, and independently resolved when contested, before an error becomes someone else’s loss.
Where Overlap fits
Overlap receives claims and resolves them through a network of independent validators.
A claim may originate from a user, an application, or an agent. It may be submitted directly for assessment, or arise from an agent response that another agent challenges as incorrect.
Through its verification layer, Overlap does not try to determine whether an agent is trustworthy in general. It answers the narrower question that matters in production:
Can this particular claim be relied on?
To answer it, Overlap submits the claim to its resolution process, where a network of independent validators examines the available evidence and establishes a verdict. Their incentives are tied directly to the accuracy of that verdict: their stake, rewards, and future earning potential all depend on the correctness of their decisions.
Find out more about the mechanism and how to join the validator network at docs.overlap.fi.
Overlap: the resolution layer for AI claims.