UpShaqo
Intelligence desk
Agent Trust & Security Source-backed analysis

The Grounding Illusion: Why Correct Facts Still Cite the Wrong Source

A new verification method exposes a quieter failure than hallucination: MCP agents that state true facts while pointing to the wrong tool output entirely, a gap most faithfulness scores never catch.

UpShaqo Editorial IntelligenceSeptember 29, 20266 min read
Intelligence standard

Independent UpShaqo analysis built from fresh, attributed sources. We explain the impact instead of repeating the announcement.

Read for leverage: focus on the workflow change, the customer problem, and the next action—not only the product announcement.

Picture a support agent telling a customer: "According to the account record, this plan includes a 30-day refund window." The refund window is real. It just doesn't live in the account record — it lives in a separate policy document the agent also had open. Nothing in that sentence is false. The citation is. Researchers at Multiverse Computing call this cross-source conflation, and their new paper argues it's the failure mode most teams building tool-using agents haven't priced in yet.

The Comfortable Assumption

Here's the obvious read on agent reliability in 2026: teams running Model Context Protocol (MCP) agents — the ones that call search tools, query databases, and pull structured records mid-conversation — assume that if a faithfulness checker gives a green light, the answer is grounded. RAGAS, MiniCheck, AlignScore, SummaC: these tools ask whether a claim is supported by the evidence the agent retrieved, pooled together. If the fact shows up anywhere in that pool, the claim passes.

For a single-passage RAG system, that's a reasonable proxy for trustworthy. For a multi-tool MCP agent juggling patient records, research articles, and account databases in the same turn, it's a much weaker guarantee than it looks. The research packet from Multiverse Computing makes the case plainly: a source-blind verifier sees support in the pooled evidence and passes the claim, even when the answer attributes it to a source that never said it.

Why "Somewhere in the Evidence" Isn't the Same as "Right There"

The clinical example in the paper sharpens the stakes. A patient-specific medication detail pulled from a patient-history tool becomes misleading the instant an answer presents it as a finding from general medical literature. The fact is accurate. The provenance is fabricated. In a data-sensitive setting — healthcare, finance, account management — a wrong attribution can do as much damage as a wrong fact, because the citation is often what a human reviewer or downstream system actually trusts.

This is the contrarian turn: the industry's working definition of "grounded" has quietly split into two different things. One means supported-by-any-tool-output. The other means supported-by-the-source-the-answer-named. Most eval pipelines today answer the first question and assume it settles the second.

What ProvenanceGuard Actually Checks

The paper's proposed fix, ProvenanceGuard, is a post-generation verification layer that sits on top of a black-box MCP agent without retraining it. It reads the captured MCP trace — tool outputs plus their source IDs — and refuses to pool the evidence into one anonymous context. Instead it runs five steps in sequence: break the answer into claims, find the source most relevant to each one, check whether that source actually supports the claim, compare that source against the one the answer names or implies, then emit both a per-claim source verdict and a global allow-or-block decision.

The results, tested on 281 real traces from a medical agent using patient records and research tools, are notable less for raw accuracy and more for what they measure. On 361 held-out claims checked by human experts, ProvenanceGuard caught 138 of 139 claims experts said should not pass, letting only one through. It also held back 67 claims experts considered supported, routing them for review — a conservative bias the researchers describe as intentional for data-sensitive review, where getting the source right matters more than speed. Against four other checkers on the same claims, ProvenanceGuard scored 0.802 on the paper's reject/block F1 measure, ahead of MiniCheck (0.783), RAGAS Faithfulness (0.758), AlignScore (0.662), and SummaC-ZS (0.436) — and it was the only one of the five that emitted a claim-to-source ID at all.

Where the Method Still Struggles

A fair contrarian analysis has to sit with the paper's own hard test, not just its favorable one. When the researchers ran ProvenanceGuard against traces with several similar sources — the realistic case where two documents plausibly could have produced the same claim — the block-decision F1 held at 0.846, but exact source identification dropped to 50.3%. In other words, the system stays reliable at deciding whether to trust a claim, but telling which of several look-alike sources actually produced it is meaningfully harder. A separate controlled test, where researchers deliberately swapped the named source in 50 cases while leaving the underlying evidence intact, found ProvenanceGuard caught all 50 swaps — proof it can spot a clean attribution error, even as the murkier many-plausible-sources case remains unresolved.

That gap matters for anyone evaluating this class of tool: it's strong at catching obvious misattribution and weaker at disambiguating near-identical sources, which is precisely the scenario a large enterprise knowledge base tends to produce.

The Repair Question: Blocking Isn't the Whole Job

A verifier that only says "no" is a liability generator if there's no next step. Wired to a RARR-style repair loop, the full-trace run resolved all 173 blocked answers — but 144 of them ended in fallback text rather than a substantive rewrite. Read charitably, that's the system declining to manufacture an answer it can't verify, which is the safer failure mode for a support or clinical context. Read skeptically, it means over 80% of blocked answers in that run didn't get fixed; they got withdrawn. Teams adopting this pattern should expect a real tradeoff between answer availability and answer trustworthiness, not a free upgrade to both.

The Business Consequence Founders Should Plan Around

For a company shipping an MCP-based support, ops, or clinical agent, this research reframes the release gate question. The practical ask, as one commenter on the original post put it, is whether your eval means supported-by-any-tool-output or supported-by-the-source-the-answer-named — because those are different release gates, and only the second one catches cross-source conflation before a customer, auditor, or clinician trusts the citation.

As analysis, not fact from the paper: teams with agents that touch more than two distinct source types — records, documents, external APIs — are the ones most exposed to this failure, since conflation risk scales with the number of similar-looking sources an agent can draw from in a single answer.

What to Do With This Before It Bites

Three concrete moves follow directly from the research rather than from speculation:

  • Audit whether your current faithfulness or RAG evaluation tooling checks attribution at all, or only pooled support — most named tools in the paper's comparison do not.
  • If your agent operates across genuinely similar sources (multiple policy versions, near-duplicate records), budget extra review time; the paper's own stress test shows source identification accuracy drops sharply in that regime even when block decisions stay reliable.
  • Decide in advance whether a blocked answer should trigger a rewrite attempt or a safe fallback — the repair-loop data suggests most blocked answers in a conservative setup end up as fallbacks, not fixes, and that's a product decision, not just an engineering one.

The method has already been adapted once outside its original medical setting, folded into an optional grounding-verification stage for a finance agent that checks completed answers against SEC excerpts without altering the underlying rollout. That's the more interesting signal than the benchmark numbers: source-aware verification is being treated as a bolt-on layer, not a rebuild, which is the realistic path for most teams weighing whether to add it.

Sources

Hugging Face — "Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents"

#MCP agents#source attribution#AI verification#enterprise AI trust#hallucination detection#RAG evaluation

Two doors. Pick one.

Hire the team.
Or become it.