Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Hugging Face introduces ProvenanceGuard, a verification layer for MCP-based LLM agents that tracks which tool output supports each claim to prevent cross-source conflation errors in data-sensitive settings.
Most factuality checkers for LLM answers pool evidence and verify whether a claim is supported anywhere in the combined context, which can mask incorrect source attributions. ProvenanceGuard instead preserves the identity of each MCP tool output and checks whether the specific source cited in the answer actually supports the claim. This prevents errors where a fact from one source is presented as coming from another, such as a patient record detail misattributed to medical literature.
The system operates as a post-generation layer that processes the captured MCP trace, including tool outputs and source IDs, without retraining the agent. It decomposes the answer into claims, identifies the relevant source for each, verifies support, compares the source to the one named in the answer, and issues per-claim verdicts and a global allow or block decision. Blocked answers can undergo repair and re-verification.
In experiments using local models, ProvenanceGuard evaluated 281 real traces from a medical agent and tested 361 claims from 40 held-out answers. Experts flagged 139 claims as unsupported, and ProvenanceGuard blocked 138 of them while allowing one through. It also held 67 supported claims for review or repair, reflecting a cautious policy prioritizing second looks over false positives.
ProvenanceGuard achieved an F1 score of 0.846 for blocking decisions in a harder test with similar sources and correctly identified the exact source in 50.3% of claims. It detected all 50 cases of wrong attribution in a controlled swap test, demonstrating sensitivity to source errors. When integrated with a RARR-style repair loop, it resolved all 173 initially blocked answers, though 144 ended in fallback text to avoid unverifiable rewrites.