Conformal Factuality Control for Multi-Hop Retrieval-Augmented Generation
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Conformal Factuality Control for Multi-Hop Retrieval-Augmented Generation
Abstract:Retrieval-augmented generation (RAG) can ground large language models in external evidence, but retrieved context does not guarantee that generated claims are factually supported. This problem is especially relevant in multi-hop RAG, where retrieval and reasoning proceed through multiple dependent stages. We study whether claim-level conformal factuality control, previously developed for RAG, remains effective in this setting. We apply split-conformal claim filtering to multi-hop RAG and evaluate it on HotpotQA, Natural Questions, and TriviaQA using Llama 3.1 8B and GPT-4o-mini, together with a single-hop reference experiment. Across all six multi-hop model-dataset configurations, increasingly stringent conformal targets consistently increase the fraction of responses whose retained claims are fully supported. At the 95% target, this rate ranges from 95.80% to 97.20%, compared with 55.60%-76.03% without filtering. However, the improvement is strongly selective: only 4.41%-31.09% of generated claims are retained and 9.70%-51.40% of responses remain non-empty at the 95% target. These results show that conformal factuality extends to multi-hop RAG, while demonstrating that nominal reliability must be interpreted jointly with claim retention and abstention.
| Comments: | 15 pages, 2 figures. Code available at this https URL |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2609.38222 [cs.CL] |
| (or arXiv:2609.38222v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38222
arXiv-issued DOI via DataCite
|
Submission history
From: Muhammad Aimal Rehman [view email][v1] Sun, 27 Sep 2026 23:55:03 UTC (670 KB)
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
Current browse context:
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
Large Language Models are Approximate Survival Estimators
Oct 1
-
TomasuLLM: Out-of-Order Speculative Execution for LLM Agents
Oct 1
-
Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling
Oct 1
-
The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
Oct 1
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.