Paper: How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms (2603.08274) Published: 9 Mar 2026. Learn more on Emergent Mind: https://www.emergentmind.com/papers/2603.08274 arXiv: https://arxiv.org/abs/2603.08274 Sign up for our free trending papers email digest: https://www.emergentmind.com/subscribe Follow us on X: https://x.com/EmergentMind Join our Discord: https://discord.gg/BhfTC4mTXq This presentation examines a large-scale study of hallucination in language models performing document question-answering tasks. Using the RIKER framework—a ground-truth-first evaluation methodology—researchers tested 35 open-weight models across 172 billion tokens, varying context lengths up to 200K tokens, temperatures, and hardware platforms. The study reveals that no current model is free from hallucination, with fabrication rates climbing as context grows and surprising disconnects between a model's ability to retrieve correct facts versus its tendency to invent false ones.
The information provided is not trading advice. kdj.com does not assume any responsibility for any investments made based on the information provided in this article. Cryptocurrencies are highly volatile and it is highly recommended that you invest with caution after thorough research!
If you believe that the content used on this website infringes your copyright, please contact us immediately (info@kdj.com) and we will delete it promptly.