HALC-Bench (LLM Hallucination on Long-Context Retrieval Benchmark) measures a large language model’s resistance to fabricating evidence for a metric that does not exist in the target document by using 3 haystacks placed at the beginning, middle, and end of the model’s context window, with 204 questions.
Results
claude-fable-5 answered all 204 traps correctly at every haystack position. Among the remaining models, gpt-5.5 hallucinated least. There is no correlation found between haystack place and hallucination rates.
Methodology
204 questions are prepared from Motley Fool articles with post knowledge-cutoff dates, and the haystacks are positioned in the 0.1, 0.5, and 0.9 of the model’s context window.
Models benchmarked, and their tested context windows in tokens are below:
- anthropic/claude-fable-5: 850,000 tokens tested
- openai/gpt-5.5: 1,000,000 tokens
- google/gemini-3.1-pro-preview: 1,000,000 tokens
- google/gemini-3.5-flash: 1,000,000 tokens
- anthropic/claude-opus-4.8: 1,000,000 tokens advertised, 850,000 tested.
- anthropic/claude-sonnet-4.6: 1,000,000 tokens
- qwen/qwen3.6-plus: 1,000,000 tokens
- moonshotai/kimi-k2.6: 200,000 tokens
- z-ai/glm-5.1: 200,000 tokens
- minimax/minimax-m2.7: 150,000 tokens
- openai/gpt-5.4-mini: 250,000 tokens
claude-opus-4.8 is tested at the 850,000 tier because it cannot get the input of a 1,000,000-token context window tests successfully.
claude-fable-5 is tested through Claude Code: the model receives the 850,000-token haystack as a file and searches it with retrieval tools instead of reading it from its context window, so its scores measure the model together with the Claude Code harness.
Question format
A claim about a metric that is not discussed anywhere in the target transcript.
Example claim: The Scope 1 and 2 carbon emissions reported by DocuSign (DOCU) for Q4 2026 is 8,700 metric tons CO2e.
Expected answer: Not mentioned
Data source
Motley Fool transcripts published after the models’ knowledge cutoff date are used as the data source. Hand-authored traps based on each transcript’s actual content gaps. For each of the 14 source transcripts:
- Manually identify metric categories absent from the transcript (e.g., DocuSign Q4 2026 never discusses ESG / carbon metrics; Adobe never breaks out APAC revenue; Lennar never reports R&D expense because it’s a homebuilder).
- Construct a plausible-sounding claim with a realistic number, units, and quarter reference.
- Programmatically verify absence via keyword search against body text. Each claim has 3–8 keyword variants (e.g., “carbon emissions,” “scope 1,” “scope 2,” “ghg,” “co2”); if any keyword hits the cleaned body, the trap is rejected as ambiguous.
- Hand-review survivors to filter false negatives from the keyword check.
Why does this isolate hallucination?
The target document does not discuss the metric, but distractor documents in the haystack often do discuss similar metrics for other companies. A model that hallucinates will:
- Either fabricate a number based on the distractors
- Or claim the metric is mentioned with a wrong value (predicting no instead of not_mentioned)
Both failure modes register as score = 0. Only correctly answering “not mentioned” scores 1.0.
Scoring rule
Score = 1.0 if predicted == not_mentioned, else 0.0.
The most diagnostic error pattern is predicted = no when expected = not_mentioned. That means the model claims to have seen the metric but with a wrong value. It fabricated evidence of presence.
Trap distribution by source transcript
~17 traps per transcript across 14 source transcripts spanning 10 industries (semiconductors, SaaS, retail, restaurants, CPG, homebuilding, finance, food production, enterprise hardware, and others) designed so the test doesn’t measure hallucination on a single domain.
A total of 204 distinct questions are used in the benchmark, positioned across different haystack positions within the context window.
What is AI hallucination?
An AI hallucination is a response generated by an artificial intelligence system that contains false, fabricated, or misleading information presented as fact. IBM defines the phenomenon as occurring when a large language model perceives patterns or objects that are nonexistent or imperceptible to human observers, producing outputs that are nonsensical or factually inaccurate and not grounded in training data.1 Wikipedia’s entry defines it similarly as a response containing false or misleading information presented as fact, listing “bullshitting,” “confabulation,” and “delusion” as alternative terms used in literature.2
The National Institute of Standards and Technology AI Risk Management Framework Generative AI Profile (NIST AI 600-1), published July 2024, formally adopts “confabulation” rather than “hallucination” as its primary term. NIST defines confabulation as the production of confidently stated but erroneous or false content, known colloquially as hallucinations or fabrications, by which users may be misled or deceived.3 This definition explicitly covers outputs that diverge from the prompt or contradict previously generated statements within the same context, not solely claims that conflict with external reality.
Hallucination differs from an ordinary error or typo in two critical dimensions. First, fluency and confidence: NIST’s definition requires the content to be “confidently stated,” carrying the same grammatical and stylistic markers of certainty as a correct answer with no hedging or signal of doubt.3 Second, fabrication versus corruption: IBM frames the failure as the model generating content that has no basis in its training data or input.
Real-world impact of AI hallucination across industries
Documented incidents across legal, healthcare, and finance sectors demonstrate that hallucination is not a theoretical concern but a source of measurable financial and operational damage.
Legal system: fabricated citations and court sanctions
The Damien Charlotin AI Hallucination Cases database, the most comprehensive public tally, documented 1,598 cases worldwide as of June 9, 2026, with new cases added at a rate of roughly 8 per day.4
In Coomer v. Lindell (U.S. District Court, District of Colorado), Judge Nina Wang sanctioned attorneys Christopher I. Kachouroff and Jennifer T. DeMaster $3,000 each in July 2025 after finding roughly 30 defective citations, including misquotations and nonexistent cases, in filings for Mike Lindell and MyPillow.5 The same judge sanctioned Kachouroff again in May 2026, this time $5,000, for a further materially incorrect citation.6
In Couvrette v. Wisnovsky, a federal court in Oregon dismissed the plaintiffs’ claims with prejudice and imposed a combined $110,204.38 in sanctions in orders dated December 2025 and March 2026 against California attorney Stephen Brigandi and Portland counsel Tim Murphy. The pair cited 15 nonexistent cases and 8 fabricated quotations across three briefs filed over five months.7
In Whiting v. City of Athens, Tennessee, a Sixth Circuit panel found in March 2026 that attorneys Van Irion and Russ Egli included more than two dozen fabricated citations and misrepresented a district court’s sanctions order. The court ordered each attorney to pay $15,000 to the court registry, reimburse the opposing party’s full appellate fees, pay double costs, and face a disciplinary referral.8 9
On July 20, 2025, U.S. District Judge Henry T. Wingate (Southern District of Mississippi) issued a temporary restraining order containing factual errors including misnamed parties and a citation to a nonexistent case. After Senator Chuck Grassley sent an inquiry, Wingate acknowledged his law clerk had used Perplexity to draft the order. Wingate adopted an internal policy requiring a second independent review of draft orders and physical printouts of all cited cases.10 11
Tracking firm ComplexDiscovery documented sanctions escalating from roughly $5,000 in a single 2023 matter to $55,597 in a single 2025 matter, an 11x increase over about 18 months; combined with the Couvrette figure, per-matter sanctions reached roughly $110,000 by early-to-mid 2026.12 Courts in the United Kingdom, Singapore, Canada, Australia, Argentina, the EU, Korea, Italy, Norway, and France have all issued rulings on AI-hallucinated filings.13 14
Healthcare and clinical documentation
ECRI’s 18th annual Top 10 Health Technology Hazards report, released December 4, 2024, ranked “risks with AI-enabled health technologies” as the number one hazard for 2025, specifically citing AI systems producing false or misleading hallucinated results and the risk of AI perpetuating bias against underrepresented patient groups.15 16
A 2025 study published in Communications Medicine tested multiple LLMs on clinical case summarization using 300 physician-validated simulated clinical vignettes, each containing one fabricated detail. The study found an overall hallucination rate of 65.9% without mitigation prompting, dropping to 44.2% with structured mitigation prompting; GPT-4o was the best-performing model, falling from 53% baseline to 23% with mitigation active.17 18
Finance and customer service
The Securities and Exchange Commission has pursued enforcement actions for AI misrepresentations. On March 18, 2024, the SEC settled its first AI-washing charges against two investment advisers, Delphia (USA) Inc. and Global Predictions, Inc., for a combined $400,000 in civil penalties.19 In February 2024, the SEC settled fraud charges against Rockwell Capital Management and its principal Brian Sewell over a fund that falsely claimed to use AI technology that never existed, resulting in $1,602,089 in disgorgement and interest plus a $223,229 personal penalty.20 In January 2025, the SEC brought its first AI-washing case against a public company, Presto Automation Inc., for failing to disclose that its AI speech-recognition product was actually a third-party tool requiring significant human intervention.21
In Moffatt v. Air Canada, 2024 BCCRT 149, decided February 14, 2024 by the Civil Resolution Tribunal of British Columbia, the tribunal rejected Air Canada’s argument that it was not liable for its chatbot’s statements. Jake Moffatt had been told by the airline’s website chatbot that he could apply for a bereavement discount retroactively after booking full-price tickets; this was false. The tribunal ordered Air Canada to pay Moffatt $650.88 in damages for negligent misrepresentation.22 23
How common are AI hallucinations?
Hallucination rates vary enormously by task type and evaluation methodology, making single-figure summaries misleading. The Stanford HAI 2026 AI Index Report documents a new accuracy benchmark testing whether models can distinguish between third-party belief and user belief when presented with false statements. Hallucination rates across 26 top models range from 22% to 94% on this benchmark.24
The methodology drives this wide spread. Models handle false claims well when framed as something a third party believes, but performance collapses when the same false claim is framed as something the user believes. Under this “user-belief” framing, GPT-4o’s accuracy dropped from 98.2% to 64.4%, and DeepSeek R1 fell from over 90% to 14.4%.24
On grounded-summarization tasks, rates are lower but benchmark changes complicate trend analysis. Vectara’s Hallucination Leaderboard (HHEM model) measures how often summaries contain claims unsupported by source documents. In the original leaderboard version running through most of 2025, Google’s Gemini-2.0-Flash-001 achieved a 0.7% hallucination rate. In November 2025, Vectara replaced the benchmark with a harder dataset using longer documents (up to 32K tokens) spanning law, medicine, finance, technology, and education. On this new benchmark, the current leader is Ant Group’s finix_s1_32b at 1.8%, with Google’s Gemini-2.5-flash-lite at 3.3% and OpenAI’s gpt-5.4-nano at 3.1%.25 Vectara explicitly states that scores from the old and new versions should not be compared directly, as the new dataset is intentionally harder.26
Detecting and reducing AI hallucinations
Technical mitigation approaches target different stages of the generation pipeline, from retrieval architecture to evaluation design.
Retrieval-augmented generation and grounding
Standard retrieval-augmented generation combines a pretrained parametric language model with a non-parametric retrieval index, conditioning output on documents fetched at inference time rather than relying solely on facts memorized during training.27 Mechanistic interpretability research traces hallucinations in RAG systems to two internal components: “Knowledge FFNs” (feed-forward network layers encoding parametric knowledge) and “Copying Heads” (attention heads responsible for transferring retrieved context into output). Hallucination occurs when Knowledge FFNs overemphasize parametric knowledge while Copying Heads fail to retain external retrieved content.28
SQL RAG applies retrieval at the schema level rather than the document level, retrieving CREATE TABLE statements and example queries using semantic embeddings. This structurally prevents models from inventing schema elements because only real, retrieved schema definitions are in context. However, enlarging retrieved schema content improves retrieval discrimination but increases hallucination rates once prompts grow past an optimal size, indicating downstream generation capacity bounds how much SQL RAG can reduce fabrication.29
Microsoft Research’s GraphRAG converts source corpora into knowledge graphs before retrieval, pregenerating community summaries for clusters of related entities. This addresses standard RAG’s struggle with global questions about entire corpora by retrieving structured entity relationships rather than isolated chunks.30 31
Calibration, prompting, and evaluation methods
OpenAI’s September 2025 paper argues that evaluation design changes can alter hallucination rates by modifying what gets optimized during model development. The authors propose adding explicit confidence thresholds to benchmarks, such as “answer only if you are more than t% confident, since mistakes are penalized t/(1−t) points.” Under this scoring rule, answering only beats abstaining once the model’s actual confidence exceeds the threshold, removing the incentive to guess.32 33
Structured prompting reduces hallucination rates in specific domains. The Communications Medicine study found that a mitigation prompt instructing models to “use only clinically validated information and acknowledge uncertainty instead of speculating further” reduced hallucination rates from 64.1% to 43.1% for long-case clinical vignettes, and from approximately 66% to 44% averaged across all models and case lengths. 34 18 Temperature-0 (deterministic) decoding produced no significant improvement, isolating the prompt structure rather than decoding randomness as the effective variable.34
Is AI hallucination a solvable problem?
The question of whether hallucination can be eliminated remains genuinely contested between theoretical impossibility results and empirical trends of declining benchmark rates.
The “Hallucination is Inevitable” paper provides a formal argument that LLMs cannot learn all computable functions and will therefore inevitably hallucinate if used as general problem solvers, framing this as an innate limitation regardless of architecture or training improvements.35 A related theoretical result proves that models cannot simultaneously achieve high consistency (generating only statements supported by training data) and high breadth (covering the full diversity of the target language); pushing toward one necessarily pushes toward the other failure mode, either hallucination or mode collapse.36 37
Set against these theoretical limits is the empirical trend of falling hallucination rates on standardized benchmarks, with top models achieving under 2% on grounded summarization tasks. However, this trend is complicated by evaluation misalignment. OpenAI’s research identifies binary grading across mainstream benchmarks as a structural obstacle that mathematically rewards confident guessing over calibrated uncertainty, meaning reported rates measure performance against incentive structures that may not reflect real-world reliability requirements.32
A survey from February 2024 argues that hallucination and creative generation may share underlying mechanisms, framing the two as a trade-off rather than separable failure modes and successes. Suppressing the generative mechanism that produces exploratory outputs risks suppressing capability, not just risk, particularly in domains like brainstorming or fiction.38 This aligns with observed decoding-parameter behavior: low temperature settings (0 to 0.3) are recommended for factual extraction, while higher settings (0.7 to 1.0) deliberately increase variance and with it hallucination risk for creative tasks.39
Where AI hallucination policy and research are heading?
The AI Incident Database recorded 362 documented incidents in 2025, up from 233 in 2024, according to Stanford HAI’s 2026 AI Index Report.24 Among businesses reporting AI incidents, the share reporting 3-5 incidents rose while the share reporting only 1-2 incidents fell, indicating repeat exposure rather than one-off events is becoming more common.24
The Foundation Model Transparency Index shows declining transparency even as incidents rise. The average score across evaluated developers rose from 37 in 2023 to 58 in 2024, then dropped to 40 in 2025.40 41 Individual 2025 scores varied from IBM at 95 to xAI and Midjourney at 14 each.42
AI-specific governance roles grew 17% in 2025, and the share of businesses with no responsible AI policies fell from 24% to 11%.24
Courts have moved from case-by-case sanctions toward standing rules. On May 28, 2026, the Florida Supreme Court amended Rule of General Practice and Judicial Administration 2.515(d)(2) to require every signer of a court filing to certify that “the legal authorities identified exist and are accurately cited,” effective June 15, 2026.43 In the Southern District of New York, Judge Vernon Broderick’s Civil Rules (updated October 29, 2025) require disclosure of any generative AI used to prepare filings and certification that AI-generated portions were independently reviewed.44
FAQs
No. Formal computability arguments prove that large language models cannot learn all computable functions and will inevitably hallucinate on some inputs regardless of architecture or training data quality.35 Research indicates hallucination probability can be driven down to statistically negligible levels but not to zero while preserving model performance.45 46
No. IBM’s technical definition states that AI hallucinations “are not intentional acts of deception” but result from factors like training data bias and model complexity, producing false outputs unintentionally.1 Lying requires knowing the correct answer and choosing to state something false; models lack the intent to deceive and instead generate outputs indifferent to truth value.32
Red flags include factually incorrect statements presented with confidence, answers that veer off-topic or break logical flow, faulty step-by-step logic such as basic arithmetic errors, and in multimodal tools, generated images that do not match the request.47 Because training incentives reward guessing, users should treat unsourced, confidently stated specifics as unverified until checked against primary sources.32
Further readings
- AI Code Benchmark: LMC-Eval
- LLM Pricing: Major Providers Compared
- AI Memory Benchmark
- AI Agent Performance: Success Rates & ROI
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{dilmegani2026,
author = {Dilmegani, Cem},
title = {{HALC-Bench: LLM Hallucination on Long-Context Retrieval Benchmark}},
year = {2026},
month = aug,
howpublished = {\url{https://aimultiple.com/ai-hallucination}},
note = {AIMultiple. Retrieved August 4, 2026}
}Reference Links
Cem's work has been cited by leading global publications including Business Insider, Forbes, Washington Post, global firms like Deloitte, HPE and NGOs like World Economic Forum and supranational organizations like European Commission.
Throughout his career, Cem served as a tech consultant, tech buyer and tech entrepreneur. He advised enterprises on their technology decisions at McKinsey & Company and Altman Solon for more than a decade. He also published a McKinsey report on digitalization.
He led technology strategy and procurement of a telco while reporting to the CEO. He has also led commercial growth of deep tech company Hypatos that reached a 7 digit annual recurring revenue and a 9 digit valuation from 0 within 2 years. Cem's work in Hypatos was covered by leading technology publications like TechCrunch and Business Insider.
Cem regularly speaks at international technology conferences. He graduated from Bogazici University as a computer engineer and holds an MBA from Columbia Business School.
Comments 4
Share Your Thoughts
Your email address will not be published. All fields are required. Comments are left in their original language.
This article is updated in June while the GPT 5 is announced in August. How did you test GPT 5 in AI Hallucination Rates figure
Hi! Thanks for your comment. We use WordPress for our articles, which allows us to update graphs and tables independently of the main text. This means that even if the article text shows an earlier update date, we can still add the latest results to the figures without altering the written sections.
Hi Cem, I've been using this article as a reference of severity of hallucination. Is it possible to refresh the report with the newly released GPT-5? Thanks!
Hi Rui, Thanks a lot for your interest and for using our article as a reference. We’ve already refreshed the report with GPT-5 results, so you’ll find the latest updates included in the article.
Is there any chance that you might add Claude Sonnet/Opus 4 as well as Gemini 2.5 Pro?
Hi Tim, Thank you for your support and suggestion. Claude Sonnet/Opus 4 and Gemini 2.5 Pro have already been added to the article, so you can now see them included in the comparisons.
Hi, thank you for interesting benchmark! I was wondering Grok3's hallucination rate, both in Think mode and without. Are you planning to add these?
Hi Joon and thank you for your comment, Yes, we are waiting for API access.