Research
How big is the hallucination problem, why hasn't it been solved, and what we're doing about it.
The scale of the problem
LLM hallucinations are not rare edge cases — they are a continuous, global byproduct of every deployed model. Even under conservative assumptions, AI systems produce on the order of billions of hallucinated outputs per week.
We can estimate global hallucination volume with a simple model: H ≈ p · N, where p is the hallucination rate and N is total output volume.
ChatGPT alone generates ~18 billion messages per week. Total global AI usage across all systems is plausibly 2–4× that — ~36 to 72 billion outputs per week.
Reported rates vary sharply by domain:
- ~1–2% in tightly constrained workflows (e.g. clinical documentation)
- ~5–15% in typical consumer usage
- ~10–30%+ in high-risk domains (e.g. legal queries)
A blended global rate sits somewhere in ~3–15%.
- Lower bound: ~1 billion hallucinations per week
- Mid-range: ~3–5 billion per week
- Upper bound: ~7–10 billion+ per week
These are rough, order-of-magnitude estimates, not directly measured. The goal is to understand scale, not precise counts. Hallucination rates are gradually improving as models improve, but total usage is growing faster — implying total hallucination volume is still increasing over time.
Why it's hard to solve
Open vs. closed domain
Open-domain hallucinations — invented entities, fake citations, self-contradictions — can in principle be caught by examining the response against general background knowledge. Closed-domain hallucinations happen when the LLM was given a specific source (a document, a retrieval result) and its response contradicts that source. Detecting these requires access to the original context, not just the response. Most detectors (including ours, today) work only on open-domain cases. See our Scope section for how this affects usage.
The deductive horizon
LLMs do a fixed amount of reasoning per forward pass. When a question requires more chained inference than that budget allows, the model produces a confident-looking but unsupported conclusion — the deduction exceeded its horizon. These failures are particularly hard to catch because the output looks fluent and internally consistent. Progress here probably requires new architecture, not better post-hoc filtering.
Post-hoc retrieval isn't enough
A popular approach is to re-query a retrieval index after generation and check whether the response agrees. This works for factual claims that happen to be in the index, but misses invented reasoning, subtle misattributions, and anything outside the index. It also doesn't scale — every detection becomes another retrieval call.
Our approach
Komplex AI's detector classifies the kind of uncertainty in an LLM response — distinguishing fabricated facts from self-contradiction from misleading framing, and so on — and reports a typed result (a probability plus the most-likely hallucination type) rather than a bare score.
See the Performance page for accuracy by hallucination type, and the About page for a longer-form description of the method.
Why we focused on these hallucination types first
The initial taxonomy — fabricated facts, fake citations, self-contradiction, misleading framing, underspecified claims — was chosen deliberately, not by convenience.
They are the most common hallucinations in the wild
When we looked at what real users encounter — journalists fact-checking AI-assisted drafts, lawyers using LLMs for research, doctors summarizing records, students writing papers — these five types covered the vast majority of failures. Rarer academic categories (symbolic reasoning errors, pragmatic failures) exist but are not where most real harm is coming from today.
They have the highest societal impact
- Fake citations erode trust in scholarship. AI-generated legal briefs have been sanctioned for citing nonexistent cases; academic papers have been retracted for fabricated references.
- Misleading framing pollutes journalism and marketing — technically defensible sentences that collectively create a false picture.
- Self-contradiction is one of the clearest signals that a response was not actually reasoned through. Catching it is a baseline quality check for any automated LLM pipeline.
- Fabricated facts make it into legal filings, clinical notes, and news articles where the stakes are highest. We included this category even though it is currently our weakest, because it is one of the most damaging classes.
- Underspecified claims are the subtle ones — they feel like answers but convey no verifiable information.
They have distinct internal signatures
These types arise from different internal failure modes — knowledge gaps, reasoning breakdown, framing distortion, deductive horizon overrun. Designing a detector that covers all of them pressure-tests the full architecture, not just one channel. A system that only catches self-contradiction has a narrow moat; one that addresses the full spread is much harder to replace.
Planned extensions
The current release focuses on general-context natural language hallucinations in English. Future releases will extend detection to code hallucinations (incorrect syntax, invented APIs, logic errors in generated code) and closed-context hallucinations (responses that contradict a document, retrieval result, or system prompt provided at inference time).
Have questions about how the detector works or what it covers? See the FAQ →
Sources
- OpenAI economic research report on ChatGPT usage — cdn.openai.com
- Stanford HAI — AI on trial: legal models hallucinate 1 in 6 or more benchmarking queries — hai.stanford.edu
- Nature — clinical hallucination study (2025) — nature.com
- Exploding Topics — global AI usage statistics — explodingtopics.com
- Ridder & Schilling (2024) — The HalluRAG Dataset: distinguishes open-domain vs closed-domain / contextual hallucinations
- Wikipedia — Hallucination (artificial intelligence)