Legal AI’s Hallucination Problem Is Getting Better — But It Isn’t Going Away
Legal AI is becoming dramatically more capable, and sophisticated systems now use retrieval, authoritative databases, citations and automated verification to reduce errors. But hallucinations remain a stubborn problem — and the most dangerous mistakes are increasingly the ones that look almost right.
If you have spent any meaningful amount of time testing generative AI for legal work, you have probably encountered a hallucination. Sometimes it is obvious: a case that does not exist, a quotation that cannot be found, or a statutory provision that appears to have been invented from scratch. Those mistakes are embarrassing, but in some respects they are the easier ones.
The more difficult problem is an answer that is almost right. The AI identifies the correct contract provision but overlooks an exception. It finds a real case but slightly misstates the holding. It correctly summarizes most of a regulation while missing the qualification that changes the result. It extracts 19 of the 20 material issues in a diligence review and fails to identify the one provision the buyer actually needed to know about.
Generative AI has improved enormously. Legal AI platforms increasingly use retrieval-augmented generation, citation verification, proprietary legal databases, structured workflows and increasingly capable reasoning models. OpenAI itself says newer models hallucinate significantly less than earlier generations while acknowledging that hallucinations remain a fundamental challenge.
Yet the evidence does not support the conclusion that hallucinations have been solved. A new August 2026 study examining eight legal retrieval-augmented generation systems found substantial variation: the strongest systems hallucinated in fewer than 10% of responses, while the weakest approached hallucination rates of one-half. False-premise questions were particularly difficult.
Hallucinations are increasingly becoming a problem to engineer around rather than a problem we can assume the next model release will simply eliminate.
What exactly is a hallucination?
At its simplest, an AI hallucination is information generated by a model that appears plausible but is false or unsupported. OpenAI describes hallucinations as plausible but false statements produced by language models and argues that they persist partly because model training and evaluation can reward systems for guessing instead of acknowledging uncertainty.
In legal work, however, the term can hide several very different failure modes. An AI system might invent a case entirely, cite a real case but attribute a proposition to it that the judgment does not support, provide a quotation that resembles judicial language but does not actually appear in the decision, identify the correct contractual clause but misunderstand what it does, or confidently state that no relevant provision exists when the provision appears elsewhere in an amendment or schedule.
Fabrication: the authority, fact or quotation does not exist.
Misattribution: the authority exists, but the proposition attributed to it is wrong.
Misinterpretation: the source is correct, but the AI misunderstands its significance.
False-premise acceptance: the AI accepts an incorrect assumption in the user’s question.
Unsupported inference: the AI presents a plausible deduction as though it were established fact.
Those errors are not equally dangerous and they are not equally easy to detect. A fake citation looks like a hallucination. A subtly wrong interpretation can look like legal work.
Why do advanced models still hallucinate?
It is tempting to assume hallucinations are simply a symptom of immature AI and will disappear as models become smarter. The underlying problem is more complicated. Large language models are not traditional databases that retrieve a verified fact and display it. At their core, they are trained to predict likely continuations of text. Fluency is therefore native to the system; truthfulness has to be engineered and reinforced around it.
OpenAI’s research on hallucinations argues that part of the problem originates in pretraining itself. Models learn patterns from enormous amounts of text without every proposition being labelled true or false. Low-frequency or arbitrary facts cannot always be inferred reliably from those patterns. Evaluation can compound the problem when a system is rewarded for attempting an answer instead of appropriately acknowledging uncertainty.
Legal work is unusually unforgiving of plausible mistakes
Legal practice creates a particularly difficult environment because legal correctness often depends on small distinctions. Jurisdiction matters. Dates matter. Procedural posture matters. Defined terms matter. Exceptions and amendments matter. The difference between “may,” “shall,” “unless” and “except” can materially change an outcome.
This means an answer can contain accurate ingredients and still produce an incorrect legal conclusion. ABA Formal Opinion 512 recognizes this risk, noting that generative AI tools can combine otherwise accurate information in unexpected ways to produce false or inaccurate results and that lawyers’ uncritical reliance on those outputs can result in inaccurate advice or misleading representations.
Legal hallucination risk is therefore not merely a question of whether an AI system knows the law. It is also a question of whether it correctly connects the law, facts, documents and context of the particular matter.
Retrieval helps enormously — but RAG is not a truth machine
One of the most important developments in legal AI has been retrieval-augmented generation, usually shortened to RAG. Instead of asking a model to answer entirely from what it learned during training, a RAG system retrieves relevant material from an authoritative collection and then asks the model to generate an answer grounded in those sources.
That architecture has materially improved legal AI. It gives systems access to authoritative and current material, allows answers to be accompanied by citations and gives lawyers something they can verify. But retrieval does not eliminate hallucination.
A Stanford study published in the Journal of Empirical Legal Studies in 2025 tested then-current legal research products using more than 200 preregistered questions. The specialized RAG-based systems substantially reduced hallucinations relative to GPT-4 but still returned false or misleading information on a meaningful share of the queries tested.
Those 2025 results are a historical empirical benchmark, not current accuracy scores for today’s versions of the products. Legal AI systems have evolved rapidly since the testing was conducted.
The broader finding nevertheless remains important: grounding a model in authoritative legal material materially reduces hallucination risk without guaranteeing correctness. The August 2026 study of eight legal RAG systems reinforces that conclusion, finding wide variation between systems and continued hallucinations even in retrieval-grounded architectures.
False-premise questions expose a deeper problem
Imagine asking an AI what the Supreme Court held in a case that does not exist, or asking what consent is required under a change-of-control clause when the contract contains no such clause. A robust system should challenge the premise. A weaker system may accept the assumption and attempt to satisfy it.
That matters because lawyers themselves can be wrong. We can misremember a case. Clients can give us incorrect facts. A junior lawyer may frame a research question around an assumption that has not actually been established. Both the Stanford research and newer legal RAG testing have identified false-premise questions as a meaningful vulnerability.
The most useful legal AI should not merely answer questions well. It should know when to push back.
The most dangerous hallucination may be a real citation used incorrectly
Fake cases have understandably dominated discussion of legal AI. They are dramatic and relatively straightforward to detect. But as legal AI improves, lawyers should increasingly worry about errors that are less visible: a real case whose holding is exaggerated, a genuine quotation taken from the wrong opinion, a correct statute that has been superseded, or a contract provision whose meaning changes because of another clause elsewhere in the agreement.
Those outputs can survive superficial review because everything looks plausible. The citation works. The document exists. The paragraph sounds professional. The mistake only becomes apparent when someone actually reads and understands the authority.
Courts are no longer treating this as a novelty
The legal profession has now had several years to learn that generative AI can fabricate authority, and courts are increasingly unwilling to accept ignorance as an excuse. In June 2026, the U.S. Court of Appeals for the Ninth Circuit imposed sanctions in LNU v. Blanche after lawyers filed briefs containing nonexistent cases, misattributed quotations and serious misrepresentations of real authorities.
The court’s point was especially important: the problem was not simply that AI had been used. The professional failure occurred when lawyers signed and filed inaccurate material without fulfilling the obligations that already apply to legal filings.
A separate 2026 study of hallucinated legal citations reported identifying more than 1,000 filings containing fabricated citations and found that the number was increasing year over year. The researchers also tested whether AI could detect those errors and found that newer systems performed much better but still struggled with subtle categories of citation problems.
The real risk is false confidence
Imagine an AI-generated diligence report containing 100 findings. Ninety-five are correct and five are wrong. A 95% success rate sounds excellent. But what if one of the five mistakes is the agreement requiring consent before closing, the customer contract containing a termination right triggered by the acquisition, or the provision affecting the company’s most important commercial relationship?
The usefulness of a legal AI system therefore cannot be measured solely by average accuracy. Legal risk is often concentrated rather than evenly distributed. A tool can be highly accurate overall while still missing the one issue that matters most.
The better AI becomes, the easier it can become to trust it too much.
AI also creates a review paradox
Generative AI can produce legal work much faster than humans can review it. If a lawyer previously produced five research memos in a week and can now generate 50, the organization has dramatically increased production without necessarily increasing its capacity for quality assurance.
The same issue arises in due diligence. AI can review enormous contract populations rapidly, but no law firm will preserve the productivity gain if lawyers must manually reread every document merely to confirm that the AI performed correctly.
The future of legal AI therefore cannot depend on the instruction “verify everything manually.” The more scalable objective is to build systems that make the important parts of the work easier to verify.
The goal should be verification by design
Instead of generating an answer first and worrying about accuracy afterward, legal AI workflows should make important claims traceable from the beginning. For research, that means linked citations to authoritative sources. For contract analysis, it means showing the underlying contractual language beside the extraction. For diligence, it means tracing material findings back to the relevant agreement and provision.
For knowledge systems, it may mean constraining answers to approved repositories. For drafting, it could mean distinguishing precedent-derived language from newly generated text. For agents, it means maintaining logs of actions, sources and decision points so that lawyers can understand how the work product was created.
The quality of the verification architecture may become just as important as the quality of the model.
Law firms should distinguish between different types of AI work
Another mistake is treating every AI output as if it presents the same level of risk. Asking an AI system to format an internal chronology is very different from asking it to state the governing law. Extracting a known field from a contract is different from predicting how a court will resolve an unsettled question.
Lower-risk work
Formatting, basic classification, summarization of noncritical internal material and straightforward field extraction may justify relatively light review, particularly when results can be automatically checked.
Medium-risk work
Contract review, due diligence, document comparison, playbook application and internal legal research may require source citations, confidence thresholds, exception review, sampling and lawyer approval.
High-risk work
Advice to clients, novel research, court filings, regulatory conclusions and strategically significant decisions should attract substantially stronger lawyer verification.
Citation verification should become infrastructure
Lawyers are frequently told to check citations generated by AI. That remains correct, but the technology itself should make the correct behavior easier. Modern legal research products increasingly link AI-generated propositions back to authoritative sources and layer traditional citation-validation systems over the generated response.
The closer verification sits to generation, the lower the friction required to supervise the work. A lawyer should be able to move from a proposition to the underlying source immediately, read the relevant passage and determine whether the cited authority actually supports what the AI says.
Law firms need their own evaluations
Vendor benchmarks are useful, but they cannot fully answer whether a system works for a particular firm. A model that performs extremely well on generalized legal questions may still struggle with the firm’s jurisdictions, contracts, drafting conventions or specialist workflows.
Firms adopting AI at scale should therefore build representative test sets from their own work. Transaction teams can test known contractual outcomes. Litigation teams can maintain research questions with senior-lawyer-verified answers. Knowledge teams can test whether the system answers correctly from approved precedents and whether it refuses unsupported questions.
The important metric is not whether the product looks impressive in a demonstration. It is whether the firm understands how the product fails.
Firms should reward abstention, not confident guessing
Lawyers generally prefer an AI system that says it cannot determine an answer from the available materials to one that confidently fabricates a solution. Yet AI products have historically been optimized around responsiveness. A refusal can look like product failure even when it is the safer legal outcome.
OpenAI’s hallucination research suggests that evaluation methods themselves can encourage guessing when incorrect attempts are not sufficiently penalized. Legal AI should move in the opposite direction: uncertainty should be surfaced, calibrated and escalated instead of hidden behind fluent language.
Human review is not disappearing — but it should change
ABA Formal Opinion 512 makes clear that lawyers cannot simply outsource professional judgment to generative AI. Lawyers need a reasonable understanding of the tools they use and an appropriate degree of verification depending on the circumstances.
But saying that lawyers must review AI output should not end the conversation. The objective should be to reduce the amount of review required while increasing the quality of that review. Lawyers may increasingly inspect linked source passages rather than reread entire documents, review low-confidence results rather than every extraction, or intervene only at predefined high-risk decision points in an agentic workflow.
The hallucination rate is only one reliability metric
There is also danger in reducing system reliability to one hallucination percentage. A system that refuses most difficult questions might achieve a very low hallucination rate while being of limited practical value. Another system may answer almost everything but introduce unacceptable risk.
Legal reliability should therefore be measured across multiple dimensions: accuracy, completeness, source fidelity, citation correctness, calibration, appropriate uncertainty, consistency and the severity of errors when they occur.
In legal work, what kind of mistake a system makes can matter more than how often it makes one.
The hallucination problem is becoming a systems problem
When lawyers first encountered generative AI, hallucinations looked primarily like a model problem: ask the chatbot a question, receive a fabricated answer, and wait for a better model. Modern legal AI is far more complex. A legal AI platform may combine a foundation model, retrieval, proprietary legal information, firm knowledge, citation verification, workflow rules, permissions, evaluation systems, human review and agents using several tools.
Hallucination risk therefore depends increasingly on the architecture surrounding the model. A less powerful model operating inside a tightly constrained, well-grounded and highly verifiable workflow may be safer for a particular legal task than a stronger frontier model given broad autonomy and weak controls.
The key question is becoming less “Which model hallucinates least?” and more “How do we design the workflow so that a hallucination is unlikely to survive long enough to matter?”
The future is not hallucination-free legal AI
Hallucinations will almost certainly continue to decrease. Models will become better calibrated, retrieval systems will improve, citation checking will become more automated, and specialized legal systems will become better at recognizing when a question cannot safely be answered.
Agents may increasingly verify their own work against authoritative materials before presenting it to lawyers. Automated detection systems will catch more citation errors. Firms will become better at testing models against representative workflows rather than relying on generic benchmarks.
But the evidence available today gives little reason to assume that hallucinations simply disappear. Even sophisticated legal retrieval systems still produce them. Courts continue to encounter fabricated citations. And newer AI systems that are themselves designed to detect hallucinations still struggle with subtle errors.
That does not make legal AI unsuitable for serious work. It changes what responsible adoption looks like.
The winners in legal AI may not be the companies that promise to eliminate hallucinations. They may be the companies that make errors visible, traceable, measurable and easy to verify.
The most sophisticated law firms likewise will not be those that instruct lawyers never to trust AI. They will be the firms that understand exactly where trust is justified, where verification is required and how to build workflows in which one plausible mistake cannot quietly become the client’s problem.
Hallucinations may never disappear entirely.
Last updated: August 2026. Analysis and educational content only; not legal advice.