AI IN LEGAL PRACTICE

AI Hallucinations in Legal Research: Why LLMs Invent Cases

Why LLMs fabricate case citations, the court sanctions on record, a five-step verification protocol, and which tools ground answers in real databases.

Book a Free Strategy Call →

Why AI Invents Case Citations That Don't Exist

Large language models hallucinate legal citations because they generate text by predicting the next plausible token, not by querying a database of real cases. A citation like "Varghese v. China Southern Airlines, 925 F.3d 1339 (11th Cir. 2019)" — one of the fabrications in Mata v. Avianca — is statistically shaped like a real citation: plausible party names, a real reporter, a real circuit, a real-looking pincite. The model has learned the form of legal authority from millions of documents without any mechanism to check whether a specific case exists. Worse, when asked to verify, an ungrounded model simply generates more plausible text: the Mata lawyers asked ChatGPT whether its cases were real, and it affirmed they were.

The scale of the problem is documented. Stanford's RegLab and HAI researchers found general-purpose chatbots hallucinated on 58%–82% of legal queries tested, and a 2024 follow-up found even leading legal-specific research tools gave incorrect or incompletely supported answers in roughly one in six to one in three responses — better, but nowhere near "trust without checking." Hallucination is not a bug being patched out next quarter; it is intrinsic to how generative models work, and it is most dangerous precisely where output looks most authoritative.

Case citations are only the most visible failure. The same mechanism fabricates statutory section numbers, invents subsections of real regulations, misstates procedural deadlines, attributes positions to secondary sources that never took them, and produces confident summaries of legal tests with one element quietly wrong or missing. These errors are more dangerous than fake cases precisely because they are harder to catch: a lawyer who would never file an unverified citation may still absorb a hallucinated four-part test into a memo, and the mistake surfaces months later in a strategy built on it. The verification mindset therefore has to cover every load-bearing legal proposition an AI supplies, not just the ones with reporter citations attached.

For lawyers, the professional stakes flow through familiar rules: competence (ABA Model Rule 1.1; FLSC Model Code 3.1-2), candor to the tribunal (Model Rule 3.3), and Rule 11's certification that filings are warranted by existing law. None of these rules mention AI — and none of them need to. A fabricated citation in a filing is the lawyer's fabrication, whatever tool produced it. This article is part of our AI in Legal Practice library, alongside the broader question of whether lawyers can use ChatGPT at all.

The Sanctions Record: A Line of Cases Every Lawyer Should Know

The documented consequences now form a genuine line of authority in both countries:

The phenomenon is now tracked systematically: researchers and legal bloggers maintain running databases of AI-hallucination court incidents, and the count crossed into the hundreds of documented matters worldwide during 2025 — including filings by large firms, government lawyers, and a striking number of self-represented litigants. Judges have responded predictably: many now run suspicious citations themselves before hearings, court staff flag unverifiable authority, and several courts have said openly that the era of good-faith assumptions about citation accuracy is over. For practicing lawyers the message is double-edged — the risk of getting caught is effectively 100% in a contested matter, and opposing counsel checking your authorities with a citator is now the cheapest sanctions motion in the building.

Two patterns run through every decision. First, the sanction is rarely for using AI — it is for failing to verify. Second, courts punish the cover-up harder than the mistake: lawyers who doubled down when confronted, or blamed the tool, fared far worse than those who promptly admitted the failure and corrected the record. The lesson is procedural, not technological.

The Verification Protocol: Citator Checking as a Non-Negotiable Step

A defensible verification protocol takes minutes per authority and should be firm policy for any AI-assisted research:

Assign the protocol explicitly. In firms that have had near-misses, the failure was almost never that no one knew to check — it was that everyone assumed someone else had. A one-line certification in the file ("all authorities verified in [database] on [date] by [initials]") costs nothing and creates the record you will want if a court ever asks.

Grounded Tools: Which Systems Actually Look Cases Up

The structural fix for hallucinated research is retrieval-augmented generation (RAG): systems that search a real legal database first and generate answers only from the retrieved documents, citing them with links. Westlaw's AI-assisted research (CoCounsel), Lexis+ AI, and vLex's Vincent AI all work this way, grounding output in their underlying case law collections; in Canada, tools built on CanLII's corpus and the Canadian offerings of the major platforms do the same. Grounding dramatically reduces fabricated citations — the linked case is real because it was retrieved, not imagined.

Match the tool tier to the task in firm policy. A sensible three-tier rule: general chatbots may inform and draft but never supply authority; grounded legal research tools may supply candidate authority that a lawyer verifies per protocol; and only lawyer-verified authority appears in anything filed, sent, or relied upon. Writing the tiers down matters because the failure mode inside firms is informal escalation — a summary generated for background quietly becomes the research memo, which becomes the brief. A bright-line rule that authority has exactly one permitted source chain (real database → lawyer verification → work product) is easy to train, easy to audit, and eliminates the ambiguity that produced every sanctions case on the record.

Grounding reduces the problem; it does not eliminate the verification duty. The Stanford follow-up study found grounded legal tools still misdescribed holdings, over-claimed support, and occasionally mis-grounded — so the proposition and citator checks above remain mandatory even with premium tools. The right mental model: grounded tools move you from "assume everything is fake" to "assume the cases are real but the characterizations need reading." General chatbots never earn even that.

One more institutional note: hallucination risk is also why regulators fold AI into technological competence. Understanding why the tool fails is now part of using it competently — see our survey of bar and law society rules on AI. And the same retrieval-vs-generation distinction explains which marketing content AI engines will cite about your firm — a topic our AI comparison guides cover from the visibility side. To see how AI engines currently describe your own firm, run the free AI Visibility Checker, or talk to LexScale.ai about a complete AI strategy.

Put AI to Work in Your Firm — the Right Way

LexScale.ai helps law firms across Canada and the United States adopt AI for growth — from client-facing intake and content systems to the visibility that puts your firm inside AI answers.

Book a Free Strategy Call →

Unfamiliar with a term here? Our plain-English legal AI glossary defines the concepts firms encounter most.

Frequently Asked Questions

Why does ChatGPT make up fake case citations?
Because language models predict plausible text rather than querying a database. A fabricated citation is statistically shaped like a real one — correct reporter format, plausible parties — but nothing in the model checks whether the case exists.
How often do AI tools hallucinate legal information?
Stanford researchers found general chatbots hallucinated on 58%–82% of legal queries, while leading grounded legal research tools still erred in roughly one in six to one in three answers — far better, but never verification-free.
What sanctions have courts imposed for AI-fabricated citations?
Documented consequences include the $5,000 Rule 11 sanction in Mata v. Avianca (S.D.N.Y. 2023), a Second Circuit discipline referral in Park v. Kim (2024), personal costs against counsel in Zhang v. Chen (B.C.S.C. 2024), and struck filings and regulator referrals since.
How do I verify an AI-generated case citation?
Pull the case by citation and party name in Westlaw, Lexis, or CanLII; confirm each quotation at its pinpoint; read the case to confirm the stated holding; then run KeyCite, Shepard's, or CanLII noteup to confirm it remains good law.
Which legal AI tools ground answers in real databases?
Retrieval-augmented tools such as Westlaw's CoCounsel AI research, Lexis+ AI, and vLex Vincent search a real case-law corpus and cite retrieved documents. They largely eliminate fake cases but can still misstate holdings, so verification remains required.
Is a grounded legal AI tool safe to cite without checking?
No. Grounded tools return real cases, but studies show they still over-claim support and misdescribe holdings. Treat their output as a research memo from a junior: real sources, characterizations to be read and confirmed by the responsible lawyer.

This article is general information, not legal or ethics advice. Professional-conduct rules on AI are evolving and vary by jurisdiction — always verify current requirements with your state bar, law society, or regulator before adopting any AI workflow.

Related Articles

Bar Rules on AI for Lawyers: US & Canada Guidance  ·  Can Lawyers Use ChatGPT? Ethics, Risks & Workflow  ·  Client Confidentiality and AI Tools for Lawyers  ·  Court Rules on AI in Filings: US & Canada Guide  ·  Billing Ethics for AI-Assisted Legal Work  ·  AI Contract Drafting for Lawyers: What Works

Ready to grow your firm with AI?