Large language models hallucinate legal citations because they generate text by predicting the next plausible token, not by querying a database of real cases. A citation like "Varghese v. China Southern Airlines, 925 F.3d 1339 (11th Cir. 2019)" — one of the fabrications in Mata v. Avianca — is statistically shaped like a real citation: plausible party names, a real reporter, a real circuit, a real-looking pincite. The model has learned the form of legal authority from millions of documents without any mechanism to check whether a specific case exists. Worse, when asked to verify, an ungrounded model simply generates more plausible text: the Mata lawyers asked ChatGPT whether its cases were real, and it affirmed they were.
The scale of the problem is documented. Stanford's RegLab and HAI researchers found general-purpose chatbots hallucinated on 58%–82% of legal queries tested, and a 2024 follow-up found even leading legal-specific research tools gave incorrect or incompletely supported answers in roughly one in six to one in three responses — better, but nowhere near "trust without checking." Hallucination is not a bug being patched out next quarter; it is intrinsic to how generative models work, and it is most dangerous precisely where output looks most authoritative.
Case citations are only the most visible failure. The same mechanism fabricates statutory section numbers, invents subsections of real regulations, misstates procedural deadlines, attributes positions to secondary sources that never took them, and produces confident summaries of legal tests with one element quietly wrong or missing. These errors are more dangerous than fake cases precisely because they are harder to catch: a lawyer who would never file an unverified citation may still absorb a hallucinated four-part test into a memo, and the mistake surfaces months later in a strategy built on it. The verification mindset therefore has to cover every load-bearing legal proposition an AI supplies, not just the ones with reporter citations attached.
For lawyers, the professional stakes flow through familiar rules: competence (ABA Model Rule 1.1; FLSC Model Code 3.1-2), candor to the tribunal (Model Rule 3.3), and Rule 11's certification that filings are warranted by existing law. None of these rules mention AI — and none of them need to. A fabricated citation in a filing is the lawyer's fabrication, whatever tool produced it. This article is part of our AI in Legal Practice library, alongside the broader question of whether lawyers can use ChatGPT at all.
The documented consequences now form a genuine line of authority in both countries:
The phenomenon is now tracked systematically: researchers and legal bloggers maintain running databases of AI-hallucination court incidents, and the count crossed into the hundreds of documented matters worldwide during 2025 — including filings by large firms, government lawyers, and a striking number of self-represented litigants. Judges have responded predictably: many now run suspicious citations themselves before hearings, court staff flag unverifiable authority, and several courts have said openly that the era of good-faith assumptions about citation accuracy is over. For practicing lawyers the message is double-edged — the risk of getting caught is effectively 100% in a contested matter, and opposing counsel checking your authorities with a citator is now the cheapest sanctions motion in the building.
Two patterns run through every decision. First, the sanction is rarely for using AI — it is for failing to verify. Second, courts punish the cover-up harder than the mistake: lawyers who doubled down when confronted, or blamed the tool, fared far worse than those who promptly admitted the failure and corrected the record. The lesson is procedural, not technological.
A defensible verification protocol takes minutes per authority and should be firm policy for any AI-assisted research:
Assign the protocol explicitly. In firms that have had near-misses, the failure was almost never that no one knew to check — it was that everyone assumed someone else had. A one-line certification in the file ("all authorities verified in [database] on [date] by [initials]") costs nothing and creates the record you will want if a court ever asks.
The structural fix for hallucinated research is retrieval-augmented generation (RAG): systems that search a real legal database first and generate answers only from the retrieved documents, citing them with links. Westlaw's AI-assisted research (CoCounsel), Lexis+ AI, and vLex's Vincent AI all work this way, grounding output in their underlying case law collections; in Canada, tools built on CanLII's corpus and the Canadian offerings of the major platforms do the same. Grounding dramatically reduces fabricated citations — the linked case is real because it was retrieved, not imagined.
Match the tool tier to the task in firm policy. A sensible three-tier rule: general chatbots may inform and draft but never supply authority; grounded legal research tools may supply candidate authority that a lawyer verifies per protocol; and only lawyer-verified authority appears in anything filed, sent, or relied upon. Writing the tiers down matters because the failure mode inside firms is informal escalation — a summary generated for background quietly becomes the research memo, which becomes the brief. A bright-line rule that authority has exactly one permitted source chain (real database → lawyer verification → work product) is easy to train, easy to audit, and eliminates the ambiguity that produced every sanctions case on the record.
Grounding reduces the problem; it does not eliminate the verification duty. The Stanford follow-up study found grounded legal tools still misdescribed holdings, over-claimed support, and occasionally mis-grounded — so the proposition and citator checks above remain mandatory even with premium tools. The right mental model: grounded tools move you from "assume everything is fake" to "assume the cases are real but the characterizations need reading." General chatbots never earn even that.
One more institutional note: hallucination risk is also why regulators fold AI into technological competence. Understanding why the tool fails is now part of using it competently — see our survey of bar and law society rules on AI. And the same retrieval-vs-generation distinction explains which marketing content AI engines will cite about your firm — a topic our AI comparison guides cover from the visibility side. To see how AI engines currently describe your own firm, run the free AI Visibility Checker, or talk to LexScale.ai about a complete AI strategy.
LexScale.ai helps law firms across Canada and the United States adopt AI for growth — from client-facing intake and content systems to the visibility that puts your firm inside AI answers.
Book a Free Strategy Call →Unfamiliar with a term here? Our plain-English legal AI glossary defines the concepts firms encounter most.
This article is general information, not legal or ethics advice. Professional-conduct rules on AI are evolving and vary by jurisdiction — always verify current requirements with your state bar, law society, or regulator before adopting any AI workflow.
Related Articles
Bar Rules on AI for Lawyers: US & Canada Guidance · Can Lawyers Use ChatGPT? Ethics, Risks & Workflow · Client Confidentiality and AI Tools for Lawyers · Court Rules on AI in Filings: US & Canada Guide · Billing Ethics for AI-Assisted Legal Work · AI Contract Drafting for Lawyers: What Works
Ready to grow your firm with AI?