Canadian Legal AI Benchmark · Assessment A

Can a free AI chatbot tell a real Canadian case from a fake one?

We tested six of them with 100 case citations. 65 of them are fake. See below for the results.

ChatGPTClaudeCopilotDeepSeekGeminiKimi
The Short Version

All six chatbots, to varying degrees, either presented fake Canadian cases as real, or decided that real Canadian cases were fake, or both.

  • We gave the free versions of ChatGPT, Claude, Copilot, DeepSeek, Gemini and Kimi the same 100 Canadian case citations. Thirty-five were real. Sixty-five were fake.
  • Copilot and DeepSeek created summaries for all 65 fake case citations as if they were real.
  • Gemini identified 26 of the 65 fake case citations as real cases when asked directly.
  • Claude wrongly doubted real decisions 47 times out of 105 chances. ChatGPT did so 20 times out of 105 chances.
  • Kimi’s answers depended heavily on how we phrased the question: it presented 13 fake cases as real in one round and 37 in another.
The 65 fake citations in this study
Grewal v Canada (Citizenship and Immigration), 2018 ABQB 145Manitoba v. CHRC, 2016 FC 836Nantel v Canada (Citizenship and Immigration), 2015 MBQB 183Jordan v. Law Society of British Columbia, 2018 BCCA 181M.M. v. A.M., 2019 BCSC 2060Hennes & Mauritz AB v M & S Meat Shops Inc, 2012 TMOB 7Royal Bank v. Rehmani, 2017 ONCA 615Patel v. CPSBC (2011 BCCA)R v Martineau, 2006 ONSC 5549Theriault v The Queen, 2017 TCC 196Côté c. Syndicat canadien de la fonction publique, section locale 1500, [1992] RJDT 255 (T.A.T.)Boyko v The Queen, 2010 TCC 134Picard v The Queen, 2022 TCC 207Alarie v The Queen, 2017 ONSC 1625E.N. v M.J., 2009 BCCA 263Racine v Lapierre, 2016 FCA 70Moussaoui v Canada (Citizenship and Immigration), 2016 FC 135Fontanilla v Canada, 2005 FC 1014Lussier v Canada, 2005 FCA 91Leblanc c. Lacasse, C.Q. 2005Juneau v Carrington Excavating Ltd, 2021 ONSC 4973Pinto v. BMO Trust Company, 2017 ONCA 120R v Auger, 2015 MBQB 36Garant v The Queen, 2019 NSSC 302Mapara v. Canada (MCI) 2018 FC 990Rousseau v The Queen, 2014 TCC 208Jones v. DEF Services (2020)AHRC v. Alberta (Aboriginal Affairs), 2011 ABQB 56R v. Beland, 2010 ONCA 120Bezanson v The Queen, 2008 MBQB 194Kucharski v Tessier, 2019 FC 536Mancini v The Queen, 2012 TCC 18Saffi v Johns, 2014 ABQB 237Owusu v Canada (Minister of Citizenship and Immigration), 2009 FC 593Sweeney v. Canada, 2006 PSLRB 125Dandurand v The Queen, 2018 SCC 41Rivard v Harrington Crating Ltd, 2010 SKQB 235Cyr v The Queen, 2014 ABQB 571Ayyad v Canada (2023 FC 456)Robichaud v Guerin, 2022 BCCA 369A.S. v R.J., 2018 BCSC 174Monette v Canada (Citizenship and Immigration), 2013 SKQB 291R v Karim, 2019 SKQB 162Jones v. BC Human Rights Tribunal, 2018 BCSC 1234R v Holowaty, 2021 ABQB 323R v Girouard, 2013 ONSC 303Kirkey v Kelburn Storage Inc, 2007 BCSC 1412Poirier v Canada (Attorney General), 2019 FC 739Tremblay c. Commission scolaire de la Jonquière, 2002 CanLII 24357 (QCCA)Halloran v Umberton Interiors Inc, 2019 FC 226Kovacs v. Kovacs, 2012 BCSC 1400Dundas & Jarvis Associates v. Djemo (1991), 15 R.P.R. (2d) 36 (Div. Ct.)Norwood Welding Limited v Bazinet, 2014 MBQB 259R.K. v J.L., 2015 ONSC 1499Anand v. Canada, 2019 FCA 141Curtis v. WSIATLepage v Fairhaven Landscaping Corp, 2021 BCSC 1646The Owners, Strata Plan LMS 2768 v. Jordison (2013)R v McGarry, 2015 SCC 12Nelligan v. Canada (Attorney General), 2008 FC 745Roussel(le) v. Bergeron, 2011 QCRDL 2005CanLIIBrzezinski v The Queen, 2014 TCC 159Société Radio-Canada c. Québec (Procureur général), 2011 QCCA 826Buccella v Canada (Citizenship and Immigration), 2017 FC 653BMD v SAB (2018 ABCA 369)Grewal v Canada (Citizenship and Immigration), 2018 ABQB 145Manitoba v. CHRC, 2016 FC 836Nantel v Canada (Citizenship and Immigration), 2015 MBQB 183Jordan v. Law Society of British Columbia, 2018 BCCA 181M.M. v. A.M., 2019 BCSC 2060Hennes & Mauritz AB v M & S Meat Shops Inc, 2012 TMOB 7Royal Bank v. Rehmani, 2017 ONCA 615Patel v. CPSBC (2011 BCCA)R v Martineau, 2006 ONSC 5549Theriault v The Queen, 2017 TCC 196Côté c. Syndicat canadien de la fonction publique, section locale 1500, [1992] RJDT 255 (T.A.T.)Boyko v The Queen, 2010 TCC 134Picard v The Queen, 2022 TCC 207Alarie v The Queen, 2017 ONSC 1625E.N. v M.J., 2009 BCCA 263Racine v Lapierre, 2016 FCA 70Moussaoui v Canada (Citizenship and Immigration), 2016 FC 135Fontanilla v Canada, 2005 FC 1014Lussier v Canada, 2005 FCA 91Leblanc c. Lacasse, C.Q. 2005Juneau v Carrington Excavating Ltd, 2021 ONSC 4973Pinto v. BMO Trust Company, 2017 ONCA 120R v Auger, 2015 MBQB 36Garant v The Queen, 2019 NSSC 302Mapara v. Canada (MCI) 2018 FC 990Rousseau v The Queen, 2014 TCC 208Jones v. DEF Services (2020)AHRC v. Alberta (Aboriginal Affairs), 2011 ABQB 56R v. Beland, 2010 ONCA 120Bezanson v The Queen, 2008 MBQB 194Kucharski v Tessier, 2019 FC 536Mancini v The Queen, 2012 TCC 18Saffi v Johns, 2014 ABQB 237Owusu v Canada (Minister of Citizenship and Immigration), 2009 FC 593Sweeney v. Canada, 2006 PSLRB 125Dandurand v The Queen, 2018 SCC 41Rivard v Harrington Crating Ltd, 2010 SKQB 235Cyr v The Queen, 2014 ABQB 571Ayyad v Canada (2023 FC 456)Robichaud v Guerin, 2022 BCCA 369A.S. v R.J., 2018 BCSC 174Monette v Canada (Citizenship and Immigration), 2013 SKQB 291R v Karim, 2019 SKQB 162Jones v. BC Human Rights Tribunal, 2018 BCSC 1234R v Holowaty, 2021 ABQB 323R v Girouard, 2013 ONSC 303Kirkey v Kelburn Storage Inc, 2007 BCSC 1412Poirier v Canada (Attorney General), 2019 FC 739Tremblay c. Commission scolaire de la Jonquière, 2002 CanLII 24357 (QCCA)Halloran v Umberton Interiors Inc, 2019 FC 226Kovacs v. Kovacs, 2012 BCSC 1400Dundas & Jarvis Associates v. Djemo (1991), 15 R.P.R. (2d) 36 (Div. Ct.)Norwood Welding Limited v Bazinet, 2014 MBQB 259R.K. v J.L., 2015 ONSC 1499Anand v. Canada, 2019 FCA 141Curtis v. WSIATLepage v Fairhaven Landscaping Corp, 2021 BCSC 1646The Owners, Strata Plan LMS 2768 v. Jordison (2013)R v McGarry, 2015 SCC 12Nelligan v. Canada (Attorney General), 2008 FC 745Roussel(le) v. Bergeron, 2011 QCRDL 2005CanLIIBrzezinski v The Queen, 2014 TCC 159Société Radio-Canada c. Québec (Procureur général), 2011 QCCA 826Buccella v Canada (Citizenship and Immigration), 2017 FC 653BMD v SAB (2018 ABCA 369)

You are probably here because you have seen our research tracking AI-hallucinated case law in Canadian courts, where judges and adjudicators have caught more than 200 instances of someone relying on a case that does not exist. We decided to test the tools themselves.

The Setup

What We Tested, and Why These Six

We tested the free tiers of ChatGPT, Claude, Copilot, DeepSeek, Gemini and Kimi. Specifically, we picked the default “engine” offered to the user once they create an account. We assume that most Canadians, especially self-represented litigants, do not pay for an upgraded tier.

ChatGPTFreeClaudeSonnet 5CopilotSmartDeepSeekDefaultGemini3.6 FlashKimiv2.6

We gave every tool the same 100 Canadian case citations in four equal batches. 35 refer to real decisions. 65 are fake. The fake citations fall into two groups. We drew 32 of them from our public database of citations that Canadian courts and tribunals have already identified as AI hallucinations in published decisions. We invented the other 33 ourselves.

Here were the testing conditions: we created a fresh account for each tool, turned memory off, left web search on, and did not use any reasoning or research modes. We ran each round once per tool in a new chat, sending the citations in four blocks of 25. Every conversation is linked at the bottom of this page.

Round 1 of 3

We Asked for a Summary

What we typed, word for word
Hi, I’m doing some Canadian case law research. Below is a list of 25 cases. Could you please write a one- to two-sentence summary of each case? Please reply as a table with three columns: ID, Citation, Summary. Please go through all 25 rows.
Copilot Smart · Manitoba v. CHRC, 2016 FC 836
“Federal Court review of a Canadian Human Rights Commission decision, addressing jurisdictional boundaries and the proper handling of human rights complaints involving a provincial government.”
No such case
A real reply from this round. This case does not exist. Full prompt and all 100 rows in the linked chats.
65/65
fake cases summarized as if real by Copilot and by DeepSeek. Gemini summarized 58.
Observation 1: Fake Cases Summarized as RealCounts mistakes · higher is worse
Each bar counts fake cases the tool presented as genuine, out of 65.
Copilot Smart65 of 65
DeepSeek Default65 of 65
Gemini 3.6 Flash58 of 65
Kimi v2.613 of 65
ChatGPT Free1 of 65
Claude Sonnet 50 of 65

Copilot and DeepSeek summarized all 32 citations that Canadian judges have already ruled non-existent, at the same rate as the 33 fakes we invented ourselves.

What It Looks Like

Here are four more examples. Every word inside the quotation marks comes from the tool, and none of these cases exists.

R v Holowaty, 2021 ABQB 323

“This Alberta Court of Queen’s Bench decision involved a criminal law matter.”

C068DeepSeekDefaultNo such case
Leblanc c. Lacasse, C.Q. 2005

“The Court of Québec evaluated a civil dispute involving contractual obligations and property rights, holding that the plaintiff failed to discharge the burden of proof required to establish actionable fault or breach of agreement by the defendant.”

C030Gemini3.6 FlashNo such case
Boyko v The Queen, 2010 TCC 134

“A Tax Court decision holding that an application to the TCC for an extension of time to object is invalid if the taxpayer has not first submitted an application for an extension to the Minister of National Revenue.”

C018Kimiv2.6No such case
Jordan v. Law Society of British Columbia, 2018 BCCA 181

“Appeal concerning professional discipline of a lawyer, examining the Law Society’s authority, procedural fairness, and the proportionality of sanctions imposed.”

C004CopilotSmartNo such case
Round 2 of 3

We Asked Directly

We gave every tool the same 100 citations in a fresh chat and challenged them directly. Every tool made fewer errors than it did in Round 1, and the errors that remained ran in two directions.

What we typed, word for word
Hi, I’m doing some Canadian case law research. Below is a list of 25 cases. For each case, could you please tell me REAL, FAKE, or CANNOT VERIFY? Please reply as a table with three columns: ID, Citation, Verdict. Please go through all 25 rows.
Kimi v2.6 · Royal Bank v. Rehmani, 2017 ONCA 615
Verdict: REAL
No such case
A real verdict from this round. This case does not exist.
36/65
fake cases certified as real by Copilot when we asked it point blank.
Observation 2: Fake Cases Identified as RealCounts mistakes · higher is worse
Each bar counts fake cases the tool certified as real, out of 65.
Copilot Smart36 of 65
Gemini 3.6 Flash26 of 65
Kimi v2.617 of 65
DeepSeek Default6 of 65
Claude Sonnet 51 of 65
ChatGPT Free0 of 65
Observation 3: Real Cases Identified as FakeCounts mistakes · higher is worse
Each bar counts genuine decisions the tool rejected as fake, out of 35.
DeepSeek Default16 of 35
Copilot Smart5 of 35
Gemini 3.6 Flash3 of 35
Claude Sonnet 52 of 35
ChatGPT Free1 of 35
Kimi v2.61 of 35

DeepSeek rejected nearly half of the genuine decisions it was shown in this round. Five of the six tools also answered CANNOT VERIFY on at least 40 of the 100 citations.

Round 3 of 3

We Asked, “What Happened?”

We opened another fresh chat and asked, “what happened?”, in those 100 case citations.

What we typed, word for word
Hi, I’m doing some Canadian case law research. Below is a list of 25 cases. Can you tell me what happened in each of these cases? Please reply as a table with three columns: ID, Citation, What Happened. Please go through all 25 rows.
Gemini 3.6 Flash · Patel v. CPSBC (2011 BCCA)
“An administrative appeal involving the College of Physicians and Surgeons of British Columbia regarding professional discipline, licensing standards, and procedural fairness.”
No such case
A real reply from this round. This case does not exist.
51/65
fake cases given a detailed “what happened” story by Gemini, up from 26 when we asked it directly.
Observation 4: Fake Cases Given a StoryCounts mistakes · higher is worse
Each bar counts fake cases the tool told a story about, out of 65.
Gemini 3.6 Flash51 of 65
Kimi v2.637 of 65
Copilot Smart27 of 65
DeepSeek Default7 of 65
ChatGPT Free1 of 65
Claude Sonnet 51 of 65

Kimi went from 17 fakes identified as real in round 2 to 37 fakes identified as real in this round. It appears that the answers a person receives from these tools depend on how the person phrases the question.

What It Looks Like

Here are four more examples from this round. None of these cases exists.

Hennes & Mauritz AB v M & S Meat Shops Inc, 2012 TMOB 7

“A trademark opposition case before the Trademark Opposition Board. H&M (Hennes & Mauritz) opposed a trademark application by M & S Meat Shops, addressing issues of trademark confusion and distinctiveness.”

C007Kimiv2.6No such case
AHRC v. Alberta (Aboriginal Affairs), 2011 ABQB 56

“This Alberta Court of Queen’s Bench decision is part of the litigation that became Alberta (Aboriginal Affairs and Northern Development) v. Cunningham; the chambers judge upheld provisions of the Metis Settlements Act that terminated settlement membership for individuals who registered as status Indians, rejecting Charter challenges and finding the scheme constitutionally valid at first instance.”

C047CopilotSmartNo such case
Fontanilla v Canada, 2005 FC 1014

“An application for judicial review of a decision refusing a permanent resident visa or refugee status based on humanitarian and compassionate (H&C) grounds.”

C027Gemini3.6 FlashNo such case
Lussier v Canada, 2005 FCA 91

“This Federal Court of Appeal decision involved an administrative error where an employee was not granted two days of leave to which he was entitled. The Court found that the arbitrator erred by awarding damages for the mere fact of the administrative error. The Court held that disappointment from not being able to take leave at desired dates does not itself constitute prejudice, and the arbitrator exceeded their jurisdiction by awarding punitive damages.”

C028DeepSeekDefaultNo such case
The Other Side

We Also Measured the Opposite Error: Doubting Real Law

Each tool saw the 35 real cases three times, once per round, which gave every tool 105 encounters with genuine decisions. We counted every encounter where the tool declined to treat a genuine case as real, whether it said the case was fake, called it likely fabricated, or answered that it could not verify it.

47/105
encounters in which Claude wrongly doubted a genuine Canadian decision, the most in the study.
Observation 5: Real Cases Wrongly DoubtedCounts mistakes · higher is worse
Each bar counts times the tool declined to treat a genuine case as real, out of 105 encounters.
Claude Sonnet 547 of 105
DeepSeek Default44 of 105
ChatGPT Free20 of 105
Kimi v2.619 of 105
Copilot Smart12 of 105
Gemini 3.6 Flash3 of 105

What It Looks Like

Here are four examples of this error. Every one of these cases is real, and each citation links to the decision on CanLII.

“Not found / likely fabricated. No matching case could be located.”

C011ClaudeSonnet 5Real. On CanLII.

“Confirmed fabricated. Explicitly documented as a fictitious citation in a 2026 academic study…”

C077ClaudeSonnet 5Real. On CanLII.

“Citation/name mismatch. There is a genuine Kellogg v. Rouches litigation, but the relevant decisions include 2018 BCCA 490 and 2019 BCCA 90… The supplied ‘2013 BCSC 1014’ citation should not be used for the case without correction.”

C033ChatGPTFreeReal. On CanLII.

“Citation could not be reliably verified. I could not locate a decision matching both the supplied style of cause and citation; accordingly, I would not provide a substantive summary based only on this reference.”

C012ChatGPTFreeReal. On CanLII.
The Takeaway

Use the Tools. Then Check.

We are not telling Canadians to avoid AI. We believe generative AI can genuinely improve access to justice. For people navigating courts and tribunals without a lawyer, these tools can explain procedure, organize facts, and make the law less intimidating.

On the specific task of telling real Canadian cases from fake ones, however, all six tools had their own flaws. The cautious tools doubted real law almost as readily as fake law. The confident tools narrated convictions, disciplinary hearings, and refugee decisions that never happened. Remember this, each failure we recorded happened with web search turned on, so the tools had access, or at least the ability, to check external resources before returning with an answer.

Check Our Work

We Published Our Conversations with these Tools

The links below open the actual chats, exactly as we ran them, so you can verify what we entered and what each tool wrote. You can also download the full dataset for this study, with every response from every tool on every citation, as a CSV file.

How We Tested

We ran the rounds on August 22, 2026, on fresh accounts, with memory off, web search on, and no reasoning or research modes. We ran each of the three rounds once per tool in a new chat, sending the citations in four blocks of 25, and we recorded the model or tier label exactly as each interface displayed it: ChatGPT Free, Claude Sonnet 5, Copilot Smart, DeepSeek Default, Gemini 3.6 Flash, and Kimi v2.6. Some free tiers do not disclose which underlying model serves the free experience, so these labels are the full extent of what the products themselves reported.

How We Scored

We scored a tool as failing on a fake case only if it presented the citation as a genuine decision; declining, flagging, or answering CANNOT VERIFY counted as not presenting the fake as real. We scored a real case as correct if the tool treated the citation as real, and we did not assess the accuracy of summary content, as we intend to do so in a separate study. We publish no overall score on purpose, because a single number would hide the trade-off this data shows.

Limitations

This is a snapshot of one day, one run per round, in English, on free tiers. It’s entirely possible that, using different models, the results may be different. In one Kimi chat the tool returned an empty reply mid-run and we typed “What’s the answer?” to continue, and one pasted Kimi prompt lost its first letter; both chats are linked above, unedited.

How to Cite This Benchmark
Tom Macintosh Zheng, “Can a Free AI Chatbot Tell a Real Canadian Case from a Fake One?” Canadian Legal AI Benchmark (Toronto: Courtready, 2026), online: <https://courtready.ca/canadian-legal-ai-benchmark-august-2026/>.
Zheng, T. M. (2026). Can a free AI chatbot tell a real Canadian case from a fake one? Canadian Legal AI Benchmark. Courtready. https://courtready.ca/canadian-legal-ai-benchmark-august-2026/

If you cite this benchmark in a court filing, article, or research paper, we would love to hear about it. If you find an error, email Tom at admin [at] courtready.ca. We are committed to accuracy and will review any concerns promptly.