The GenAI verification burden in legal practice has been highlighted by Joshua Yuvaraj in a recent paper on ‘The Verification-Value Paradox: A Normative Critique of Gen AI in Legal Practice‘ [PDF].

The Abstract reads:

It is often claimed that machine learning-based generative AI products will drastically streamline and reduce the cost of legal practice. This enthusiasm assumes lawyers can effectively manage AI’s risks. Cases in Australia and elsewhere in which lawyers have been reprimanded for submitting inaccurate AI-generated content to courts suggest this paradigm must be revisited. This paper argues that a new paradigm is needed to evaluate AI use in practice, given (a) AI’s disconnection from reality and its lack of transparency, and (b) lawyers’ paramount duties like honesty, integrity, and not to mislead the court. It presents an alternative model of AI use in practice that more holistically reflects these features (the verification-value paradox). That paradox suggests increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers. The paper then sets out the paradox’s implications for legal practice and legal education, including for AI use but also the values that the paradox suggests should undergird legal practice: fidelity to the truth and civic responsibility.

This was met on LinkedIn with, on the whole, acknowledgement and support for this proposition.

However, Antti Innanen did not see it that way. He commented:

It would make a great paradox if it were true.

The supposed paradox – that increases in efficiency from AI use in legal practice will be met by a correspondingly greater need to manually verify outputs, making the net value of AI negligible – is not supported by any evidence.

The paper offers nothing beyond a few anecdotes and old, tired GPT-3.5 hallucination rates.

Meanwhile, massive evidence points in the opposite direction.

It might sound like a responsible thing to say, and it might even be fun to debate, but the data don’t support it.

The claim belongs in the same category as Brian Inkster’s original blog post: funny but not true.

I assume the reference to my original blog post is to Inkster’s Law:

The 2:1 ratio of time needed to check for accuracy of AI generated material

's Law - AI Content Review Principle - Image of the scales of justice showing the balance between GenAI content and the verification burden of checking it.

Although long before that I was advocating the need for The Legal Hallucinatory Detectorist over Legal Prompt Engineers.

I responded to Antti Innanen:

Please do provide to us citations for the massive evidence that points in the opposite direction.

He replied:

To counter a claim “increases in efficiency from AI use in legal practice will be met by a correspondingly greater need to manually verify outputs, making the net value of AI negligible”?

If someone argues that the “net value of AI is negligible in legal practice”, the burden of proof lies entirely on them.

Outrageous claims require evidence, not the other way around.

There’s already substantial proof that AI delivers meaningful efficiency gains when applied thoughtfully. Too tired to post the obvious here.

I countered:

So no “massive evidence” after all. As I thought! The “outrageous claims” concern AI providing huge efficiency gains for lawyers. As you say, this requires evidence. I will always be interested to see such evidence if anyone can actually produce it.

Antti Innanen responded:

Ok, I will bite.

PERFORMANCE:

Vals AI (October 2025) benchmark shows legal AI outperforming human lawyers by 9 percentage points on core legal tasks.

OpenAI’s GDPVal (September 2025) evaluated AI on authentic professional tasks across 44 occupations, including legal work, with Claude achieving near-human baseline performance while completing work much faster.

Then you have the Stanford LegalBench evaluations.

One of my favorite evals is Legalbenchmarks, by Anna Guo. The new study is a gem.

STUDIES:

“Better Call GPT” showed AI surpassing junior lawyers on contract analysis while operating at much lower cost.

“Building a Better Lawyer” demonstrated task completion time reduction with maintained quality.

“Generative AI and Legal Aid” field study showed that legal professionals reported increased productivity.

“AI-Powered Lawyering” showed that o1-preview increased productivity by a significant number.

Oldie but goldie (and non-legal specific): Harvard/BCG’s “Navigating the Jagged Technological Frontier” established 12.2% more tasks completed and 40% higher quality outputs with AI assistance.

ADOPTION:

LegalBenchmarks 2025: 97% of lawyers use AI tools for legal work. 83% actively use two or more platforms.

EFFICIENCY GAINS:

Thomson Reuters surveyed 1,700+ professionals (lawyers included) and found adoption is driven by time-savings.

Everlaw (vendor) reports legal teams saving up to ~32.5 working days per lawyer annually after deploying generative AI in review and drafting workflows.

Juros “State of In-house 2025”: 93% of CEOs & CFOs want their legal teams to increase AI adoption; 99% believe AI will change their legal roles within a year.

Clio 2025 Legal Trends Report says most of those using AI are seeing improvements in the quality of their work, responsiveness to clients, and work capacity.

All the studies that I mentioned point to efficiency gains.

Then there are countless papers from consulting firms:

BCG, McKinsey, Bain, Thomson Reuters’ 2025 Future of Professionals report, Deloitte’s survey of Chief Legal Officers, Forrester’s Economic Impact studies.

MARKET VALIDATION:

Harvey and Legora just both raised $150M with valuations at $8B and $1.8B.

These are just from the top of my head. I’m surely forgetting many more.

On the other side, Brian Inkster, we have your blog post (funny, but untrue) and this study that says:

“However, the actual paradox – based on the above analysis of the verification cost – is fundamentally theoretical, even though anecdotal evidence seems to bear it out. The paradox is therefore a hypothesis for further examination.”

Like your post, it’s an interesting way to point out that verification matters when using AI tools, and that verification costs can be high in legal work. As they should be, this is serious work!

But if you’re truly arguing that “the net value of AI use often negligible to lawyers”, you need to come up with evidence.

I just presented mine and it’s…MASSIVE.

Now it’s your turn to present yours.

But I suspect we’ll get 🦗 🏏 because there isn’t any.

Let me repeat it: there is no evidence that “increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers”.

It’s just a fun thing to say, not a substantiated claim.

Later in the thread Antti states:

Fermi style: where are all the lawyers manually verifying AI outputs all day. Where are they?

I replied to Antti’s “massive” evidence:

That’s not “massive”! It is predictable GenAI hype that does not actually address the point in hand. You won’t get 🦗 🏏 from me. You should know that by now 😉. But having not long got off a 15 hour flight with an 8 hour time difference you will have to wait until I get over the jet lag for a full and detailed response, which might warrant a complete blog post!

I am now over the jet lag and able to do what did, indeed, warrant the complete blog post that you are now reading.

Given how much faith Antti puts in GenAI, I asked Copilot:

Is there evidence that for legal related matters the time spent verifying GenAI output can outweigh the benefits of using it?

Copilot answered:

There is evidence suggesting that the time spent verifying Generative AI (GenAI) outputs can indeed outweigh the benefits, particularly in legal contexts where accuracy and authenticity are paramount.

Challenges of AI-Generated Evidence

The increasing use of AI-generated evidence in legal proceedings presents significant challenges. Judges and attorneys must determine the authenticity, validity, and reliability of digital content that may have been artificially created or manipulated. The rapid advancement of GenAI technology has outpaced the development of reliable detection tools, making it difficult to distinguish between genuine and AI-generated evidence.

Verification Time and Legal Risks

Increased Verification Time: Legal professionals often find themselves spending considerable time verifying the outputs generated by GenAI tools. This is particularly true when the outputs are used as evidence in court. The need for rigorous testing and validation of AI-generated content can lead to delays in legal processes, potentially outweighing the initial time savings offered by using GenAI.

Case Examples: There have been notable cases where attorneys faced sanctions for submitting legal documents that included fictitious citations generated by AI tools. For instance, a New York attorney was fined for failing to verify the authenticity of AI-generated citations, highlighting the professional risks associated with unverified reliance on GenAI outputs. Such incidents underscore the necessity for attorneys to exercise due diligence and maintain human oversight when incorporating AI into their practice.

Conclusion

In summary, while GenAI can enhance efficiency in legal research and documentation, the time and effort required to verify its outputs can be substantial. This verification process is crucial to uphold the integrity of legal proceedings and avoid potential sanctions or ethical violations. Therefore, legal professionals must weigh the benefits of using GenAI against the time spent ensuring the accuracy and reliability of its outputs.

Sources given by Copilot were:

It would be rather ironic if Antti now seeks to debunk what GenAI says on the matter!

Let me now also go into some detail by breaking down, and critically examining, Antti Innanen’s “massive” evidence.

PERFORMANCE:

  • Vals AI (October 2025) benchmark shows legal AI outperforming human lawyers by 9 percentage points on core legal tasks.

As reported by Bob Ambrogi, this particular benchmark also shows that:

Lawyers outperformed AI in roughly one-third of question categories, particularly those requiring deep interpretive analysis or nuanced reasoning, such as distinguishing similar precedents or reconciling conflicting authorities.

These areas underscore, as the report put it, “the enduring edge of human judgment in complex, multi-jurisdictional reasoning.”

It should also be noted that the benchmarking did not include the three largest AI legal research platforms: Thomson Reuters, LexisNexis and vLex.

  • OpenAI’s GDPVal (September 2025) evaluated AI on authentic professional tasks across 44 occupations, including legal work, with Claude achieving near-human baseline performance while completing work much faster.

However, with regard to legal work the report states:

The current version of the evaluation is also one-shot, so it doesn’t capture cases where a model would need to build context or improve through multiple drafts—for example, revising a legal brief after client feedback or iterating on a data analysis after spotting an anomaly. Additionally, in the real world, tasks aren’t always clearly defined with a prompt and reference files; for example, a lawyer might have to navigate ambiguity and talk to their client before deciding that creating a legal brief is the right approach to help them.

  • Then you have the Stanford LegalBench evaluations.

However, these evaluations have been criticised. It has been argued by Ryan McDonough that LegalBench focuses heavily on structured legal reasoning tasks (e.g., statutory interpretation, rule application) that may not reflect the messy, ambiguous, and pragmatic reasoning lawyers often perform in real-world settings.

This could lead to overestimating LLMs’ actual legal competence, since passing benchmark tasks doesn’t necessarily mean a model can handle real legal practice.

While LegalBench tasks are crafted by legal experts, Ryan McDonough has also noted that they are still synthetic and may not capture the complexity, nuance, or unpredictability of real legal documents or litigation scenarios.

For example, tasks often involve binary outputs (e.g., “Is this hearsay: Yes or No?”), which oversimplifies the interpretive nature of legal analysis.

  • One of my favorite evals is Legalbenchmarks, by Anna Guo. The new study is a gem.

Anna Guo has openly acknowledged that public benchmarks—even rigorous ones—can be manipulated. She cites examples like Meta allegedly gaming Hugging Face rankings to illustrate how leaderboard performance doesn’t always reflect real-world capability.

LegalBenchmarks are designed to test legal AI tools on structured tasks like contract drafting or statutory interpretation. However, Guo and others note that these tasks don’t always reflect how lawyers actually work—especially in messy, ambiguous, or collaborative environments.

This report was also aimed specifically at in-house Counsel and information extraction tasks.

STUDIES:

  • “Better Call GPT” showed AI surpassing junior lawyers on contract analysis while operating at much lower cost.

An article by Phuoc Nguyen suggested that Saul might disagree:

The comparison between the cost and time efficiency of LLMs and paralegals, while gleefully underscoring the supreme efficiency of LLMs over human lawyers (what a groundbreaking revelation, right?), conveniently skips over the crucial expense and necessity of having a lawyer double-check the LLM’s work. Unless, of course, you’re feeling particularly bold and ready to join the ranks of those pioneering souls who’ve flirted with professional jeopardy by putting too much faith in their digital counterparts, like that New York lawyer who got quite a stir for cozying up too closely with ChatGPT.

  • “Building a Better Lawyer” demonstrated task completion time reduction with maintained quality.

The study used law students to test AI-assisted legal task performance. Law students are, of course, not representative of practicing lawyers, especially in terms of experience, judgment, and workflow familiarity.

Lars Daniel, in Forbes article, pointed out:

Despite the positive results, the study identified several limitations. Neither AI tool consistently improved accuracy in legal research, with o1-preview showing a tendency to hallucinate sources in some cases.

Additionally, both tools were less effective for transactional work. The one assignment involving drafting a non-disclosure agreement showed no significant improvements in either quality or speed when using AI.

He went on to state:

While AI shows tremendous promise in improving legal work, human verification remains essential for final outputs to ensure no hallucinations or errors slip through. This balance between leveraging AI’s capabilities and maintaining human oversight will be key to harnessing the full potential of AI in law.

  • “Generative AI and Legal Aid” field study showed that legal professionals reported increased productivity.

The study involved 91 participants using paid generative AI tools over two months, with a subset receiving “concierge” support.

Colleen Chien and Miriam Kim acknowledge limitations in their own study. They note that:

  • The sample size was small and self-selected.
  • The study didn’t measure client outcomes or long-term impacts.
  • There were disparities in AI uptake across gender and roles

There is a worry that promoting AI tools in legal aid could lead to overreliance, especially among under-resourced practitioners who may lack time to verify outputs.

There is also, of course, concern that hallucinations or subtle errors in AI-generated legal advice could harm vulnerable clients if not caught and corrected.

  • “AI-Powered Lawyering” showed that o1-preview increased productivity by a significant number.

This was another study that used law students to test AI. As indicated previously, law students are, of course, not representative of practicing lawyers, especially in terms of experience, judgment, and workflow familiarity.

It was again a study that found the AI tools hallucinated.

The six legal tasks used in the study were controlled and standardised, which helps with measurement but may not reflect the messy, ambiguous, and context-rich nature of actual legal practice.

Real cases often involve incomplete information, conflicting precedents, and strategic judgment—factors that are hard to simulate.

  • Oldie but goldie (and non-legal specific): Harvard/BCG’s “Navigating the Jagged Technological Frontier” established 12.2% more tasks completed and 40% higher quality outputs with AI assistance.

As non-legal specific I will discount this entry from Antti as a no go rather than an oldie or a goldie. However, in any event, you may wish to read Conor E. Doherty’s take on why this study is no goldie but “deeply flawed and potentially dangerous“.

ADOPTION:

  • LegalBenchmarks 2025: 97% of lawyers use AI tools for legal work. 83% actively use two or more platforms.

If that AI tool is spell check in Word then this high percentage figure has no doubt been true for some time.

I asked Copilot if this figure was correct. It responded:

While the numbers are impressive, some experts have raised concerns about overgeneralization and the methodology behind the 97% figure. A critical review in the National Law Review suggests that claims of AI “universality” may be premature and that real-world usage varies by firm size, region, and practice area.

This review pointed out that the survey involved only 72 respondents and that quite possibly solicitors contacted who did not use AI simply opted out of taking part “for any number of reasons, including not wanting to appear out of date or due to embarrassment”.

I also asked Copilot for it’s thoughts on a more accurate percentage. It responded:

46% of lawyers now use AI tools, up from just 11% in 2023 — a 318% increase in under two years.

84% of legal professionals are either using or planning to adopt AI by the end of 2025.

55% of lawyers already use generative AI for tasks like research, drafting, and client communication.

61% of UK lawyers report using generative AI, though only 11% use it heavily for day-to-day work.

In-house legal teams lead adoption at 72%, compared to just 31% in the public sector.

Personal injury firms are the most aggressive adopters, with 56% automating tasks like medical record summarization.

Copilot gave its sources for this as aiqlabs and allaboutai.

Whilst not highlighted by Copilot the second of those two sources states that 21% of law firms use GenAI showing a distinction between law firm and individual use. This would then pose the question, is that individual use approved by the law firms that the individuals work for?

Some further (better perhaps) research via Google gave me from Legalit insider that ‘LexisNexis report: Over 60% of UK lawyers now ‘use GenAI’, but law firm culture slows progress‘. In particular this report shows that whilst 60% may say they are using GenAI only “11% of lawyers said they are using GenAI heavily for day-to-day work and that number drops to 7% at large and medium-sized firms”. Furthermore “the most common response in describing their organisation’s AI culture was ‘we’re experimenting but progress is slow.’ 19% reported interest but little investment.”

However, usage figures do not counter the GenAI verification burden. Many solicitors will be playing with GenAI rather than using it for real legal work. Many will, as they have done for many years, be using AI for spell check in Word or other such tasks that are taken for granted but are not GenAI use cases. Some will be using it when it is completely daft to do so. For example, I spotted just the other day this post on LinkedIn by Aaron Krigelski:

Someone told me they used AI to highlight every instance of a word in their document.

“It took some prompt engineering, but I got it to work!”

Ctrl+F. That’s it. That’s the tool.

Been around since the 1970s. Works every time. Takes zero seconds. No prompt engineering required.

But here we are, using artificial intelligence to do what a keyboard shortcut solved decades ago.

I see this constantly. Firms buying AI solutions for tasks their existing software already handles. Like buying a Ferrari to drive to your mailbox.

Last month, a legal assistant spent hours getting ChatGPT to count pages in PDFs. There’s literally software called PDF Page Counter. Does one thing. Does it perfectly.

The week before that? Someone used AI to rename files in sequence. Bulk Rename Utility. Free. Since 2000.

Here’s the thing: AI is incredible for certain tasks. I use it daily. But when you’re using a sledgehammer to push in a thumbtack, you’re not innovative. You’re inefficient.

The best tool isn’t always the newest tool. Sometimes it’s the boring one that’s worked perfectly for 20 years.

Know what problem you’re solving. Then pick the right tool. Not the trendy one.

Your Ctrl key is feeling neglected.

EFFICIENCY GAINS:

  • Thomson Reuters surveyed 1,700+ professionals (lawyers included) and found adoption is driven by time-savings.

The report also disclosed that:

However, even with usage skyrocketing, many of the firms, departments, and agencies in which these professionals work still have a way to go to fully extract value from GenAI. The report shows that few organizations are capturing return-on-investment (ROI) metrics, regularly training staff on GenAI updates, or integrating GenAI use into their policies. In addition, few professionals say their firms and their clients are having conversations around GenAI use. And even more worrisome, questions about GenAI’s impact on billing rates and costs remain unanswered.

  • Everlaw (vendor) reports legal teams saving up to ~32.5 working days per lawyer annually after deploying generative AI in review and drafting workflows.

This relates to ediscovery software which was, of course, doing this for lawyers (using AI) long before GenAI came along.

  • Juros “State of In-house 2025”: 93% of CEOs & CFOs want their legal teams to increase AI adoption; 99% believe AI will change their legal roles within a year.

This is rather than hire more staff. The CEOs & CFOs hope that AI can replace staff. The CEO of Klarna found out that was not the case:

Months after touting AI’s potential to replace human work, Klarna CEO Sebastian Siemiatkowski is backtracking and reversing an AI-induced hiring freeze to bring on more human staff.

Siemiatkowski, 43, told Bloomberg on Thursday that Klarna is hiring human workers again to ensure that customers always have a human presence to talk to, if needed.

“From a brand perspective, a company perspective, I just think it’s so critical that you are clear to your customer that there will always be a human if you want,” Siemiatkowski told the outlet.

Siemiatkowski tells Bloomberg that the AI-focused strategy Klarna employed for the past few years wasn’t the right path. He says that while AI customer service chatbots were cheaper to employ than human staff, they resulted in a “lower quality” output.

  • Clio 2025 Legal Trends Report says most of those using AI are seeing improvements in the quality of their work, responsiveness to clients, and work capacity.

The report also references the risks of AI hallucinations, particularly in the context of legal accuracy and professional responsibility.

Acknowledged Risk: The report notes that while AI tools offer significant productivity gains, they also pose risks – especially when they generate inaccurate or fabricated legal information (hallucinations).

Professional Implications: Lawyers are cautioned to maintain oversight and verify AI-generated outputs to avoid ethical breaches or reputational harm.

Training and Guardrails: The report emphasises the need for proper training and the implementation of safeguards to mitigate hallucination risks in legal workflows.

Contextual Framing: Hallucinations are discussed as part of a broader concern about cognitive load and decision-making quality – highlighting that AI should support, not replace, human judgment.

Thus the GenAI verification burden is acknowledged.

It was also noted recently that Kim Kardashian (who has been studying law) blamed ChatGPT for failing her law exams. So she didn’t see an improvement in the quality of her work by using it! This also shows that the verification burden is real if she was relying mainly on ChatGPT to get her through her exams without such verification.

Kim Kardashian blames CHatGPT for failing law exams

However, not all students are like Kim Kardashian. William Scates Frances, pointed out, just the other day, on LinkedIn that:

A significant number of university students hold a strong stance against generative AI and refuse its use outright or in part.

One of the five top reasons given for this is:

They don’t trust AI outputs. They view chatbots as unreliable and not worth the effort required to babysit them. It’s more trouble than just doing things the “old-fashioned” way.

Thus, the verification burden is being recognised in higher education.

  • Then there are countless papers from consulting firms: BCG, McKinsey, Bain, Thomson Reuters’ 2025 Future of Professionals report, Deloitte’s survey of Chief Legal Officers, Forrester’s Economic Impact studies.

Then there is the fact that Deloitte produced a $439,000 (AUS) report for the Australian Government riddled with GenAI hallucinated errors!

MARKET VALIDATION:

  • Harvey and Legora just both raised $150M with valuations at $8B and $1.8B.

Which does not validate in any way any argument that there is not a GenAI validation burden.

Furthermore, many (including Bill Gates) believe that such valuations are all pointing to a GenAI bubble that may soon burst.

Indeed, it has been suggested that OpenAI are looking for Government bailouts when the bubble bursts.

Just yesterday: Wall Street drops to one of its worst days since April as AI bubble fears return

LAWYERS MANUALLY VERIFYING?

  • Fermi style: where are all the lawyers manually verifying AI outputs all day. Where are they?

Well they should be at their desks doing so all day if they are in fact using GenAI. No doubt they aren’t broadcasting the fact to the world as they record their time privately on their time sheets.

If they are not we have a problem. That problem is borne out by the number of reported court cases where hallucinated citations have been put forward without verification (more on that below – under ‘Funny but Untrue?’)

However, for Antti’s information there is a GenAI legal technology company that is using lawyers to manually verify AI outputs all day. That is LawY:

Ask your legal questions and unlock matter-specific instant AI-answers, with optional verification by qualified lawyers. LawY blends cutting-edge legal AI with human expertise, to enhance your firm’s efficiency with confidence.

That shows recognition of the verification burden but is at the same time, in my opinion, quite ridiculous. It is a tool being sold to law firms who then use LawY lawyers (outside their own firms) to verify the GenAI output and then they should still, of course, do so themselves (as no PI insurance from LawY) thus doubling the verification burden!

Before I knew about LawY, the concept of it was an April Fool’s joke on this blog!: Elwood launches claiming to be Real AI for Law

Funny but Untrue?

Antti says that ‘Inkster’s Law’ is funny but untrue. However, he provides no real evidence to counter it. The foregoing “massive” evidence from him is merely pro GenAI sound bites (many from vendors incorporating GenAI into their products) that can easily be taken down. Indeed, many of them meanwhile actually acknowledge the problem of hallucinations and hence arguably back the principle of the GenAI verification burden.

Antti also considers Joshua Yuvaraj’s academic paper to be funny but untrue. However, again, he provides no real evidence to counter it. If anything he provides evidence in support of it, when he says:

verification matters when using AI tools, and that verification costs can be high in legal work. As they should be, this is serious work!

Yes, verification matters as can be seen by the numerous cases (545 identified so far by Damien Charlotin  who tracks AI Hallucination Cases) where GenAI hallucinated citations have been used in court without such verification. It should be noted that was 532 this time last week. So Damien has clocked 13 new ones just in the space of one week. These are in the public domain. It is frightening to think, on that basis, of what is being hallucinated behind closed doors in law offices around the world daily.

It should be clear to anyone who practices in court (I do and, I assume, Antti does not) that the time that would be spent verifying such erroneous output would outweigh (as Joshua Yuvaraj says) or far outweigh (as I say) any benefits of using GenAI for legal research in the first place.

I talk about this with Kayode Alabi in a recent interview. Although, at that time I had read of 122 reported cases (not the 545 reported, as of today’s date, by Damien Charlotin):

Guidance issued by The Bar Council (England & Wales) stresses the need for verification:

  • Due to possible hallucinations and biases, it is important for barristers to verify the output of LLM software and maintain proper procedures for checking generative outputs.
  • ‘Black box syndrome’ – LLMs should not be a substitute for the exercise of professional judgment, quality legal analysis and the expertise that clients, courts and society expect from barristers.

The verification burden is borne out by comments on Joshua Yuvaraj’s LinkedIn post, such as:

Haward Soper (Honorary Professor Of Law at University of Leicester):

Another must read! It reflects my data too. I have asked around ten AI tools to draft an exclusion of consequential loss for me. Redlining their work would have taken quite some time. This should be published later this year in my upcoming book on excluding consequential loss (which includes a note on the issue in Australia ).

Dr Eliza Mik (IT Lawyer, Academic):

100% true. LLMs accelerate text production- not reading / rewriting time.

Julian Webb (Professor of Law at The University of Melbourne):

Thanks for the HT Joshua Yuvaraj, I’m very much looking forward to reading this. From the summary, I think we will be in substantial agreement!

Jonathan Crass (Privacy, Cyber and AI Lawyer | CIPP/E):

On the reading list and thank you Joshua Yuvaraj for putting in the time to dig deeper.

Sounds like it confirms what I have found in practice and what other colleagues have told me they have found.

Marco Rizzi (Associate Professor, Graduate Research Coordinator and Co-Director of the Centre for Health Law & Policy at UWA Law School):

This aligns with my personal and very anecdotal experience. I have tried to use LLMs to generate scenarios, mostly for teaching purposes. The reality is that the amount of work required to tailor the output to the actual teaching needs eats up pretty much all the time ‘saved’ by outsourcing the initial thinking process. In fact it slows down the editing and re-writing process, because it is not text that I have generated in the first place, so I have less familiarity with it.

Alice Hewitt (Legal Information & Knowledge Management Champion | Discoverability & Accessibility Advocate | Information Literacy Tour Guide | Research Ninja | Interrogator of Mysteries (& databases & AI & IA & the why…):

A interesting read – and supports a lot of what I know many law librarians know (or at least strongly suspect).

Jane Cornwell (Senior Lecturer in Intellectual Property Law at University of Edinburgh Law School):

What a great a paper on generative AI, legal practice and legal education…

As someone who spent many years in legal practice – not only producing my own outputs in terms of advice, correspondence, pleadings and submissions, but also reviewing the work of others and signing off on that work in the name of the firm – I worry that misjudging the proper response to AI in legal education risks setting up students to fall short in the discharge of their professional obligations as they move forward in their careers. This paper’s conclusions really chimed with me – particularly the response to the argument that incorporating AI into legal education is essential because the use of AI will be all-pervasive.

I have previously written about the limitations/dangers of GenAI in legal practice, including the verification burden (although I hadn’t previously referred to it as that):

The GenAI verification burden is also borne out by others writing on the topic:

The Siren’s Song of GenAI: Why legal practitioners still fall for fabricated content‘ – Armin Almardani refers to the phenomenon of ‘verification drift’. He says that “this occurs when users initially approach AI-generated content cautiously, aware of the risks of inaccuracy and the potential for hallucination. However, as they engage with the content, they gradually become overconfident in its reliability and find verification less necessary. This misplaced trust may stem from GenAI’s authoritative tone and ability to present incorrect details alongside accurate, well-articulated data. This suggests that the challenge is not just a lack of awareness but a cognitive bias that lulls users into a false sense of security.” He also refers to the ‘verification burden’ and states that “for certain tasks, relying on generative AI followed by exhaustive verification may be less time-efficient than conducting traditional legal research.”

AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries‘ – A study reveals that “bespoke legal AI tools still hallucinate an alarming amount of the time: the Lexis+ AI and Ask Practical Law AI systems produced incorrect information more than 17% of the time, while Westlaw’s AI-Assisted Research hallucinated more than 34% of the time”.

Hallucinations in Legal Practice: A Comparative Case Law Analysis‘ – Bakht Munir concludes that “it shifts the onus to the end user to counter-verify the generated content otherwise face the music.

Law, Lies, and Language Models: Responding to AI Hallucinations in UK Jurisprudence‘ – Tahir Khan of The Barrister Group discusses the epistemological and regulatory threats that hallucinations pose in UK legal practice, including the need for manual review and correction.

Beyond Guesswork: Reducing Hallucinations in Legal GenAI Tools‘ – Ryan McDonough highlights how hallucinations force lawyers to spend time validating AI outputs and suggests strategies to reduce this overhead.

Law Firms Leveraging AI: Maximizing Benefits and Addressing Challenges‘ – Logan Lathrop explores the tension between automation and accuracy, showing how hallucinations increase legal risk and workload.

AI-hallucinated cases end up in more court filings, and Butler Snow issues apology for ‘inexcusable’ lapse‘ – Debra Cassens Weiss reports on Butler Snow acknowledging that “there is no excuse for using ChatGPT to obtain legal authority and failing to verify the sources it provided”.

When AI hallucinations hit the courtroom: Why content quality determines AI reliability in legal practice‘ – Thomson Reuters highlight how poor-quality AI content leads to courtroom errors, requiring more human oversight and correction.

The risks of using GenAI for legal research‘ – Alex Heshmaty explores whether GenAI should be used for legal research (NB I am quoted within this article).

Jurists of the Gaps: Large Language Models and the Quiet Erosion of Legal Authority – Philippe Tritto and Ilsse Ortega consider that Large Language Models (LLMs) are not merely tools to assist legal professionals—they represent a deeper epistemic and normative challenge to the foundations of legal authority. While LLMs allow humans to produce outputs that convincingly simulate legal reasoning, they lack the embodied judgment, ethical intentionality, and contextual awareness that define legitimate legal decision-making. Their paper argues that the social legitimacy of the legal profession relies on capacities that are not reproducible through computational systems.

These are just a small selection of articles on the topic available online. There are countless more. Just Google it or, dare I say, ask ChatGPT. I tried the latter (requesting a list of 10) and it produced five articles with sources given and five without sources. When I asked for the sources for the latter five it admitted that it could not find articles with the titles it had previously cited! Thus, these were hallucinations creating a verification burden for me before an accurate list of 10 could be reliably cited. Indeed, I found better articles through a Google search than that produced by ChatGPT and only two of its five (non-hallucinated ones) made the final cut. Those two also appeared via a Google search. I would have saved time had I just used Google in the first place.

It is also blatantly obvious. If you use GenAI to research a legal issue and it produces, say, 12 fake citations and misses out goodness knows how many correct ones, then the verification burden is going to be huge. Furthermore, not all case law is digitalised. Even the best software will not necessarily find what you are looking for. A search in an old fashioned law library may well be necessary.

OpenAI’s unified Usage Policies (effective October 29, 2025) explicitly prohibit using their services for:

provision of tailored advice that requires a license, such as legal or medical advice, without appropriate involvement by a licensed professional

automation of high-stakes decisions in sensitive areas without human review … [including] legal [and] medical

Specific GenAI legal technology vendors followed suit, as Antti Innanen acknowledged:

Antti Innanen on Harvey and Legora not providing legal advice

These are all effectively acknowledgements of the verification burden that clearly exists.

It is also an acknowledgement of what GenAI really is. Joshua Yuvaraj, in his paper, puts it like this:

AI models are fundamentally probabilistic – they learn from input data, including its biases and omissions, and return outputs that are statistically likeliest to reflect what is requested by users. They are not structurally linked to reality: namely, factual accuracy, and valid links between ‘factual propositions…[and] relevant legal documents.’ However accurate the training data, a machine learning model does not learn the facts underlying that training data, but reduces that data to patterns which it then ingests and seeks to reproduce with variations depending on the input of a user (instructions/prompts).

The reality flaw means hallucinations – outputs that are ‘false, incorrect, or outright nonsensical’, no matter how plausible-sounding – and other errors occur frequently enough to warrant significant concern. One study of ‘public-facing’ models like GPT-4/3.5 (OpenAI), PaLM2 (Google) and Llama 2 found that, in response to ‘a direct, verifiable question about a randomly selected [US] federal court case’, 58%-88% of responses were hallucinations. Similarly, a study of GPT-4o and Llama-3-8B legal analysis documented that 80% of responses had hallucinations.

One response to this problem is to enhance the quality of the training data. In one sense this can help address concerns about omissions, biases, mistakes and other failings in the training data that can impact the output. Yet even with more high-quality training datasets, or bespoke datasets built for particular contexts, the hallucination problem is not immediately resolved. Studies have repeatedly found ‘hallucinations’ even in machine learning tools built for the legal context. Even where leading legal research companies like Westlaw and Lexis have built AI into their search functions, they remain unreliable.

Or as Jim Amos put it recently on LinkedIn:

Understand this one thing about Large Language Models if you want to stay grounded and sane:

LLMs do not think. They do not reason. They do not understand a single word they output. They don’t understand cause and effect or right from wrong or true from false.

They’re language calculators: they use Natural Language Processing algorithms to intepret prompts then generate statistically plausible answers rendered from algebraic vectors — afterimages of training data compressed into non-euclidian space.

LLMs are a feat of science and a breakthrough in computation, but they are not an emergent intelligence.

It’s a card trick, not real magic. Don’t let anyone convince you otherwise.

I like the simple stochastic parrot metaphor which can be summed up in this image:

ChatGPT v Parrot

The dangers of not verifying GenAI output is summed up nicely in this image, that Mitch Kowalski shared on LinkedIn (thinking I would like it – I do):

GenAI and edible mushrooms - Image of GenAI telling someone a mushroom was edible and once they die confirming it was poisonous and would they like to learn more about poisonous mushrooms

Mapping ‘Inkster’s Law’ against the ‘Verification-Value Paradox’

When I asked Copilot to map ‘Inkster’s Law’ against Joshua Yuvaraj’s ‘verification‑value paradox’ I got this output:

Inkster's Law v Verification-Value Paradox

⚖️ How They Align

  • Both highlight the same tension: GenAI promises speed, but lawyers must verify everything, which can outweigh the benefits.
  • Inkster’s Law gives a catchy, numerical shorthand (2:1 ratio).
  • Verification‑Value Paradox provides the theoretical and ethical framework explaining why that ratio exists.

🧩 Why This Matters

  • Together, they show that the verification burden is not just anecdotal but systemic.
  • Inkster’s Law makes the issue memorable in practice; the Verification‑Value Paradox grounds it in academic critique.
  • Both reinforce that lawyers cannot abdicate responsibility — GenAI outputs must be checked as if the lawyer wrote them.

✅ So, Inkster’s Law is the practitioner’s shorthand, while the Verification‑Value Paradox is the academic framing. They converge on the same point: GenAI may not save time if verification duties outweigh drafting efficiencies.

A little verification burden on that would be to point out that the citations for the ‘Verification-Value Paradox’ may not be correct and the ‘Verification-Value Paradox’ does not necessarily agree with the 2:1 ratio of Inkster’s Law. Otherwise, Copilot has grasped the similarities between the two.

Conclusions

In my opinion the evidence is very clearly stacked against Antti Innanen’s viewpoint and what ‘evidence’ he has produced to counter Inkster’s Law and the findings by Joshua Yuvaraj in his paper. The verification burden is real and has to be taken very seriously. It cannot simply be dismissed as “funny but untrue”.

What that verification burden may be, will vary from case to case. Sometimes it will be a cancelling out of any benefit as Joshua Yuvaraj suggests. Other times it may be a 2:1 ratio as I suggest in Inkster’s Law. However, the verification burden could be higher than that in certain instances. Inkster’s Law is, I consider, a good average and warning point.

That is not to say there might be other uses of GenAI in legal practice that would not carry the same risks or verification burden. This may be the case in legal marketing, for example. Although, I would not personally use it for that due to the bland nature of the output and lack of any clear unique voice to associate with your brand.

Reactions to The GenAI Verification Burden in Legal Practice

On LinkedIn the following comments have been made:-

Antti Innanen (⚫ Ⓜ️ 💯):

Good stuff!

Remember what we are debating: “Paradox suggests increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers.”

We are not debating whether AI outputs should be verified. They should. We are not debating whether verification takes time. It does.

The question is whether the verification burden “renders net value of AI use often negligible to lawyers.”

It doesn’t.

The blog post is mostly a collection of criticism of the studies I brought up. And the criticisms might be valid in some situations.

But I didn’t see anything that would prove this “paradox.”

Actually, it shouldn’t be that difficult to test.

The question is simple: “do you have to use so much time to verify AI outputs that the net value of AI is often negligible?”

If a large group of lawyers (something like half) said yes, the paradox would make sense.

But the more realistic answer is: Yes, verification takes time. No, it doesn’t make the net value of AI negligible.

Me:

You say “It doesn’t”. I say “It does”. In my own experience the time spent does negate the use of it and usually far outweighs that use. Others agree as cited in my blog post. The evidence is fairly clear especially where legal research is involved. As I say the ratio will vary from case to case but it will be rare in legal practice currently that the time involved does not cancel any perceived benefits of using GenAI in the first place.

Antti Innanen:

I have to admit that I actually thought “Inkster’s Law” was a joke.

Like an insightful quip that reminded us that verification is a burden and that we should always verify the outputs.

And as such it was fun and insightful. A bit absurd too.

😅 But now we are arguing that it is actually true.

Me:

That’s what I thought about the metaverse 😉

It was never a joke. It was based on real life experience and my thoughts on that in 2023. Gav Ward coined the term following a debate on the topic on LinkedIn. Two years later and I don’t think a lot has really changed.

Still can’t see why it would be “a bit absurd too”. My blog post sets out in detail why it is not. “Funny but untrue” simply does not cut it.

Yes, we are arguing that it is actually true because I believe it is and you do not.

ChatGPT will tell you that the “2:1 ratio is flexible; complex or poorly prompted AI outputs may require even more verification.” Likewise, those of the variety suggested by Jennifer Marsh, in this thread, will likely require less.

I have personally used GenAI for a timeline summary that was quick and fairly accurate and drew all its information from blog posts I had written over several years. I was able (from my expert knowledge on the subject and as author of the source material) to quickly verify it. It saved me time and the 2:1 ratio did not apply. But for legal research (and especially where you do not have subject expertise) the story is very likely to be a different one.

-+-+-+-+-+-

Mark Bennett (Associate Professor of Law at Te Herenga Waka | Victoria University of Wellington):

This is a weird debate.

The question has to be answered empirically.

I see lawyers who use AI answer that it does help their efficiency and effectiveness. I find the same benefits with my own academic research and administration work. The use cases – legal and otherwise – are described and demonstrated all over the place, and can be tested by trying to use AI yourself.

Why do we need the debate? If using AI wastes your time, don’t do it. If it saves you time and improves your work, do it.

Me:

Many don’t realise the pitfalls involved. Debates like this hopefully raise awareness amongst the uninitiated. Then, as you say it is their choice, but armed with information to assist that choice.

-+-+-+-+-+-

Jennifer Marsh (Product Leader | Former Practicing IP Attorney):

It’s true in general BUT the best use cases don’t require you to validate the output. For example, rephrase this sentence to x, explain this to me in plain English, conversational intake workflow, tell me why I might be wrong, etc.

Me:

But those are not the ‘use cases’ that are going to replace lawyers 😉

-+-+-+-+-+-

Daniel Yim (Founder @ Sideline (track changes in Outlook) | Helping Law Firms Develop Their Lawyers | Legal Tech & Innovation):

Yeah but more verification time = more billable hours = good. 👍

Anyway, everyone’s mileage will vary depending on the task and context. I think use the AI if it’s useful for your work, don’t use it if it’s not, and definitely don’t use it just because big tech (or hangers-on to big tech) tells you that you have to use it.

Me:

Indeed. And they told us AI would kill the billable hour!

-+-+-+-+-+-

John McCarthy (Helping Law Firm & Legal Tech Owners unlock £100K extra profit in 6 months, build a highly profitable, valuable business & plan for exit with the PROFIT System):

It comes down to risk- benefit ratio too Brian Inkster

Using AI may save time (good) but increase risk for lawyers if the appropriate checks and balances are not in place (bad).

My thoughts are that we will move towards AIs use but it should be done with caution ⚠️

Me:

Approach with caution indeed. The problem is many lawyers are throwing caution to the wind as per the 545 cases identified so far by Damien Charlotin.

-+-+-+-+-+-

Ryan McDonough (Head of Software Engineering – AI in Legal: Building, Governing, and Owning the Tech):

Thanks for including my work Brian, fantastic article! The bit I keep seeing in real workflows is that the drag isn’t mistakes, it’s the clean paragraphs with no real idea of how they were produced, which just pushes lawyers back into doing the reasoning themselves.

Accuracy always grabs the headlines, but predictable behaviour is what would actually reduce verification time. That feels like the real gap the next wave of tools need to look to close

Me:

And clearly a real danger in those clean paragraphs if you don’t know where they came from. Many will, unfortunately, accept them without question.

-+-+-+-+-+-

Chantal McNaught (That LegalTech Redhead | 🎙️People in Legal | Answering the question “how lawyers can navigate the conflicts between law as a profession and law as a business?”):

The marketing “lie” about generative AI in legal practice is that it is merely a “productivity tool”. It isn’t. It fundamentally changes what the practice of law looks like, and simply validating the outputs will absolutely add to the overall time it takes to complete legal work product. Lawyers should consider the high-value vs low-value legal work products in the delivery of legal services, and design their processes accordingly.

-+-+-+-+-+-

Robert Dean, J.D. (Director, Litigation | UnitedLex | eDiscovery for in-house and law firms | Legal data annotation for AI):

It takes your associate 10 hours to draft a memo. It takes you an hour to check the cites / sources before filing. Or, it takes your AI Copilot 10 minutes to spin up a brief. It takes you an hour to check the cites / sources before filing. I don’t see the burden – what am I missing?

Antti Innanen (⚫ Ⓜ️ 💯):

According to Inksters law you spend only 20 minutes verifying so it takes only half an hour 🤣

Robert Dean, J.D.:

That’s right. The act of “verifying” outputs is what lawyers in firms have been doing since the invention of associates. That the output is AI-generated doesn’t change the profession. It’s part of doing good work, not an argument against adopting more efficient tools.

Me:

The Associate probably won’t make stuff up (hallucinate) and their cites/sources will be easier/quicker to check. Your AI Copilot might well spit out fabricated nonsense and miss much that you would want included along the way. As a result you just might spend a lot longer checking the AI Copilot output than you would have to with the Associate output.

Robert Dean, J.D.:

Hallcuination itself has been a fundamental part of machine learning since Rosenblatt in the 1960s. A model is always limited by the training data, and the hyperplane will be drawn anywhere within the bias range. It is within that range where the machine “hallucinates” as it has to guess on which side of the plan the input data might fall. Nothing wrong with that; you can fill in the bias gap with fine tuning to add more data. It’s just math – and, the size of the models (ie the availability of more data to calibrate the hyperplane and reduce gaps) is only getting more robust. Kimi K2 just released a trillion (!) parameter model.

But your premise that “cites/sources” are easier to check with an associate generates the output versus the an AI Copilot is just not true. Ever generate a brief using GPT-5.1 with Deep Research? It spins up a brief in less than 20 minutes. You get Bluebook citations and can spend the same amount of time checking the cites as you would any other work product.

Bottom line, verifying is what lawyers have always done. Whatever extra time is spent verifying AI work product is easily captured by the efficiency of using AI in the first place.

Me:

It will be true whenever the Associate’s output is better than the Copilot’s output. You cannot say that will never be the case. Whilst hallucination may be a fundamental part of machine learning, I don’t think it has ever been a fundamental part of Associate learning!

-+-+-+-+-+-

Anna Guo (📕 Lawyer | Legal AI Researcher):

I honestly don’t see how this can be properly debated, because the tradeoff is different for everyone and depends on what they’re using AI for.

I agree with what Jennifer Marsh said. Some tasks, like using AI to brainstorm ideas or for polishing, require very little verification, whereas using AI to summarize open-ended queries can require much more in a high-risk scenario.

And let’s not forget the study published by METR earlier this year, which found that engineers did spend more time than expected verifying and correcting AI-generated code, but also that the cognitive burden of reviewing AI output was lower than writing the code themselves. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

Both halves of that result matter.

At the end of the day, it’s important to understand both the failure modes and the benefits of using AI assistance.

That, I think, is the consensus we can all come to here.

Me:

It has been fully acknowledged that the trade-off is different for everyone and depends on what they’re using AI for.

The debate arose from a suggestion that the GenAI verification burden in legal practice was “funny but untrue”.

Allowing lawyers to think so creates the danger of perpetuating the non-verification seen in the 545 cases identified so far by Damien Charlotin which are rising in number by the day.

Debates like this are necessary for educating the uninitiated. I trust that my blog post, and the comments arising on here from that, will do just that.

If that results in a greater exercise within legal practice of the caution that John McCarthy refers to, then it will have been a worthwhile debate.

And this assists, as you say, “to understand both the failure modes and the benefits of using AI assistance.”

Then, yes, I too think there is the consensus we can all come to here.

Antti Innanen (⚫ Ⓜ️ 💯):

No, we are not debating verification burden in general.

We are debating if ”increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers.”

Lets not move the goalposts.

Me:

I’m not moving any goal posts. That debate is about the verification burden and the weight of it. You dismiss that weight, I do not.

Antti Innanen:

You are.

What I am saying that verification is important and it surely takes time.

But it doesn’t ”render the net value of AI use often negligible to lawyers”.

Me:

Not always (I accept that), but more often than not (especially when actual legal tasks are involved).

Sam Davidoff (CEO at Align | Former Trial Lawyer | Electronic Binders for Litigators | Litigate with your iPad | Using my 20+ years of litigation experience to give lawyers practical tools and advice that will make them more effective):

Anna Guo – something else about the METR article that seems relevant to this debate is that it emphasizes that people are bad at assessing their own efficiency. It’s not just that the devs WERE 19% LESS efficient, it’s that they also THOUGHT they were 20% MORE efficient.

It always makes me skeptical of claims that run along the lines of “but I use it and I’m more efficient” (whether made individually or in aggregate via surveys). I’m not even saying those claims are wrong. I’m just saying I have no idea how to evaluate how reliable that evidence is coming from lawyers. (Though I will note I have some innate skepticism from my time as a lawyer and seeing how bad lawyers are at estimating how long something SHOULD take to do.)

Josh McBride (Barrister at Richmond Chambers):

Brian Inkster – honestly this is like saying Microsoft Word is complicated and takes ages to learn before you will realise any efficiencies, so you should just keep writing with a fountain pen!

Me:

Not a good analogy. What we are debating here is the inherent hallucinatory problems of GenAI (not to mention its other drawbacks) and the verification burden placed on detecting and correcting those hallucinations (and filling in the gaps it will simply not output). This is not an issue when using Microsoft Word.

Josh McBride:

it’s an excellent analogy. “Hallucinations” are only an issue with generative AI if you don’t know how to use it. Just like Word won’t correct typos if you don’t know how to use spellcheck.

Me:

Spellcheck runs for me in Word every time I use it. I didn’t realise there was a way of switching it on and off. Just as well it has always been on! Do please tell us how you turn hallucinations off.

Abigail Hope S. (Empowering GCs across APAC to lead strategy, digital transformation & innovation | Senior Manager, Legal Transformation at SCG | MIT Sloan MBA | JD/LLM):

Anna Guo – I love your point about context-dependent tradeoffs, Anna and it deserves more attention than it gets. Most corporate AI policies today are blanket AI policies. For our company it’s “use it everywhere” and “verify everything equally” policies, where both miserably fail. High-stakes summarization for a tax opinion from legal needs different guardrails than UI copy generation from corporate branding, but most AI governance frameworks treat verification as binary. The failure modes aren’t just different in degree, they’re different in kind.

-+-+-+-+-+-

Robyna May (CIO – Specialising in Innovation, Strategy, Change & Leadership – Writer & Speaker – AI enthusiast | McInnes Wilson Lawyers.):

AI hallucinates. It’s not a problem that’s been solved. Therefore, where using it generate content, it has to be checked carefully. The analogy to an eager intern rings true. Where we use AI differently, say summarising a document, it can still hallucinate but the burden is lower. Where I see the burden being the absolute highest is where the other side is using AI and it has not been checked whether that’s a self represented litigant, or a lawyer who doesn’t quite understand AI.

Me:

Many think it is unlikely ever to be solved. An inbuilt ‘feature’ apparently! There is, I believe, a difference between GenAI and an eager intern though. An eager intern may well make mistakes but is unlikely to make up case law that never existed. Furthermore, they will hopefully learn from those mistakes as they are corrected by a supervisor. However, GenAI will not and you are always back to square one when using it.

Robyna May:

I think the analogy is more helpful just in terms of it HAS to be checked. Rather than the nature of the mistakes made. That said, there is much more reward in guiding a new lawyer to think about their reasoning rather than picking out completely fabricated cases. And that perhaps is a whole other discussion entirely.

Me:

Agreed! Although, I have seen some using the analogy as though there is no difference between the two when, of course, there is.

Jennifer Marsh (Product Leader | Former Practicing IP Attorney):

Robyna May – It will not be solved because that’s how GenAI works. It creates the best response, not the correct response. It needs to be used in such a way where that is what you need.

Robyna May:

Agreed. It’s trying to replicate what a great response looks like, and it fills in gaps. It’s interesting that RAG hasn’t closed the gap.

-+-+-+-+-+-

Gerard Stegmaier (Practical Problem Solver, Trusted Consigliere + Lawyer):

Begs and important question about whether AI-native lawyers will have the other skills required for this human-in-the-loop work and where it will come from. Filing of pleadings and issuance of decisions with hallucinations happens for a reason and that reason is merely a new symptom of an older problem—quality assurance.

Me:

Indeed. Might it be the case that those skills will diminish as there is more reliance placed on AI?

-+-+-+-+-+-

Teresa Villa (A trauma-informed technology ecosystem restoring transparency, due process, and civil rights protections through ethical AI. Includes The Justice Engine™, Notary of Justice™, BitOfJustice™, Solutions4Value™, LegalFlow):

What do I think ?
This is exactly why I Built Justice Engine and why it matters.
People verifying legal AI can’t afford:
• hallucinations
• fabricated citations
• made-up rules
• “almost true” answers
The stakes are TOO high.
lived experience fighting wrongful garnishments, misapplied payments, and fake registry numbers

Do you see hallucinations in GenAI output? Here’s my real-world answer transparent, no fluff:
Yes. GenAI does hallucinate.
All models do. Even the best ones still occasionally produce:
• incorrect facts
• outdated info
• misattributed citations
• overconfident answers
• “logic leaps” that sound right but aren’t
Not because they’re malicious but because they’re probability machines, not truth machines.

Do I find myself having to verify GenAI
yes people do have to verify GenAI output.
Especially when:
• the topic is legal
• the conversation involves benefits
• the data affects someone’s rights
• the system must handle trauma safely
• the stakes are high

How much time is that taking me?
Verification can take anywhere from 30 seconds to 30 minutes, depending on:
• whether the user knows the topic
• how complex the question is
• how sensitive the consequences are

Me:

Thanks. Good lists. Also important is what it misses out. The verification burden of ascertaining that can be high.

Teresa Villa:

You’re right the verification burden is enormous.
And that’s actually the part I built JusticeTree around.

Most tools treat verification as an afterthought.
In JusticeTree, verification is the architecture:
• Every output is tied to a source, statute, or record trace
• If the model can’t show its grounding, it isn’t used
• Human review isn’t optional it’s the governing layer
• Overrides create a documented trail, not a productivity penalty
• Survivors and claimants stay in control of their own narrative

So instead of burying people under “prove the AI wrong,” the system reduces that load by design.

In legal contexts, verification shouldn’t be a burden it should be a right.

-+-+-+-+-+-

Michael Lawrence (CTO | Private Intelligence Operating Systems Architect):

This entire debate assumes a world where AI is used as a loose tool giving unstructured answers.
In a governed system — where outputs are jurisdiction-bound, rule-constrained, logged by design, and routed through human oversight at strategic checkpoints — hallucination is engineered out and the verification burden drops to near-zero.
The problem isn’t AI.
The problem is architecture. Juris by CalyxOS

Me:

Indeed. On the whole that, unfortunately, appears to be the world that most lawyers are living in at the moment. I have seen hashtag#LegalTech vendors add a GenAI layer to their product without testing, requesting permission, providing warnings, not giving any training on its use etc. What was previously a governed system suddenly became an ungoverned one. Some vendors are creating rather than solving the architecture problem.

-+-+-+-+-+-

Brian Belt (Founding Shareholder | Experienced Corporate, RE & Hospitality Attorney | Legal Tech Enthusiast):

The article by Joshua Yuvaraj was purely theoretical- no real world testing of the theory. I see no recent real world testing of the supposition. This also does not match our experience. It’s too much easier to verify output from a quality, legal AI vendor than from an associate, since documents and clauses are so readily accessible (hovering over or clicking on links). There is no comparison with the time necessary to verify another human’s output. Everything that I have read, is that when legal AI is applied to the correct use cases, a considerable net time savings results.

Me:

How do you measure “quality”. There are currently many hashtag#LegalTech AI vendors and new ones appearing on the block all the time. Many lawyers are using more readily available GenAI tools often not realising the issues associated with those. Specific LegalTech AI tools are not without problems as research cited by me demonstrates. Before GenAI, LegalTech tools were available to research case law. These were possibly more reliable than they are now with a GenAI layer added. Just because you can doesn’t mean you need to. You would hope a good Associate would cite their sources (even hyperlink them in this day and age). If not, they have not been well trained. And what are the “correct use cases”?

-+-+-+-+-+-

Kevin Keller (General Counsel and Tech Builder | Inventor | Investor | Advisor | co-Founder | Board Member – Helping develop hardware, software and SAAS solutions for 25+ years; developing applications to enhance learning using AI):

This is what validation/verification of AI output looks like in the system we’ve built – every reasoning step and bit of output has provenance and a confidence score – human expert validated facts/reasoning approaches get scored higher, web search and synthetic data lower. You can look at the overall scores and what goes into each step to determine where to dive in more yourself to understand whether and how much to trust the output.

Kevin Kellar - validation - verification of AI output in a system

Me:

Thanks. We need more of this built within GenAI systems.

-+-+-+-+-+-

Paul Correa (Experienced legal leader):

Reductionist arguments on both sides. Neither capture the nuance or value. 1. You are already supposed to “verify” by reading cases, reading documents, knowing your case deeply, and using your expertise. (I require all my legal writers to use pinpoint citations with a parenthetical quote from the case. )So, the “verify” work already existed, if you were doing your job right. 2) Now, AI uncovers more data, it researches more than we were able to before, it brings new insights. That’s additive, adds time, improves quality. 3) Ultimately, if it takes longer to use AI, but the output is better, that’s a gain. So, the added time cost does not cancel out the benefits! You have more work to do and you have better output. The robots aren’t here to do your job for you.

Me:

It does not necessarily follow that “AI uncovers more data, it researches more than we were able to before, it brings new insights”. It depends on what the AI has been trained on (does it cover your jurisdiction, legal subject area and has all the law in that area been digitalised and given to the LLM in question?). It also depends on the prompts given to it. “New insights” might be the hallucinated ones 😉 If the output is better that will be a gain. If it is worse it will not be.

-+-+-+-+-+-

Josh McBride (Barrister at Richmond Chambers):

I’m genuinely puzzled by all this! Have you tried Notebook LM? Or other RAG tools? When deployed alongside tools like Claude, Notion, and Wispr, the efficiencies are well beyond extraordinary. To suggest it takes “longer” because I then need to check the work – produced in seconds/minutes – is risible and demonstrably untrue.

Me:

Notebook LM is for uploading specific documents to and then using it to analyse them. As stated elsewhere in this thread RAG has not eliminated hallucinations. The arguments by me and Joshua Yuvaraj, on the whole, revolve around using GenAI for legal research. When you are preparing for a court case do you rely solely on GenAI for the authorities you cite? You may have found a set of AI tools that work for you. Many lawyers have not as per the 545 cases identified so far by Damien Charlotin.

Nikolaos Papavasiliou (Legal Counsel at Ray White | Independent Legal AI Consultant | LLB(Hons); BSc(BiomedSc); GDipLegPrac; DipSusLiv):

Brian Inkster – agreed that folk unable to establish, integrate, and operate AI workflow / suites properly should not attempt each of same; however, they absolutely should learn how to do so.

advantageous to spec into new mastery where the option is (checks notes) available by asking the product to train you. time commit now is an infitismal contribution (here, a genuine 0-approach if ai progress is accepted-exponential) as against efficiency gains.

qualifying, we could wait for the products to improve before training. there’d be a shorter inflection period from novice to mastery – but that is at the cost of time gained now at a ‘promise’ to self to pick it up later. may as well hold out for entire automative duties in such case.

Josh McBride:

Brian Inkster that’s not what your post says though. You say “increases in efficiency from AI use in legal practice will be met by a correspondingly greater imperative to manually verify any outputs of that use, rendering the net value of AI use often negligible to lawyers”.

That is quite a claim! I and many others struggle with that proposition, given what we are witnessing IRL with these tools.

If you are talking only about legal research (for most lawyers only a small fraction of their working day) then I suggest you try getting Chat GPT to draft bespoke Boolean searches of legal databases for research topics.

It is a game changer. And yes, I rely on this, and yes it saves lots and lots of time.

Me:

Nikolaos Papavasiliou – Thanks for those thoughts. Important too is training on when GenAI could be used and when it probably should not be used in legal practice.

Josh McBride:

Nikolaos Papavasiliou – 💯

Me:

Josh McBride – That is what Joshua Yuvaraj says in his academic paper on the matter which I agree with when it comes to legal research. I trust you have read the paper?

Giving what we are witnessing in real life with the lack of GenAI verification in legal practice and the consequences of that, it is clear that the verification burden is high.

You have not answered my question: When you are preparing for a court case do you rely solely on GenAI for the authorities you cite?

-+-+-+-+-+-

Viraj Deshwal (AI | Physics):

The verification burden is real, but the root problem is deeper. GenAI citations are fundamentally probabilistic.

Current systems return similarity scores for source citation like “95% similar” which are mathematically sound but legally meaningless. No lawyer can present “95% similar” as evidence in court.

Regulated industries can’t work with “the AI said so,” they need deterministic, verifiable proof. Until this changes, the 2:1 verification burden will persist.

-+-+-+-+-+-

Antti Innanen (⚫ Ⓜ️ 💯):

I feel like my postion is pretty solid.

Yes, verification is important.
Yes, verification takes time.

No, it does not render the net value of AI use negligible to lawyers like the ”paradox” states.

I am done arguing this.