A new Redgrave LLP study shows generative AI beating a managed TAR review at responsiveness classification – and quietly demonstrates why classification was never the hard part. Neil Cameron reads the small print.

These numbers will be on every vendor slide by Friday. This month Redgrave LLP published a working paper, Generative AI for Complex Document Review. Robert Keeling and colleagues ran two reviews on the same set of 45,004 documents from the public Mallinckrodt opioid archive, under the same hard responsiveness test. One review used Relativity aiR for Review, a leading generative AI tool. The other used Relativity Active Learning with a 24-person review team. The AI found 88% of the responsive documents. The human team found 64%. The gap is statistically real, not noise. The AI also left fewer responsive documents behind. The trade-off was precision: the AI flagged more documents that turned out not to be responsive. But the headline is the effort. The AI workflow took 18 hours of one lawyer’s time. The human review took 24 people and 1,123 hours.

Credit where it is due: this is serious work. The test was not simple topic-spotting. A document counted as responsive only if it showed compliance with, violation of, or reckless disregard of US rules on pharmaceutical marketing and controlled substances. That is real legal judgement, applied at scale. Where a close call could have gone either way, the authors counted it against the AI. The paper treats generative AI as a new engine inside the established TAR validation framework (the Sedona Conference’s TAR 1 Reference Model), and within that frame it does very well. Anyone still saying AI cannot apply a complex review protocol to a yes/no call should drop the claim.

But the most interesting findings are in the detail, and there are three of them.

The yardstick wobbles

The whole study rests on one yardstick: a single subject-matter expert reviewed a sample of 1,000 documents, blind, and coded 73 as responsive. By definition, that expert scored perfect recall and precision – the rest of the study is measured against his calls. Then came a second test. The authors showed him the AI’s reasoning and citations for documents he had marked not responsive. He changed his mind on ten of them, and moved six more to “borderline”. So his original blind coding had missed at least 10 of 83 responsive documents – about 12%. And 12% is the floor, not the ceiling: they only checked one direction. Nobody tested whether the AI’s reasoning could have talked him out of any of his responsive calls. So, the 88% system was being graded by a marker who was himself only about 88% complete. You might say the expert was simply corrected, not wrong – but that concedes the point, because a marker who needed correcting was incomplete. The other reading, that he was persuaded rather than corrected, is worse, and we come to it below. Even for yes/no calls, “ground truth” is a human judgement, and it moves.

The validation theatre finding

The result in section 5.1 is the one to remember. The human team’s own quality check – sampling the documents it had set aside, in the standard way – reported that it had caught 100% of the responsive material. But when the whole population was sampled at random, the real figure was 64%. Here is why the check missed it: when a reviewer wrongly marks a responsive document as not responsive, that document is dropped from production and then fed back in as training data. The mistake teaches the system, and the standard check has no way to see it. So a long-established, court-accepted method, used correctly, certified as near-perfect a review that had missed more than a third of the responsive documents. To be fair, the authors are clear that the 64% reflects this particular review team and set-up, not active learning in general. That is true – and it is the point. The lesson is not that TAR validation is worthless. It is that validation has to cover the whole combined process and system, people and machine together, not just the algorithm. The law requires lawyers to make a reasonable check that their disclosure is sound, and that duty is not met by a number the tool reports about its own work. That is true on both sides of the Atlantic — Rule 26(g) in the US, the reasonableness and proportionality rules of PD 57AD in England and Wales. You can’t mark your own homework.

The persuasion finding

Back to those ten changed calls. The authors read them, fairly, as the AI catching responsive material a human reviewer had missed. But look at what the study cannot show. There was no second expert sitting above the first, so nothing tells us whether the AI’s reasoning was right or merely convincing. And the test only ran one way. The paper then suggests using the tool exactly like this in practice – putting fluent, well-cited AI reasoning in front of the very human who is supposed to be checking it. Recall and precision tell you whether the label was right. Nothing in the study, or anywhere else yet, tells you whether the reasoning was right – and that reasoning changed the expert’s mind in 16 of 151 close calls. There is a name for the well-known risk that people scrutinise something less carefully the more polished it looks: automation bias. This study has caught it happening in real conditions and presented it as a selling point.

The boundary the study cannot cross

Here is the structural point. Every result in this study is a yes/no label. Not one summary, chronology, privilege log entry, or case narrative was produced or tested – and the framework the paper relies on openly leaves those tasks out, as its own authors say. Yet the same software licence that buys the classification also buys summarising, charting, and narrative-building, one click away. Lawyers are already relying on those outputs in live cases every day, with nothing like this study behind them – and no clear way to build one, because there is no yes/no answer key to test a chronology against. Citations do not solve it either. As Maura Grossman has pointed out, a citation tells you the right document was looked at, not that the right meaning was taken from it. So the significance of the Redgrave study is not that it erases the line between TAR and what is now called HAR (human AI-assisted review). It is that it confirms the line, exactly where I have previously drawn it (see figure). The shared zone, where AI does classification, now has solid evidence behind it. The interpretive work that fills vendor marketing and conference panels has none.

The TAR/HAR relationship: binary categorisation occupies the shared zone, where statistical validation transfers; interpretive synthesis is HAR-only territory. The Redgrave study operates entirely within the overlap.

That gap – and what a framework to close it might look like – is the subject of my forthcoming article in the Federal Courts Law Review. The short version: the profession should build the answer before a judge demands one.

The authors’ Law.com headline, “Better Than TAR. Nearly Expert”, slightly overstates the case. What the study shows is that an AI classifier can beat one traditional active-learning workflow on a yes/no task – that AI can be better at TAR than some existing TAR set-ups. “Better at TAR” is accurate. “Better than TAR” is not. The study says nothing about the interpretive work discussed above.

“Redgrave’s most useful contribution may turn out to be an accident. When they showed the expert the AI’s reasoning and watched him change his mind, they built – without quite meaning to –  a rough draft of the experiment the field needs next: run it in both directions, put an independent expert above the first one, and measure when the AI’s reasoning genuinely corrects a human and when it just persuades one. Until someone runs that study, “nearly expert” cuts both ways. Generative AI has now aced the exam that TAR sat. The exams nobody yet knows how to set are the ones it takes every day.

Neil Cameron is the Lead Analyst with Legal IT Insider, and an academic. His article “Beyond TAR: Validating Human AI-Assisted Review in eDiscovery” is forthcoming in the Federal Courts Law Review. The Redgrave working paper is available at redgravellp.com.