Practice · Legal AI

eDiscovery automation when you don't have a review army

A twelve-lawyer litigation boutique and an AmLaw 50 firm receive the same 800,000-document production. One of them can put forty contract reviewers on it next Monday. The other one cannot. That asymmetry — not the law, not the merits — decides a lot of mid-market cases. eDiscovery automation is how the smaller side stops losing on capacity. But it only works if you build the defensibility record before you start, not after opposing counsel challenges you.

The math you can do yourself, and where the fit is real

Skip the vendor statistics. Do the arithmetic with your own numbers. Take a review population of 500,000 documents. A competent human reviewer sustains somewhere between 40 and 60 documents per hour on a first pass with a simple relevance call — you know your own team's rate, use it. At 50 per hour, that is 10,000 reviewer-hours. At a $60 blended contract-reviewer rate, that is $600,000, and at 35 review hours per reviewer per week you need 285 reviewer-weeks. Ten reviewers, seven months.

Now run the same population through a classifier that produces a ranked relevance score, and review only the documents above your cutoff plus a validation sample drawn from below it. If the responsive rate is 8% — high for a broad custodian pull — and your model surfaces most of that in the top 15% of the ranking, you are reviewing 75,000 documents plus a few thousand sampled from the null set. That is 1,600 hours instead of 10,000.

The saving is real. What the arithmetic hides is that the 1,600 hours are harder hours, the model costs something to build and validate, and the whole exercise collapses if you cannot explain to a magistrate judge how the cutoff was chosen. The technology is not the hard part. The record is.

First-pass relevance is the right beachhead for exactly that reason. It is high volume, the decisions are coarse, errors are recoverable through sampling, and there is a decade of case law treating technology-assisted review as acceptable — judges have been comfortable with predictive coding since Da Silva Moore in 2012. The novelty in 2026 is not that a machine sorts documents. It is that a language model can read a document, apply a written relevance standard expressed in plain English, and produce a short reason for its call: "Discusses the 2023 supplier renegotiation referenced in RFP 14; identifies pricing terms." Your reviewer confirms or overturns that in eight seconds instead of ninety. The measurable gain is in throughput on the documents you do look at, not only in the ones you skip.

  • Good fit. Relevance and responsiveness against a written standard, issue tagging, custodian and date-range triage, deduplication of near-identical email threads, flagging documents that need translation or technical review.
  • Workable with tight controls. Confidentiality designations, PII and PHI flagging for redaction, first-cut hot-document identification.
  • Assistive only. Privilege. That is where mid-market teams get hurt, so it gets its own section.

Privilege is different, and the asymmetry is brutal

Relevance errors are symmetric and cheap. Miss a responsive document, produce a junk one, and you correct it in a supplemental production. Privilege errors are asymmetric and expensive. One privileged email that goes out the door can trigger a subject-matter waiver argument covering an entire category of communications. FRE 502(d) clawback orders help — get one entered in every case, without exception — but a 502(d) order is a safety net, not a review strategy, and it does not protect you from a court finding that your screening process was unreasonable.

So the design rule is: the agent never decides privilege. It routes. Build a screen that flags candidates on the signals that actually predict privilege — participation of counsel from a maintained attorney list, domain matches for outside firms, legal-department distribution lists, legal-advice language, documents attached to flagged parent emails — and send every flagged item to a lawyer. Set the threshold to over-flag deliberately. A privilege screen with a 30% false-positive rate that catches everything is a good screen. A screen tuned for precision is how you get a waiver motion.

Two failure modes worth naming. First, attachments: an innocuous spreadsheet attached to a privileged transmittal email needs to inherit that treatment, and family-level logic is easy to get wrong. Second, the in-house counsel problem — a general counsel who also runs business operations sends thousands of emails that are not privileged at all, and no classifier reliably separates legal advice from business advice in that stream. That is a lawyer's judgment call. Keep it one.

The record you have to be able to produce

Assume you will be asked to defend your process. Not necessarily in this case, but in the one where the stakes justify a Rule 26(g) challenge. What you need to hand over is not the model. It is evidence that the process was reasonable and that you measured it.

  • The written relevance protocol. The exact standard the agent applied, versioned, with dates. If the protocol changed mid-review, you need to know which documents were coded under which version and whether you re-ran the earlier ones.
  • The validation sample and its results. A statistically valid random sample from the documents you did not review, coded by a human blind to the model's score, with recall and precision estimates and confidence intervals. This is the single most important artifact. Without it you have an assertion, not a measurement.
  • The cutoff decision. Why you drew the line where you drew it, the estimated recall at that point, and who signed off.
  • Every human override. When a reviewer disagreed with the agent, what they changed it to, and when. Overrides are your quality signal and your best evidence of meaningful attorney supervision.
  • The model version and configuration. Which model, which prompt or seed set, which date. Vendors update models silently. If your production spanned a version change and you cannot say so, you have a hole.

If your tooling cannot emit all five without a manual reconstruction project, it is not ready for a contested matter. This is the same discipline we apply everywhere: the evaluation harness gets written before the agent code. In discovery, the harness is the defensibility record. They are the same artifact, and building it second means building it twice.

Automation cuts both ways on proportionality

Rule 26(b)(1) limits discovery to what is proportional to the needs of the case, weighing the parties' resources and the burden of the proposed discovery. Mid-market teams have historically used burden as a shield: we cannot review 2 million documents, narrow the request. That argument gets weaker as review costs fall. If a competent AI-assisted first pass can process the population for a fraction of linear review, "it costs too much" is a harder position to hold, and opposing counsel will say so. Expect it.

The flip side is more useful. Automation makes you credible on the offensive side of proportionality. When you can show what a scoped review actually costs — with your own arithmetic, on the record — you can propose a concrete alternative instead of a general objection: these six custodians, this date range, this issue set, at this cost, versus their proposal at that cost. Judges respond to specific numbers. Bring a declaration with a real estimate and you usually get a better order than the party that just says "unduly burdensome."

Use the same capability on incoming productions. When you receive 400,000 documents from a larger opponent, ranking is how you find the twelve documents that matter in a week rather than a quarter. For a smaller firm, that is often where the leverage is — not on your own production, but on theirs.

How to actually run one without betting the case

Do not turn this on in a live matter as your first attempt. The most common way legal AI projects die is that a firm pilots on the case that matters and gets burned by something a dry run would have caught. Across industries, 88% of AI proofs-of-concept never reach production (IDC, 2025), and in legal the causes are usually procedural rather than technical.

Run the first one alongside your team: the agent codes documents from a closed matter alongside the coding your team already did, and nobody acts on the output. You compare. You get a real recall number against known ground truth, on your document types, before anything is at stake. A closed matter with a completed privilege log is worth more than any vendor benchmark, because the ground truth was produced by your lawyers applying your standards.

Then, on the first live matter: disclose your process to opposing counsel early and get it into the ESI protocol if you can. Courts have consistently favored transparency about TAR methodology over litigation about it after the fact. You do not have to disclose a seed set in every jurisdiction, but you should be ready to describe the validation protocol. Meet and confer on the recall target and agree on the sampling methodology up front, and you have largely eliminated the challenge before it starts.

Staffing changes too. A team of eight reviewers becomes two experienced associates doing exception handling and quality control on a ranked queue. That is a different job with different supervision needs — closer to practicing law, which is a genuine retention benefit, but it means your review lead has to understand sampling. Budget for that training. The related patterns for discovery and contract review automation apply here directly.

When this is the wrong tool

Small populations. If your review is under roughly 20,000 documents, the setup, validation sampling, and process documentation cost more than linear review. Search terms and a couple of associates will beat you on both time and defensibility. Do not build a pipeline to save four days.

Novel or shifting issue definitions. If the relevance standard is still moving because the pleadings are moving, you will re-code the population two or three times, and each re-code costs you the validation work again. Wait until the issues are stable.

Populations that are mostly non-text. Audio, video, engineering drawings, and scanned handwriting still need specialist handling, and a text classifier running on bad OCR produces confident nonsense — worse than no output, because it looks like a decision.

And privilege, again, as decision-maker rather than router. If a vendor tells you their system makes privilege calls at production quality, ask for the validation sample and the false-negative rate on in-house counsel communications. The answer tells you what you need to know about the vendor.

The realistic goal for a mid-market litigation team is not to replace review. It is to make the first pass cheap enough that the cost of discovery stops determining which cases you can take. That is a smaller claim than most of what you will hear, and it is achievable in a single matter. If you want to know whether your document types and matter mix support it, a two-week readiness assessment answers that with your own data before you commit to building anything — or just tell us what you are up against.

Frequently asked questions

Do I have to tell opposing counsel I'm using AI for document review?

There is no blanket rule requiring disclosure, and expectations vary by jurisdiction and judge. In practice, disclosing your methodology early and negotiating it into the ESI protocol is far cheaper than defending it in a motion later. Courts have generally favored parties who were transparent about their process. At minimum, be prepared to describe your validation protocol and recall estimate.

What recall rate should I target for a first-pass relevance review?

There is no legally mandated number, and anyone quoting one as a standard is overstating things. What matters is that you measured recall with a valid random sample drawn from the documents you did not review, that a human coded that sample blind to the model's score, and that you can state the estimate with a confidence interval. Negotiating the target with opposing counsel in advance protects you more than hitting any particular figure unilaterally.

Can an AI agent make privilege calls?

It should not. Privilege errors are asymmetric — one waived document can trigger a subject-matter waiver argument covering a whole category of communications. Use the agent to route candidates to lawyers, tuned deliberately to over-flag, and keep the decision with an attorney. Also get an FRE 502(d) clawback order entered in every case, but treat it as a safety net rather than a review strategy.

How small is too small for eDiscovery automation to be worth it?

Under roughly 20,000 documents, the setup, validation sampling, and process documentation cost more than running a linear review with search terms and a couple of associates. The economics improve sharply as populations grow. Reviewing incoming productions from a larger opponent is often where a smaller firm gets the most leverage per dollar.

Work with us

Scope your first legal agent

Contract review, discovery, deposition prep. We name the narrowest workflow worth owning, run it in shadow on your real matters inside 30 days, then harden it to production.

Contact us