
AI Contract Review: The Real Math
Most pitches for AI contract review start with a percentage. This one starts with a stopwatch and a spreadsheet. If you cannot write down what one contract costs your team today, in minutes, by task, by who touched it, you cannot tell whether an agent helped. You will only be able to tell whether people felt busier or less busy, which is not the same thing and has never survived a budget review.
Below is the arithmetic. Cost per contract, where the hours really go, which judgments a language model can make and which it structurally cannot, how clause libraries and playbooks actually get built, and what a defensible sign-off model looks like. Then an honest section on the contracts where you should not deploy an agent at all.
What a contract costs today
Do this before you buy anything. Take the last 100 executed agreements. Bucket them by type: NDA, MSA, DPA, SOW, vendor paper, order form. For each bucket record three numbers. How many came in last quarter. The median cycle time from receipt to signature. The loaded hourly cost of whoever touched it.
Now multiply. In-house counsel at a fully loaded $180/hour who spends 45 minutes on a vendor NDA costs you $135 per NDA. Four hundred NDAs a year is $54,000 of attorney time on the least differentiated document in the building. An MSA that takes six hours across two reviewers costs closer to $1,100. Outside counsel at $650/hour changes the denominator, not the method.
The second number matters more than the first, and almost nobody tracks it: queue time. A contract that takes 45 minutes of work but sits eleven days waiting has a cycle-time problem, not a labor problem. Automation that halves the 45 minutes and leaves the eleven days untouched moves a metric your CFO does not care about. Measure both, separately, before you scope anything. We put this instrumentation in during the AI Readiness Map precisely because so many teams discover their bottleneck is routing and approvals rather than reading.
Where the hours actually go
Sit with a reviewer and timestamp a real MSA review. The distribution is consistently lopsided, and not in the direction the software demos suggest.
- Locating. Finding the indemnity clause, the limitation of liability, the assignment provision, the governing law. On a 40-page agreement with inconsistent defined terms this is a real chunk of clock and it involves no legal judgment at all.
- Comparing. Holding the counterparty's language next to your standard and articulating the delta. Not judging it yet. Just stating precisely what changed.
- Judging. Deciding whether the delta is acceptable, acceptable with a fallback, or a walk-away. This is the part people went to law school for.
- Drafting. Writing the redline and the cover note that explains it to the business owner.
- Chasing. Emails, approvals, version control, discovering that procurement already conceded something on a call.
Locating and comparing are frequently half the clock. Chasing is workflow, and a model does not fix workflow. Judging is the irreducible core, and it is smaller than it feels. That asymmetry, a large mechanical fraction wrapped around a small judgment core, is what makes contract review automation worth doing. It is also exactly why replacing the whole review is the wrong goal.
What an agent can and cannot judge
Sort every review task by whether a correct answer exists inside the four corners of the document plus your written standards. That line is the deployment boundary.
Inside the line: clause extraction and classification, presence and absence checks against a checklist, deviation detection against a template, defined-term consistency, cross-reference integrity, date and dollar arithmetic, obligation extraction into a calendar, comparing a stated liability cap against your stated floor. These are verifiable. A reviewer can look at the output and say right or wrong in seconds, which means you can build an evaluation set and score it. In our practice the evaluation harness is written before the agent code, because you cannot improve a system whose accuracy you cannot measure.
Outside the line: whether to accept an uncapped data-breach carve-out from a strategic customer you are two weeks from closing. Whether a clause that is technically fine reads as bad faith and poisons the relationship. Whether an ambiguity is worth fighting given the counterparty's litigation history and your leverage this quarter. Whether the deal is even a good idea. None of that is in the document. The model has no access to the facts that determine the answer, and confident-sounding output on these questions is the most dangerous thing the system can produce.
The practical consequence: build the agent to produce a first-pass markup with citations, never a decision. Every flag points to a page and line. Every recommendation names the playbook rule it came from. If the agent cannot cite, it does not speak.
Clause libraries and playbooks
Here is the part vendors skip. An agent is only as good as the written standard it compares against, and most legal teams do not have one. They have a senior lawyer with twelve years of pattern recognition and a folder of past deals.
A usable clause library has, per clause type: your preferred language, your first fallback, your second fallback, and your walk-away condition stated as a test rather than a vibe. "Liability cap below 12 months of fees paid is a fallback. Below 3 months, or uncapped indemnity for our own breach, is a walk-away." Write it that way and the agent can apply it. Write "we generally prefer a reasonable cap" and you have built a random number generator.
Extracting this takes real work. The method that works: pull 50 executed contracts of one type, have the agent extract every instance of a given clause, cluster them, and show the lawyer what the company actually agreed to over two years. That session is uncomfortable and extremely productive. It reliably surfaces positions the team believed were firm and were not. You are not documenting policy, you are discovering it. Budget two to three working sessions per contract type, and start with the highest-volume type rather than the highest-value one.
One more discipline: version the playbook like code. When a rule changes, every downstream flag changes. If you cannot say which version of the playbook produced a given markup six months ago, you cannot defend the markup. Treat it the way you would any operated system. Our methodology covers how the evaluation set and the playbook move together.
The sign-off model
A lawyer signs off. Always. The question is what they are signing off on, and how long it takes them. Three tiers work well:
- Tier 1, mechanical, agent-complete, spot-checked. Extraction, checklist coverage, obligation capture, arithmetic. A human reviews a sample, not every instance. The sampling rate falls as measured accuracy holds.
- Tier 2, agent-drafted, human-approved. Deviations against the playbook, proposed redlines drawn from your own fallback language. The reviewer reads the flag, checks the citation, accepts or overrides. This is where the time actually comes out.
- Tier 3, human-only, agent assists. Anything novel, strategic, disputed, or above a dollar threshold you set. The agent may summarize and locate. It does not recommend.
Two mechanics make this hold. First, every override is captured with a reason code. That log is your training data, your playbook backlog, and your evidence that oversight is real rather than nominal. Second, the agent must have a route to "I don't know." A system that always produces an answer teaches reviewers to trust it uniformly, which is the failure mode you are trying to design out. Test for abstention on low-confidence extractions explicitly.
And run it in shadow mode first. The agent processes real contracts alongside the human team, produces output nobody acts on, and you compare. That is how you learn your true accuracy before anything is at stake. It is also a large part of why so many pilots die as pilots: IDC put the share of AI proofs-of-concept that never reach production at 88% in 2025, and the ones that cross over generally did the boring measurement work first.
Where it fails
Be direct with yourself about these. They are not edge cases you will engineer away next quarter.
Bad scans. A faxed, hand-annotated, signed amendment is a document-quality problem before it is an AI problem. An OCR error inside a numeric liability cap is silent and catastrophic. Gate on document quality and route failures to a human queue rather than letting the system degrade quietly.
Incorporation by reference. The MSA points to an online terms page that changed in March, plus a purchase order with conflicting terms, plus an amendment nobody uploaded. The agent reviews what it can see and will not tell you what is missing unless you build explicit completeness checks. For most teams the document repository, not the model, is the binding constraint.
Negotiated ambiguity. Sometimes both sides knowingly left a clause vague because clarity would have killed the deal. A model reads that as a defect and proposes a fix that reopens a settled fight.
Novel structures. A first-of-its-kind commercial arrangement has no template to compare against. The agent will confidently map it onto the nearest familiar pattern, which is precisely the wrong answer.
Regulatory overlay. Whether a DPA satisfies a specific supervisory authority's current position is a question about the world, not about the document. Same for export controls, sanctions screening, and anything where the rule changed after your playbook was written.
Adversarial drafting. Sophisticated counsel bury the operative term in a definitions section or a schedule. It is not hidden from a careful reader, but it is placed to defeat pattern matching.
So the honest scope. An agent takes real time out of locating and comparing, drafts a defensible first-pass markup on high-volume standard paper, and gives your lawyers back the hours they should have been spending on judgment. It does not review your contracts. Someone with a license still does that, and should.
If you want this arithmetic run on your own contract mix rather than a generic model, that is what the two-week readiness map produces: a costed baseline and a scoped first agent. See how we scope legal review work, or start with the numbers.
Frequently asked questions
How much does AI contract review actually save?
It depends entirely on your mix, which is why you should compute it rather than accept a vendor figure. Take your annual volume by contract type, multiply by median review minutes and loaded hourly cost, then estimate what fraction of those minutes goes to locating clauses and comparing them to a standard. That fraction is the addressable pool. Judgment time and approval queue time are not addressable by an agent, and in many teams the queue time is the larger number.
Can an AI agent replace a lawyer for contract review?
No. An agent can produce a first-pass markup with citations against a written playbook, extract obligations, check completeness, and flag deviations. It cannot weigh commercial leverage, relationship risk, litigation posture, or facts that live outside the document. A licensed reviewer signs off in every tier of the model described above. What changes is how long that sign-off takes.
What do we need in place before deploying contract review automation?
A written clause library with preferred language, ranked fallbacks, and walk-away conditions stated as testable rules. A set of past executed contracts to build an evaluation harness against. And a document repository complete enough that the agent is not silently reviewing a fragment. Most of the real project work sits in those three items, not in the model.
What is shadow mode and why does it matter for contract review?
Shadow mode means the agent runs on real incoming contracts alongside your existing team, producing output nobody acts on. You compare its markup against what the humans did, on live work, before anything depends on it. For contract review this is the only credible way to learn your true extraction and deviation-detection accuracy on your own paper rather than on a vendor's benchmark.
Scope your first legal agent
Contract review, discovery, deposition prep. We name the narrowest workflow worth owning, run it in shadow on your real matters inside 30 days, then harden it to production.
Contact us