
How to Choose a Vertical AI Agency
Buying AI implementation is unusually hard because the market has almost no failure signal. A vendor whose last four agents got quietly shelved looks identical, in a sales meeting, to one whose agents are still running. IDC put the number at 88% of AI proofs-of-concept never reaching production (IDC, 2025) — which means most of the people pitching you have a graveyard they will not mention. This is a buyer's guide for finding out which kind you are talking to, in about two meetings, before you sign anything.
We are a vertical AI agency, so read this with the obvious bias in mind. The questions below are ones we expect to be asked, and a few we would rather not be. Use them on us too.
The ten questions that actually separate builders from decks
Order matters less than insistence. Ask all ten, write down the answers verbatim, and compare across vendors afterward. Vague answers are only obvious side by side.
- What is running in production right now, and who is it running for? Not pilots. Not "in flight." Something processing real cases today with a human accountable for the output.
- What is its accuracy, measured how, on what sample? The measurement method matters more than the number. "94%" with no denominator is noise.
- Show me your evaluation harness. Ask to see the actual file, redacted. If it doesn't exist as an artifact you can look at, evaluation is a slide, not a practice.
- What did the agent get wrong last month, and what did you change? A real operator answers this in specifics within ten seconds.
- Which parts of this workflow should stay human, and why? A partner who claims full automation of a regulated workflow is either inexperienced or selling.
- What happens on the day a model provider deprecates the model you built on? Listen for whether they have already lived through it.
- Who is doing the work — name them. Get the actual engineers on a call before signing, not after.
- What does your handover package contain? Ask for the table of contents from a real one.
- What is the smallest engagement you will accept? A vendor whose minimum is $400K is a vendor who cannot afford to be wrong in front of you.
- What would make you tell a prospect not to build this? The best answer we've heard from a peer was a specific workflow type, not a philosophy.
Answers that should worry you
Some responses are disqualifying on their own. Others are only yellow flags but cluster together in the same kind of firm.
- "We're model-agnostic" as an answer to an accuracy question. Model-agnosticism is an architecture choice, not an evaluation result. It is often used to change the subject.
- Accuracy figures with no denominator or date. Ask "on how many cases, over what period, judged by whom?" The follow-up is where it falls apart.
- Named logos with no described workflow. A logo means someone at that company signed something. Ask what the agent does there, in one sentence, with a noun and a verb.
- "We'll figure out the metrics during discovery." If success criteria are deferred past contract signature, you cannot fail the project, which means you also cannot exit it.
- Refusal to run in shadow mode. Shadow mode means the agent runs on real cases beside your team, producing output nobody acts on, so you can compare accuracy before cutover. A vendor who resists this is protecting their numbers from your inspection.
- No answer to "when is AI the wrong tool?" Every honest practitioner has a list. If the workflow is low-volume, or the rules change monthly, or the ground truth is contested, an agent is a bad investment and a good vendor says so.
How to verify they have shipped anything
Sales references are curated. You need evidence they cannot stage in a week. Four checks, roughly in order of how much they tell you per minute spent.
Ask for a screen share of a running system's observability. Not the app — the logs, traces, and eval dashboard. Production systems have unglamorous artifacts: retry counts, latency percentiles, a queue of flagged cases, someone's name attached to a review. This is very hard to fake convincingly under questions.
Ask the engineer, not the founder, to describe the hardest bug. Real answers are strange and specific — a PDF encoding, a payer portal that changed a field name, a date parsing edge case in a 1994 record. Made-up answers are structural and tidy.
Ask what they cut. Every shipped agent has a scope that shrank. If the vendor cannot name something they wanted to automate and didn't, they have not been through a cutover.
Check for public work. Repos, technical writing, a product of their own. IAW builds SimplOS, Medix, Dona Legal, and Rockets AI, built for the University of Toledo campus — you can go look at those. A firm that has only ever built for clients under NDA has nothing you can inspect, and that is a real cost to you, not a neutral fact. Our own trust page exists so this check takes five minutes.
Structure the first engagement so you can exit cheaply
The single highest-leverage thing you control is engagement shape. You want the first check to buy you a decision, not a dependency.
Structure it in three stages with a real stop between each. Stage one is a short assessment — two weeks, fixed price, delivered as a document you own. Ours is the AI Readiness Map, and the test of any assessment is whether the deliverable is useful if you fire the vendor the next day. If the output is a proposal for more work, it was a sales exercise you paid for.
Stage two builds one agent for one workflow. Insist on a run-alongside milestone (agents working beside your team on live cases) at roughly day 21: the agent runs on live cases, the output goes to a comparison log, nobody acts on it. That milestone is your exit ramp. If shadow accuracy is bad, you have spent one month, not two quarters, and you have real numbers about your own data — which is worth something regardless. Our Production Agent Sprint is built around that gate; the point is the gate, not our version of it.
Stage three is operations, and it should be separately terminable on 30 days' notice. Do the arithmetic yourself before signing: take your current volume, your fully loaded cost per case, and a conservative deflection rate — say half what the vendor projects. If the engagement only works at their number and not at half of it, you are buying a forecast, not a system. The methodology page walks through the sequence in more detail.
Two more terms worth fighting for. Cap total stage-one-plus-stage-two spend at a number your CFO would approve as an experiment, not an initiative. And write the success criteria into the SOW as measurable statements before work starts — the evaluation harness should be written before the agent code, which means the criteria exist on day three, not day sixty.
Contract terms: who owns the evals, the code, the prompts
Most AI services contracts are adapted from generic software templates and are silent on the three assets that actually determine whether you can leave.
- The evaluation harness and the labeled dataset. This is the one people forget and the one that matters most. The harness encodes what "correct" means for your workflow, and the labeled cases represent real expert hours from your staff. Both should be yours outright, delivered in a portable format, with no license back to the vendor for reuse. If you own the evals, any competent successor can pick up the work. If you don't, you are locked in regardless of what the code clause says.
- Prompts, and the reasoning behind them. Prompts are frequently carved out as "vendor methodology" or "pre-existing IP." Sometimes that is legitimate — a firm's general scaffolding predates you. What is not legitimate is treating prompts written against your policy documents as vendor property. Name the distinction explicitly in the contract rather than leaving it to a boilerplate IP clause.
- Code, including infrastructure-as-code. Application code is usually granted. Deployment configuration, CI pipelines, and monitoring setup often are not, and without them the code you own does not run.
- Model and data terms. Get in writing that your data is not used to train any model, that it is not sent to subprocessors you have not approved, and what the deletion timeline is on termination. In healthcare or student-records work this is a compliance requirement, not a preference.
- Exit assistance, priced in advance. Thirty days of transition support at a stated rate, defined as a deliverable list, not "reasonable cooperation."
One note on ownership as a negotiating signal: a vendor who resists giving you the evals is telling you their business model depends on you being unable to measure them. That is worth more information than any reference call.
When not to hire a vertical AI agency at all
Three cases where the honest answer is don't. If your workflow runs fewer than a few hundred cases a month, the engineering will cost more than the labor for years — fix the process instead. If the rules change faster than you can re-evaluate, you will spend your budget maintaining accuracy rather than gaining it. And if your team cannot agree on what a correct output looks like, no agency can build an evaluation harness for you, because there is nothing to encode. That last one is common and fixable, but it is your work, not ours.
If you have gone through the ten questions and want to compare specific answers, our comparison page lays out where a vertical agency beats a platform, a generalist consultancy, or your own team — including the cases where it doesn't. And if you want to run these questions at us directly, say so.
Frequently asked questions
What is the single most important contract term when hiring a vertical AI agency?
Ownership of the evaluation harness and the labeled dataset. The harness defines what "correct" means for your workflow, and the labels represent real hours from your subject-matter experts. If you own both, any competent successor vendor can continue the work and you can measure any vendor's claims independently. If you don't, code ownership won't save you from lock-in.
How can I tell whether an AI vendor has actually put an agent into production?
Ask for a screen share of a running system's observability — logs, traces, latency percentiles, the queue of flagged cases — rather than the application UI. Then ask the engineer who built it to describe the hardest bug they hit. Real answers are oddly specific (a PDF encoding, a portal field rename); fabricated ones are tidy and structural.
What is shadow mode and why should it be in my first contract?
Shadow mode means the agent runs on your real cases alongside your human team, producing output nobody acts on, so accuracy can be compared before any cutover. Written in as a milestone around day 30, it becomes your cheap exit ramp: if the numbers are bad, you have spent one month and still learned something real about your own data.
When is a vertical AI agent the wrong tool?
Three common cases. Volume below a few hundred cases a month, where engineering cost exceeds labor cost for years. Rules that change faster than you can re-evaluate, so budget goes to maintaining accuracy rather than gaining it. And workflows where your own team cannot agree on what a correct output looks like — there is no evaluation harness to write until that is settled internally.
Find your first vertical agent
We inventory your workflows, score them for automatability, and name the one worth owning first. One workflow, live in shadow inside 30 days, one measurable outcome.
Contact us