Here's the trap. The demo shows the software writing a personalised email and it is genuinely impressive, so you extrapolate from that to the whole job. Writing the email was never the hard part. Your SDR does about nine things and writing is the one that was always going to automate first.
The software is strong at building the list, researching the account, drafting and sending, and weakest at handling the reply. Everything before the reply is volume. The reply is where a deal actually begins, and it is the step with the least evidence behind it.
So this page goes task by task through what an AI SDR does, marks each one does it, does it badly, or does not touch it, and then prices the parts separately so you can see what the bundle is charging for.
Task by task, what does it actually do?
| Task | Verdict | The catch |
|---|---|---|
| Build the list | Does it | Through a data partner, so you pay full-stack prices for a layer sold separately at 15 to 83 cents a contact |
| Research the account | Does it | No vendor publishes any measure of research quality, so the claim cannot be checked either way |
| Draft the message | Does it well | The genuine capability, and the only one every demo shows |
| Send it | Does it | Deliverability stays your problem. The domains and their reputation are yours |
| Handle the reply | The gap | The least evidenced step, and the one that converts |
| Book the meeting | Partly | Usually hands off to a scheduler rather than negotiating a time |
| Log it to the CRM | Does it | Genuinely solved and worth having |
Why is list building not the win it looks like?
Because the AI SDR is not the one finding the contacts. It buys or licenses them from a data provider, the same way you would, and the coverage question belongs to that provider rather than to the software wrapped around it.
That matters for pricing. A contact with a verified email and a mobile number costs between 15 and 83 cents to buy directly, which we set out in the cost per contact comparison. Inside a full-stack contract that layer is invisible, and you are paying for it at a very different rate.
What about research and personalisation?
It does it, and nobody can tell you how well. No vendor in this category publishes a measure of research quality, and there is no independent test. That cuts both ways: the claims cannot be verified, and they cannot be disproved either.
What can be said is that this is where the early failures were loudest. Artisan's chief executive has acknowledged, in reported comments, that the product had problems with bad hallucinations in its early period alongside low response rates. We have that through secondary coverage of an interview rather than from the interview itself, so treat it as reported rather than quoted.
Where exactly does it fall down?
The reply. A prospect answers with a question, an objection, a reschedule request or something ambiguous, and the software has to decide what that means and what to say next. That is the step your SDR is actually good at, and it is the one with the least published evidence in the whole category.
The sharpest evidence available is a customer's own assessment. ZoomInfo, which used 11x, told TechCrunch the product "performed significantly worse than our SDR employees." That is a buyer describing a deployment rather than a competitor selling against one, which is why it carries weight that most criticism in this category does not.
It is also worth noticing what is not published. No vendor here reports a reply-handling accuracy rate, a rate of correctly identified objections, or a rate of conversations escalated to a human. Those would be the obvious metrics if the numbers were good.
What do the parts cost separately?
Clay at $185 a month for the research and enrichment layer, plus Instantly at $47 for sending, comes to $232 a month, or $2,784 a year. The median full-stack AI SDR contract in our records is $45,000 a year.
That is a ratio of about sixteen to one, and it is not a like-for-like comparison. The $2,784 stack needs a person to run it. The $45,000 one is sold on the premise that it does not. So what the premium actually buys is the reply handling and the accountability, and the reply handling is the least evidenced part of the product.
Put another way: you are paying sixteen times for the step nobody will publish a number about. That is not an argument against buying one. It is an argument for asking about that step specifically, in writing, before you do.
So what still needs a person?
- Deciding who to contact. The software builds the list from criteria you define. Getting the ICP wrong produces immaculate emails to the wrong people, faster than before.
- The offer itself. No amount of personalisation rescues a message that has nothing to say. This is the most common failure and the one least often blamed on the tool, correctly.
- Anything past the first reply. Assume you are handling these until a vendor shows you a number saying otherwise.
- Domain and deliverability management. The sending reputation belongs to you whoever presses send, which is why the sending infrastructure is worth budgeting separately.
The answer to the question in the title is that it does six of the seven steps and does the seventh badly, and the seventh is the one where revenue happens. Whether that is worth sixteen times the cost of the parts depends entirely on how many meetings it holds, and our arithmetic puts break-even against an in-house SDR at about eight held meetings a month.