"Automated outbound" is sold as one thing. It is six jobs in a row, and the software is genuinely good at three, mediocre at two, and not doing the sixth at all, regardless of what the pricing page implies.
Here is the split job by job, with the evidence, plus two pieces of arithmetic that change how you should read every benchmark and every A/B test you run.
Job 1: Build the list · mostly automated
What software does: filters a contact database by firmographics, title and technographics; exports with verified emails. "Waterfall" enrichment tools query several providers per contact to raise match rates.
What you still do: define the ICP, and audit the output. Every database markets accuracy in the nineties; independently reported testing puts real-world accuracy across a much wider band, and the spread varies sharply by region and seniority. Treat claimed accuracy as marketing and measure your own.
The test that costs nothing: pull 200 records from your actual ICP (not the vendor's sample) run them through a verifier, send to them, and record the bounce rate. Then extrapolate. A 6% bounce rate on a 20,000-record list is 1,200 hard bounces, which is a domain-killing event you can detect for the price of 200 records.
Verdict: automate it, then sample it. The list is the biggest single lever on results and the easiest thing to be quietly wrong about for months.
Job 2: Research each prospect · partly automated, badly
What software does: scrapes the company site, LinkedIn, funding announcements, job postings, tech stack, and produces a paragraph of "insight" per prospect.
What you still do: decide whether the insight is a reason to email. This is the most visible failure in AI outbound. These tools reliably find a fact and unreliably find a relevant fact. "I saw you're hiring three AEs" is a fact. Whether that implies a need for your product is an inference, and it is the inference the model makes badly, usually by asserting a pain point it has no evidence for.
That assertion is precisely what makes a cold email read as generated, and it is what costs the reply.
Verdict: automate the gathering, keep the judgement. Concretely: read 20 outputs end to end every time you change the prompt, before any of that batch sends. Not a sample of the sends: a sample of the drafts.
Job 3: Write the email · automated, with a hard ceiling
What software does: produces grammatical, correctly-formatted, personalised-looking openers and bodies at any volume.
What you still do: supply the angle, and edit. The ceiling isn't grammar. It's that the model will confidently assert things it cannot know (a priority, a frustration, a reason the reader cares right now) and those assertions are the tell.
Verdict: you write the template and the angle; the tool fills variable parts; you read a sample of every batch. Never send an unread generated draft at volume.
Job 4: Send it · fully automated, and this is the actual product
What software does: rotates across mailboxes, throttles per-inbox daily volume, randomises intervals, handles timezones, stops sequences on reply, manages the suppression list, injects unsubscribe headers, warms new inboxes. Unglamorous infrastructure, and where sequencing tools genuinely earn their price.
What you still do: own the DNS and the volume decisions. Non-negotiable, and no sender tool fixes them:
- Google requires SPF, DKIM, DMARC, valid forward and reverse DNS, TLS, RFC 5322 formatting, one-click unsubscribe, and a spam rate under 0.30% for senders above 5,000 daily messages to Gmail, with the From: domain aligned to the SPF or DKIM domain.
- Microsoft, since 5 May 2025, rejects non-compliant high-volume mail with
550 5.7.515rather than junking it. Your sequencer reports that as a hard bounce, which looks like a list problem, which sends you to re-verify a list that was fine. Learn the code. - The 5,000/day threshold counts across the primary domain, subdomains included, and once crossed, bulk-sender status is permanent.
The copy-paste records are here.
Verdict: automate all of it. This is the part worth paying for.
Job 5: Follow up · fully automated, and routinely overdone
What software does: fires steps 2 through 6 on a schedule, stops on reply, branches on opens and clicks.
What you still do, two things, both easy to get wrong:
Decide the touch count by hand. Automation makes eight follow-ups exactly as cheap as three, which is why so many sequences run eight. The constraint isn't cost, it's complaint rate, and complaint rate is measured, daily, with Google's guidance being under 0.10% and never reaching 0.30%. At 0.10% you're allowed roughly one spam complaint per 1,000 sends. Touches five through eight are where complaints come from.
Stop branching on opens. Since Apple's Mail Privacy Protection and equivalent proxy-prefetching elsewhere, a large and unknowable share of "opens" are machines. A sequence that branches on open is branching on noise, and an open rate reported to two decimal places is false precision on a corrupted signal.
Job 6: Handle the reply · not automated, whatever you were told
What software does: routes replies to a shared inbox, classifies them as interested / not interested / out-of-office, sometimes drafts a response.
What you still do: all of it. The moment a human replies you are in a conversation with context, objections, timing and a calendar, and this is exactly where "AI SDR" products hand back to a person.
The evidence, rather than the assertion. TechCrunch's March 2025 investigation into 11x (one of the best-funded companies in this category, backed by a16z and Benchmark) reported that multiple companies displayed as customers were not customers; ZoomInfo stated "We did not give them permission to use our logo in any manner, and we are not a customer," and described a one-month trial in which the product performed worse than their own SDR employees. Three current and former staff told TechCrunch most early customers exercised break clauses to exit, citing the emailing product not working as expected and hallucinations. The CEO stepped down six weeks later.
That is one company, and the category is not one company. But it is the best-documented case available, it is reported rather than claimed, and the failure mode it describes (everything up to the reply works, the reply handling doesn't) is the structural one.
Verdict: do not plan a pipeline around this being automated. Budget the headcount. If a vendor promises autonomous reply handling and booking, ask specifically what happens on a reply that is neither a yes nor a no, which is most of them.
The two pieces of arithmetic that change how you read everything
1. The denominator problem: 0.45% and 3.7% can be the same campaign
Reply-rate benchmarks in this category are not comparable, because they use different denominators and nobody states which.
Belkins' 2026 study of 7,530,489 emails reports a 0.45% average reply rate: measured against total sends, excluding auto-replies, out-of-office and bounce notifications. Saleshandy publishes 3.7%, measured against emails delivered. Belkins itself notes that "a 5% reply rate against openers and a 0.45% reply rate against total sends can describe the same campaign".
These can all describe one campaign. Replies ÷ sends, replies ÷ delivered, and replies ÷ opened are three different fractions, and the third is computed on the open signal that Job 5 above just established is unreliable.
What to do: define your own denominator as replies ÷ emails sent, excluding auto-replies, and never compare your number to a published benchmark without checking what they divided by. When a vendor quotes a lift, ask for the denominator before asking for the percentage.
2. Your A/B tests are almost certainly underpowered
At a 0.45% baseline reply rate, here is what it takes to detect a difference at 95% confidence and 80% power (two-proportion test):
| Comparison | Relative lift | Sends <strong>per variant</strong> | Total sends |
|---|---|---|---|
| 0.45% → 0.90% | +100% (doubling) | ~5,200 | ~10,400 |
| 0.45% → 0.60% | +33% | ~36,400 | ~72,800 |
| 0.45% → 0.50% | +11% | ~330,000 | ~660,000 |
Now compare that to how these tests actually get run: two subject lines, 500 contacts per variant, decision made on Friday. At 500 sends and a 0.45% rate you expect 2.25 replies per arm. Three replies versus one is not a result: it is the difference between two and four coin flips, and it will reverse itself next week.
What this means in practice:
- You cannot A/B test subject lines at normal cold email volumes. Not "it's hard": the sample sizes aren't available to you.
- What you can test is big and structural: a different ICP segment, a different offer, a different channel. Changes large enough to double the rate need ~5,200 sends per arm, which is reachable.
- Everything smaller is a judgement call: make it on craft, and stop pretending it was data.
- Run one variant properly rather than four badly. Splitting 2,000 sends four ways guarantees four uninformative results.
This is the most expensive misunderstanding in outbound: teams spend months optimising copy on sample sizes that can't detect a 33% improvement, and conclude the channel doesn't work.
Where this leaves the buy decision
| Job | Automatable? | Who actually owns it |
|---|---|---|
| 1 · Build list | Yes, with sampling | Commodity software + your 200-record audit |
| 2 · Research | Gathering only | A prompt-and-review discipline, not a product |
| 3 · Write | Drafting only | Your template, model's variables, your read |
| 4 · Send | Fully | The real product. Buy it. |
| 5 · Follow up | Mechanically yes | Touch count is a policy decision disguised as a feature |
| 6 · Reply | No | A person |
An all-in-one "AI SDR" bundles 1–5 and charges for the appearance of 6. A modular stack buys 1, 3 and 4 separately, usually for less, and makes 2, 5 and 6 explicitly your job, which they are in both cases. The difference is whether that's stated on the invoice.
The minimum honest setup
- A separate root sending domain with SPF, DKIM and DMARC verified by reading a received header.
- Mailboxes ramped before real sends: manually if there are few, with the warmup your sender already bundles if there are many.
- A list bounce-tested on a 200-record sample of your real ICP.
- A template you wrote, with the model filling variable parts, and 20 drafts read per batch.
- Three to four touches, not eight.
- One-click unsubscribe (RFC 8058) honoured within 48 hours.
- Postmaster Tools open weekly (spam rate, domain reputation, delivery errors) knowing the spam-rate figure covers only personal Gmail, not Workspace recipients.
- A defined denominator for your reply rate, and no A/B test below ~5,000 sends per arm.
- A human on the inbox.
Eight of those nine are setup you do once. The ninth is the recurring cost nobody quotes.