ColdOps
All posts
Benchmarks & metrics

We Analyzed 800 Cold Email Tool Reviews. Here's What Breaks

By Mo Charradi, Founder9 min read

Every review of every cold email tool seems to be written by someone selling you a different cold email tool. So we pulled the data instead: the 100 most recent G2/Capterra reviews for each of the 8 biggest cold email senders — 800 reviews total, roughly 70% from the last 12 months — plus the low-star pages on Trustpilot and a sweep of operator communities. No affiliate links in this post, no tool we're trying to sell you. Just what the reviews actually say, including the parts the headline ratings hide.

How we ran the analysis

We collected the 100 most recent reviews per tool for Instantly, Smartlead, Apollo, Lemlist, Woodpecker, Mailshake, Reply.io, and Saleshandy through a commercial review scraper, normalized every rating to a 5-point scale, and counted complaint themes by matching the pros/cons text against keyword families (support, pricing, deliverability, bugs, UX, data quality, billing). We also recorded which reviews the platform itself flags as incentivized — left by a reviewer who received something for writing it.

Two honest caveats. Theme counts are keyword-matched, so treat percentages as directional, not gospel. And a most-recent-100 sample is deliberately recency-weighted — it tells you what users think now, which is exactly why it sometimes disagrees with a tool's all-time headline rating. That disagreement turned out to be one of the findings.

The scoreboard

ToolAvg rating (recent 100)Reviews ≤3★Marked incentivized#1 complaint theme
Lemlist4.6521%Pricing
Mailshake4.5850%Support
Instantly4.51421%Pricing
Saleshandy4.4572%Pricing
Reply.io4.4260%Support
Woodpecker4.42727%Support
Apollo4.29663%UX / learning curve
Smartlead3.292543%Deliverability

The average across all 800 reviews is 4.33 out of 5, which sounds healthy until you look at the two columns on the right.

The headline rating and the recent reviews disagree

The most striking number in the table is Smartlead's. Its all-time rating on G2 sits around 4.6. Its most recent hundred reviews average 3.29, with a full quarter of them at three stars or below — and that's with 43% of the sample marked as incentivized. The two loudest themes in those reviews are deliverability (raised in about a third) and support (raised in nearly as many). Recent sentiment and the badge on the homepage are telling two different stories.

Apollo runs the same pattern in a milder form: a 4.29 recent average held up by the highest incentivized share we measured — 63 of its 100 most recent reviews carry G2's incentivized flag. Woodpecker's positive reviews skew old; only a small fraction of its recent-100 came from the last year at all.

None of this means the tools are bad at everything. It means the number you see on a review badge is an average over years of mostly-solicited feedback, and the experience you'll have next month is better predicted by what the last hundred reviewers said.

What actually breaks, by the numbers

Counting complaint themes across all 800 reviews:

ThemeShare of reviews mentioning it
Support~17%
Pricing / billing~17%
Deliverability~13%
UX / learning curve~11%
Missing features / limits~7%
Bugs / reliability~6%
Data quality~3%

Support and pricing being the universal top two surprised us less after reading the actual text. The support complaints cluster around the same moment: something breaks mid-campaign — mailboxes disconnect, deliverability drops, sends silently stop — and the response is a chatbot, a 24–72 hour email loop, or an escalation that never lands. Support quality gets judged at the worst possible time, which is precisely when it matters.

Deliverability at 13% understates its weight, because G2 reviewers skew toward happy customers writing feature requests. On the low-star pages and in operator communities, deliverability is the reason people leave: campaigns that collapsed overnight, blacklistings discovered only after replies dried up, bounce rates that burned domains before anyone noticed.

The zero-star tail: billing is where trust dies

Read the one-star reviews across all eight tools and one pattern repeats more than any product flaw: money. Reviewers report being charged after cancelling, prepaid credits expiring or disappearing, refund requests refused on technicalities, and annual contracts that were harder to leave than to sign. We saw versions of this on nearly every tool in the sample — it's a category habit, not one vendor's sin.

If you're evaluating a tool, this is the cheapest due diligence available: skip the 5-star wall, open the 1-star tail, and read what happened when someone tried to leave. How a vendor behaves at cancellation tells you more than how it behaves at signup.

The dashboard problem

The deepest theme in the data isn't about any single tool. Reviewers keep describing the same experience from different angles: the dashboard said everything was fine, and it wasn't.

Warmup scores read 90+ while live campaigns land in spam — a contradiction so common that operators now openly debate whether warmup tools work at all in 2026, and some report better results after turning warmup off. Delivery rates show 94% while replies flatline, because "delivered" includes the spam folder. Open rates inflate past sends thanks to scanners and inbox assistants — we've written before about why open rates dropped out of nowhere. Every one of these numbers is generated by the tool grading its own homework.

The practical conclusion from 800 reviews is not "pick the tool with the best score." It's: trust campaign outcomes — replies, bounces, meetings — over any dashboard the sending tool renders about itself, and watch those outcomes closely enough to catch the break early. That's a monitoring habit, not a tool choice. For agencies running many clients across these platforms it's the whole job, which is why we built ColdOps to watch campaign reality across every workspace independently of what any sender's dashboard claims. But even doing it manually — a weekly look at reply and bounce trends per client — beats believing a green score.

What to do with this if you run cold email for clients

A few practical moves that fall straight out of the data:

  1. Weight recent reviews, ignore the badge. The last 100 reviews predict your experience; the all-time average predicts nothing.
  2. Read the 1-star tail before you enter a card number. Billing behavior at cancellation is the category's biggest trust gap.
  3. Assume support will be slowest during your worst week. Have your own runbook for deliverability incidents instead of betting on a ticket queue — the failure signs are visible in campaign data before a client notices.
  4. Treat every first-party score as marketing. Warmup scores, health badges, delivery percentages. If it can't be verified from campaign outcomes, it's a claim, not a measurement.

The tools themselves are fine — millions of cold emails go out through them every day. The reviews just make the terms of the deal unusually clear: the sending is reliable, the self-reporting isn't, and nobody is coming to tell you when it breaks.

Frequently asked

What do cold email tool reviews actually complain about?
Across 800 recent reviews of the 8 biggest cold email tools, the most common complaints were support quality (mentioned in ~17% of reviews), pricing (~17%), and deliverability (~13%). Billing and cancellation problems dominate the lowest-rated reviews on every platform we checked.
Are G2 reviews reliable for cold email tools?
Partially. G2 skews positive — most criticism comes as feature requests inside 4–5 star reviews — and a meaningful share of reviews are marked incentivized by G2 itself (63% of Apollo's recent reviews, 43% of Smartlead's in our sample). Read the most recent reviews and the 1–3 star tail, not the headline average.
Why do cold email tools have such different ratings on G2 vs Trustpilot?
G2 collects reviews through vendor campaigns, often incentivized, so scores cluster high. Trustpilot reviews are mostly unprompted, often written mid-crisis, so they carry the severe problems: billing disputes, support failures, sudden deliverability collapses. The truth usually sits between the two.
Do email warmup tools still work in 2026?
The consensus among operators has shifted: warmup scores grade engagement inside the tool's own network, not real inbox placement, and Gmail and Outlook increasingly detect warmup pools. Several operators report deliverability improving after turning warmup off. Treat warmup scores as a weak signal and campaign metrics as the truth.

Keep reading