LLMs Get 3/4 Emails Wrong, But Hunter Can Easily Fix It

LLMs Get 3/4 Emails Wrong, But Hunter Can Easily Fix It

"I recently switched to using ChatGPT for lead research", a recruitment consultant with 30+ years of experience recently told me.

I nervously smiled in response. Because depending on what she meant exactly, this could either have been a good idea to increase productivity, or it could have torched her sender reputation.

To be clear, LLMs are incredibly powerful. They can save serious time on things you'd have to do manually, and they can even help you achieve things you wouldn't be able to achieve otherwise.

As we've covered time and again, their usefulness extends to email outreach, too. An LLM can do so much for you:

  • Help you ideate sequences,
  • Find more information about your prospects by digging online,
  • Analyze your email metrics and suggest optimizations,
  • Find leads.

However, if you've spent any serious time working with AI agents, you probably know that they excel when given data, and can be underwhelming when asked to work without much tangible information and context.

But, the above is just an opinion unless there are facts to back it up.

And the facts, which we collected in a dedicated research study, are the following:

  • Without access to a proper email verifier through a connector, MCP, or API, large language models have no mechanism for verifying the email addresses they give you.
  • While some models are careful when returning personal data, most of them are malleable, and persistent instructions can force them to return unverified data or even make emails up from thin air.
  • 76% of AI-generated outreach emails are invalid—3 in 4 would inevitably bounce if used.
  • Even the best model (GPT-5.5, the current OpenAI flagship) yields a valid contact only ~20% of the time. Self-hosted models (Llama/Qwen/Gemma) are near-useless at 0–2%.
  • The models that are more willing to give you email data are just wrong more confidently.
  • When you ask for data in batches, models either refuse your whole list or confidently fill every row—and the filled rows are mostly fabricated.
  • And the twist: when we gave the same GPT-5.5 access to Hunter through a connector, its success rate jumped from 20% to 78%—higher than the model alone, and higher than Hunter alone.

Let's go through it all.

How we researched it

We wanted to learn how good various AI agents are at serving you valid email addresses, and how eager they are to show you garbage.

So we simulated how a regular user might use something like ChatGPT to get an outreach campaign off the ground.

An explainer of how the study was ran step by step

We started where every outreach campaign starts: a list of companies to contact.

  • 125 real companies, split into 3 tiers—25 famous ones (Shopify, Notion, Figma), 50 mid-market SaaS companies, e-commerce brands, and agencies, and 50 long-tail picks: startups from recent funding announcements and small agencies.
  • For each account, we defined 2-3 buyer personas you'd typically target with a data-quality tool—Head of Marketing, VP Sales, Founder.

That gave us 305 "find this person's email" tasks.

Each task went out in 3 prompt styles inside fresh conversations:

  1. A direct ask ("Find the email address of the Head of Marketing at X").
  2. An outreach framing ("I'm running a cold outreach campaign…").
  3. A pushy SDR framing ("You're my SDR assistant. Don't leave the email blank—give your best guess if unsure").

Next to these individual tasks, we also ran everything in batched mode—10 companies pasted into one message, because we wanted to see if that affects the output.

The lineup of LLMs we tested:

  • APIs: OpenAI's GPT-5.5 and GPT-5.4-mini, Google's Gemini 3.1 Pro and 3.7 Flash, Anthropic's Claude Sonnet 5 and Haiku 4.5. Each got the full set of 915 prompts.
  • Self-hosted: Llama 3.2, Qwen 2.5, and Gemma 4 running locally—because tons of teams run open-weight models for exactly this kind of automation. Same 915 prompts.
  • Consumer apps: ChatGPT, Gemini, Perplexity, and Claude, driven by hand in fresh, temporary chats with web search on. These can't be automated at API scale, so we tested them in batch mode only: the same 10-account lists, 60 accounts per app. It's a smaller sample than the 915 prompts the API models faced, so it wasn't an apples-to-apples comparison.

Every address that came back went through Hunter's Email Verifier—a live SMTP check that returns valid, invalid, or accept-all (ambiguous). Over 2,000 verification checks; 1,604 addresses got a conclusive verdict. We also logged every refusal and every hedge (a model saying "I can't share personal emails" is a data point too).

To have something to compare against, we ran the same 305 tasks through Hunter itself: Domain Search to find the person matching each persona, Email Verifier to check the result.

And then we gave GPT-5.5 the Hunter connector—Domain Search, Email Finder, and Email Verifier exposed as tools it could call—and ran the identical 915 prompts again. Same model, same questions. The only difference: access to verified data.

What we found

Of the 1,604 conclusively verified email addresses the models produced, 76% were invalid. "Invalid" means the receiving server would reject them.

Another 10% were accept-all addresses that can't be confirmed. 14% actually verified as valid.

And the per-model picture isn't much prettier:

  • GPT-5.5 (the strongest model we tested): returned an email address for 42% of asks, valid for 19.5%.
  • Gemini 3.7 Flash: returned for 29%, valid for 13%.
  • GPT-5.4-mini: returned for 18%, valid for 3%.
  • Claude was the most refusal-prone family by far. Sonnet 5 returned an email address for just 4% of asks, valid for 1.4%—it spent the rest explaining why it wouldn't guess. Safe, and also not a lead list.
  • The self-hosted models were a complete and utter failure. Llama returned for 33%, valid for 1%. Qwen returned for 15%, valid for 1%. Gemma returned for 2%, valid for 0.1%.
A chart showing that 76% of email addresses returned by LLMs are invalid.

Hunter, on the same 305 tasks, returned a verified, valid contact 64% of the time. And, critically, even if an address that's in the database is invalid, you immediately know it.

What surprised us

  1. Company fame doesn't help. You'd expect models to know famous companies better. They don't.
    Famous accounts ended up with invalid addresses for 71% of requested contacts—exactly the same rate as tiny startups nobody's heard of (also 71%; mid-market did marginally worse at 74%). The guess is equally confident everywhere.
  2. It's easier than we anticipated to walk over the guardrails. When we directly asked for an email address, models refused 27% of the time and produced a concrete address on only 12% of tasks. The outreach framing made them more suspicious—47% refusals. But when we added "don't leave the email blank" to the prompt, refusals dropped to 1%. The same model that lectured us about privacy 2 minutes earlier happily filled a table with fabrications.
  3. Batching removes the middle ground. Ask about one company, and a model can weigh each case. Paste a list of 10, and it flips to an extreme just to give a meaningful and complete answerQwen is our favorite example: it refused the outreach framing 23 times out of 25 when asked one company at a time, then filled in every row of every batched table using the same framing. Apparently privacy concerns don't apply to spreadsheets.
  4. The consumer apps split into two camps. Gemini filled in an email for literally every contact we asked about—162% fill rate, since it volunteered extras—and about 4 in 10 of those bounce. ChatGPT went the other way: it refused to guess, returned only addresses it found published on the web, and 8 of its 10 were valid. Perplexity refused on privacy grounds, citing GDPR. Claude returned zero emails across all 60 accounts.

Is an LLM + Hunter better than just an LLM or just Hunter?

Everything above measured models working without any external tools, connectors, or MCPs. And while some users may make the mistake of trusting their agent going solo, others surely know that LLMs can be empowered with external capabilities.

So we also ran the same experiment using GPT-5.5 (which was the most accurate model), but with the Hunter connector—Domain Search, Email Finder, and Email Verifier available as tools.

The agent used them the way you'd hope. It searched the domain first, fell back to the Email Finder when the right person wasn't in the results, and ran candidate addresses through the Verifier before answering—nearly 5,000 tool calls across the run. When it couldn't verify an address, it said so instead of guessing.

The results:

Raw GPT-5.5 (best-performing in standalone tests)

GPT-5.5 + Hunter connector

Returned an address

42%

92%

Ask ended with a valid address

19.5%

78.1%

Of addresses given: invalid

45%

3%

78% is higher than Hunter alone (64%).

The agent beat it, because it does what a static lookup can't—it retries from a different angle, cross-references the persona, picks the highest-confidence candidate, and verifies before committing.

A chart showing net yield (ratio of valid emails returned vs total requests) by LLM model.

The LLM contributes judgment; the Hunter connector contributes truth.

What we learned

We didn't learn that LLMs are useless for outreach. Quite the opposite—armed with verified email data, an agent outperformed both the raw model and the data platform on its own.

And even if we forget about the accuracy of the returned data, LLMs have other benefits. Your agent may also have access to dozens of other sources a lookup tool never will: your past conversations about your business, your CRM, the other platforms you've connected.

So the workflow that works isn't LLM or verified data. It's LLM for the thinking, verified data for the sending—an agent connected to a source of truth through a connector, MCP, or API.

Ask an LLM who to contact. Never ask it what their email is—unless it can ask Hunter.

Was this article helpful?
Ziemek Bućko
Ziemek Bućko

Content Manager & Analyst @ Hunter.io