Back to all experiments

Experiment 002, the list

We fed three email verifiers 100 addresses we made up.

A verifier says an address is good. We wanted to know what that is worth, so we gave three of them addresses that do not exist.

The setup

Hundreds of leads on our HVAC list were held back because our verifier graded them catch-all. Are those bad addresses, or good addresses nobody can confirm? We built a blinded test: 300 addresses in four groups, and the verifiers were not told which was which.

GroupWhat it wasn
1Catch-all domains, addresses we invented100
2Catch-all domains, real candidates from our list100
3Control: real addresses already verified as deliverable50
4Control: invented addresses on normal domains50

If a verifier can see through a catch-all, it should grade group 2 higher than group 1. That gap is the whole test.

The result

Share of each group graded as a good address:

GroupTool 1Tool 2Tool 3
1, catch-all, invented0%3%0%
2, catch-all, real0%4%0%
3, real, verified12%96%84%
4, normal domain, invented0%74%0%

Real minus fake on catch-all domains: +0, +1 and +0 points. None of the three could tell our invented addresses from real ones, because the receiving server accepts anything. That is what catch-all means. It is not a verdict on the address. It is the server refusing to answer.

Tool 2 also graded 37 of our 50 invented addresses on normal domains as valid. A confident valid on a catch-all domain is not better information. It is a guess with a green label.

The result we did not expect

Tool 1 is a self-hosted prober that asks the recipient's mail server directly. On the control group of real, deliverable addresses it scored only 12%. Split by the recipient's mail provider it resolves completely: 5 of 5 on Google, 0 of 31 on Microsoft. Microsoft rejects probes from a home connection before it ever looks at the address, so the tool returned the same answer for real and fake, and it looked like a clean verdict. Most of our list is on Microsoft. Run through that tool, most of our good leads would have been marked invalid, and it would have looked like a successful cleanup.

What we do now

Catch-all addresses never go into a campaign. If we send to them at all, they go as their own small campaign, where a real send is the only verdict there is. We judge a finder by how many of its addresses bounce, not by how many it calls valid. And a verifier running from a residential connection never gets pointed at Microsoft-hosted domains.

To be clear: an address existing says nothing about whether the email lands in the inbox. This test only asks whether the mailbox is real. The tools are not named because the cause of the false valids is unconfirmed, and because the point is the mechanism, not the vendor.

Try this

Take 20 real addresses from your list and invent 20 on the same domains. Run both sets through your verifier. If the two sets come back looking the same, you now know what its catch-all label is worth.

Want this for your outbound?

We test on our own sends first, then run it for you.

Book a 30 minute call