Experiment 002, the list
We fed three email verifiers 100 addresses we made up.
A verifier says an address is good. We wanted to know what that is worth, so we gave three of them addresses that do not exist.
The setup
Hundreds of leads on our HVAC list were held back because our verifier graded them catch-all. Are those bad addresses, or good addresses nobody can confirm? We built a blinded test: 300 addresses in four groups, and the verifiers were not told which was which.
| Group | What it was | n |
|---|---|---|
| 1 | Catch-all domains, addresses we invented | 100 |
| 2 | Catch-all domains, real candidates from our list | 100 |
| 3 | Control: real addresses already verified as deliverable | 50 |
| 4 | Control: invented addresses on normal domains | 50 |
If a verifier can see through a catch-all, it should grade group 2 higher than group 1. That gap is the whole test.
The result
Share of each group graded as a good address:
| Group | Tool 1 | Tool 2 | Tool 3 |
|---|---|---|---|
| 1, catch-all, invented | 0% | 3% | 0% |
| 2, catch-all, real | 0% | 4% | 0% |
| 3, real, verified | 12% | 96% | 84% |
| 4, normal domain, invented | 0% | 74% | 0% |
Real minus fake on catch-all domains: +0, +1 and +0 points. None of the three could tell our invented addresses from real ones, because the receiving server accepts anything. That is what catch-all means. It is not a verdict on the address. It is the server refusing to answer.
Tool 2 also graded 37 of our 50 invented addresses on normal domains as valid. A confident valid on a catch-all domain is not better information. It is a guess with a green label.
The result we did not expect
Tool 1 is a self-hosted prober that asks the recipient's mail server directly. On the control group of real, deliverable addresses it scored only 12%. Split by the recipient's mail provider it resolves completely: 5 of 5 on Google, 0 of 31 on Microsoft. Microsoft rejects probes from a home connection before it ever looks at the address, so the tool returned the same answer for real and fake, and it looked like a clean verdict. Most of our list is on Microsoft. Run through that tool, most of our good leads would have been marked invalid, and it would have looked like a successful cleanup.
What we do now
Catch-all addresses never go into a campaign. If we send to them at all, they go as their own small campaign, where a real send is the only verdict there is. We judge a finder by how many of its addresses bounce, not by how many it calls valid. And a verifier running from a residential connection never gets pointed at Microsoft-hosted domains.
To be clear: an address existing says nothing about whether the email lands in the inbox. This test only asks whether the mailbox is real. The tools are not named because the cause of the false valids is unconfirmed, and because the point is the mechanism, not the vendor.
Try this
Take 20 real addresses from your list and invent 20 on the same domains. Run both sets through your verifier. If the two sets come back looking the same, you now know what its catch-all label is worth.
Want this for your outbound?