Back to all experiments

Experiment 006, the list

The stat we stopped quoting.

For weeks we believed people on Microsoft inboxes do not reply to cold email. We had a number for it. The number was ours, and it was wrong.

The number we quoted

Recipient providerReply rate, old data
Google2.72%
Microsoft0.36%
Gap7.6x

It held across two generations of our infrastructure. It looked like a law of the market: Microsoft filters harder, so write off Microsoft.

What was wrong with it

Two things, and we missed the first for a month. The Outlook senders in that data had been warmed with settings written for Google inboxes. They were sick from day one, and Microsoft weighs sender reputation more heavily than Google does, so a sick sender shows up as a Microsoft problem. The lists were also unverified. “Microsoft recipients do not reply” and “our senders were unhealthy and Microsoft noticed” produce the identical number, and nothing in that data could tell them apart.

A second dataset seemed to confirm the gap: 1.74% against 0.43%. We withdrew that one too. It counted out of office messages as replies.

The clean read

Healthy Google senders, a verified list, every reply read by hand. Our roofing campaign:

Recipient providerSampleHuman reply rate
Microsofta few hundred recipients1.27%
Googlea few hundred recipients1.46%
GapNone we can see

Four of the first six replies in that campaign came back from Microsoft mailboxes, one of them written by a person who had read the email. That is inbox delivery to Microsoft from a Google sender, measured, not assumed.

What we do now

We never quote the 7.6x again, except as a retracted number. Our lists stay majority Microsoft, because that is where our buyers are. And when a per-provider number looks bad, we check the health of our own senders before we blame the recipient's provider. A measurement taken on sick senders is a property of the senders, not of the people they wrote to.

To be clear: the old number is un-evidenced, not disproved. We are not saying Microsoft is fine everywhere. We are saying no penalty shows up in any clean data we hold, and the data we built the claim on could never have shown one either way.

Try this

Before you compare reply rates by recipient provider, split your senders by health. If the gap only appears on the senders that were already struggling, you have measured your own infrastructure, not your audience.

Want this for your outbound?

We test on our own sends first, then run it for you.

Book a 30 minute call