Cold email open rate benchmarks for 2026 (and why opens lie)
No mailbox provider or standards body publishes one. Here is what the circulating numbers are made of, and the arithmetic that makes them incomparable.
There is no authoritative benchmark for cold email open rates, because no mailbox provider or standards body publishes one and none of them could — an open is recorded by the sender's tracking pixel, not by the receiver. Every figure in circulation is a sending platform's aggregate over its own customers. And even a correctly measured open rate is not comparable between two lists, because the number is dominated by the recipients' mail-client mix rather than by anything you wrote. Here is what the number is actually made of, and what to report instead.
Where a published cold email open rate benchmark actually comes from
Trace any open-rate benchmark back and it ends in the same place: a sending platform ran a query over the campaigns its own customers sent, and published the aggregate.
That is not dishonest, and the resulting number is not useless internally. But it fails every test you would apply to a primary source. The population is self-selected — it is whoever bought that product, sending to whatever lists they happened to build. The methodology is usually undisclosed: whether the denominator was sends or deliveries, whether repeat opens were deduplicated, whether machine opens were filtered and by what rule, whether a sequence's opens were counted per message or per contact. And no two vendors' figures are computed the same way, which is why the ranges in circulation differ by more than the effect any copy change would produce.
This site publishes a benchmark table too, on its cold email benchmarks page, and the honest description of every row on it is "a general range that people report", not "a measurement". Treating it as anything more is the error this post is about.
Read the mailbox providers' own documentation and the absence is conspicuous. Google publishes a spam-complaint threshold as a hard number. Yahoo publishes the same. Microsoft publishes authentication requirements. None of them publishes anything about opens, and the reason is structural: the receiver does not record an open. Your server does, when something fetches an image.
September 2021 split every open-rate dataset in two
Apple shipped Mail Privacy Protection with iOS 15 and iPadOS 15 in September 2021. Apple's own description of the behaviour is that Mail downloads remote content in the background by default, regardless of whether you engage with the email, and that it hides the recipient's IP address so senders cannot use it as an identifier.
Read that literally. For every recipient reading in Apple Mail with the feature on, your tracking pixel is fetched on delivery. Not on read. On delivery. Whether the message was opened, archived unread, or never looked at again.
The consequence for benchmarks is not that the number went up. It is that the number stopped measuring the same quantity. A dataset spanning 2020 to 2023 is averaging pre-MPP opens and post-MPP opens together, which are two different things sharing a name. Any figure derived from such a dataset is a weighted average of two incomparable populations, and its value depends on where in that window the campaigns fell.
This is why "the open rate benchmark used to be 20% and is now 40%" is not evidence that email got easier. It is evidence that a measurement changed definition mid-series.
A reported open rate is four populations added together
Decompose what your dashboard shows and it resolves into four groups, each contributing in a different direction.
Prefetchers that always fire. Apple Mail with MPP, and corporate link-and-attachment scanners — Microsoft Defender Safe Links, Proofpoint, Mimecast, Barracuda and similar — fetch remote resources on delivery to scan them. Every one of these registers an open regardless of human behaviour. In B2B outbound the scanner share is substantial and is systematically under-discussed.
Readers whose client loads images. Gmail's web client proxies and caches images through Google's own servers, so the first load registers and subsequent reads of the same cached message frequently do not. This group contributes real signal, undercounted.
Readers whose client blocks images. Outlook desktop blocks external images by default for senders not on the recipient's safe list — which describes every cold email ever sent. A genuine, attentive read from this group registers nothing at all.
Non-readers who register nothing. The people who deleted it unread in a client that does not prefetch.
Your reported open rate is groups one and two, plus a fraction of group three, divided by everyone. It is not "the proportion of people who read your email" and no correction factor turns it into that, because the size of each group is a property of the list rather than a constant you can divide out.
Two tells give away a scanner-heavy list. A click rate that approaches or exceeds the open rate, which is impossible for humans and routine for gateways that follow every link. And a dense cluster of opens registered within a second or two of delivery.
Two identical campaigns, 61% and 27%
The following is an illustrative worked example with assumed client mixes, not a measurement. Its point is the size of the effect, not the specific figures.
Take two campaigns, each delivered to 1,000 contacts, with identical copy, identical subject line and identical folder placement. Assume in both that exactly 30% of recipients genuinely read the message.
Campaign A — a founder list. 450 recipients on Apple Mail with MPP, all registering an open on delivery. 50 behind a prefetching security gateway, all registering. 350 on Gmail webmail, of whom 30% read and load images, giving 105. 150 on Outlook desktop with remote images blocked, giving 0.
Recorded opens: 450 + 50 + 105 = 605, or 61%.
Campaign B — an enterprise IT list. 80 on Apple Mail with MPP. 150 behind a prefetching gateway. 120 on Gmail webmail, 30% of whom load images, giving 36. 650 on Outlook desktop, giving 0.
Recorded opens: 80 + 150 + 36 = 266, or 27%.
Identical human behaviour. Identical placement. A reported spread of 34 percentage points, produced entirely by which mail clients the two audiences happen to use. Campaign A comfortably clears any "good" band in circulation; Campaign B falls below every "typical" one. Neither number tells you anything about the email.
That is the whole argument against cross-campaign open-rate benchmarking, and it does not depend on the exact percentages being right. Vary the mixes however you like; the conclusion survives, because the dominant term is a property of the audience.
Autocloz's free plan covers 5 users and 10 mailboxes with warmup, authentication monitoring and reply classification included — start free if you would rather instrument replies than argue about opens.
The denominator moves independently of the numerator
Even with an honest numerator, the divisor is doing damage, and it is doing it in three different places.
Sent versus delivered. Bounces belong in the bounce rate, not in the denominator of an engagement metric. A campaign with 8% hard bounces reports a materially lower open rate against sends than against deliveries, for a reason that has nothing to do with engagement.
Accepted versus delivered. A 250 reply at the end of an SMTP transaction means the receiving server took responsibility for the message. It does not mean anyone saw it. Dashboards that label acceptance as "delivered" are technically defensible and practically misleading, and a campaign filed entirely into spam can report a 99% delivery rate.
Per message versus per contact. A four-step sequence to 2,000 contacts dispatches far more than 2,000 messages. A per-message open rate and a sequence-level open rate for the same run differ by a large factor, and both are legitimate — but they are different questions, and half of all benchmark disagreements are two people comparing one against the other. Report both, label which is which, and never compare across the boundary.
What a defensible internal benchmark looks like
The number you can actually use is a comparison against yourself, under conditions you controlled. Four constraints make it defensible.
One mailbox, or one clearly-labelled group. Reputation, placement and throttling differ per mailbox, so pooling mailboxes averages away the thing you are trying to see.
One segment, with a stable client mix. This is the constraint that matters most, and it follows directly from the decomposition above. Comparing this month's enterprise segment against last month's startup segment measures the segments.
A denominator you wrote down. Delivered contacts, not sends; per contact or per message, stated explicitly.
A sample big enough for the difference you are claiming. A four-point movement on 200 sends is noise. The general problem — how many sends a comparison needs before a difference means anything — is worked through for reply rate in what a good cold email reply rate actually is, and the arithmetic is the same for any rate.
Under those four constraints, week-over-week movement in your own open rate carries some signal, because the measurement error is roughly stable when the population is stable. That is the strongest honest claim available for this metric.
The two questions an open rate can still answer
There are exactly two, and both are diagnostic rather than performance-related.
Did something break? A mailbox that reported around 40% for six weeks and drops to 8% across every campaign at once has had something change at the receiving end — an authentication failure, a reputation event, a blocklist. The absolute value is meaningless; the discontinuity is not. This is a genuinely useful alert and it costs nothing to configure.
Is this segment behaving differently from that one, within the same mailbox and the same week? Directional only, and only if you accept that part of the difference is client mix rather than interest.
What it cannot do is settle a subject-line test. If a large fraction of your recipients register an open on delivery regardless of the subject, the test is partly measuring a constant. Run subject-line tests on replies instead, and read how to structure a cold email test that produces a trustworthy answer before committing volume to one.
The four metrics side by side, by what actually records them
The reason open rate is the weakest of the four is visible as soon as you ask who writes the record and what can corrupt it.
Open rate. Recorded by your own server, when something fetches a one-pixel image. Corrupted by prefetchers that fire without a human, by clients that block the image, and by proxy caching that suppresses repeat loads. Nobody at the receiving end participates in the measurement, so there is no second source to check it against.
Click rate. Recorded by your own server, when something requests a redirect URL. Corrupted by the same gateways, more severely — a security scanner that follows every link produces a click for every link in the message. It shares open rate's disease with a rewriting layer on top.
Reply rate. Recorded by a message arriving in your mailbox, which is a real artefact with headers you can read. Corrupted by auto-responders and bounce notifications being counted as replies, which is a classification problem rather than a measurement problem — and classification problems can be fixed, because you can look at the message.
Inbox placement. Recorded by reading a mailbox you control and finding, or not finding, a token you put in the subject. Corrupted by the seed mailbox having no engagement history of its own, so it is a sample of how the domain is treated rather than of how your recipient is treated. Its limitation is a known bias, not an unknown one.
The pattern is that the two metrics you can audit are the two that involve a message arriving somewhere. The two you cannot audit are the two that involve an HTTP request from a machine you have never seen.
The one open-rate figure worth writing down
If you keep exactly one open-rate number, make it a per-mailbox baseline used only to detect a break.
Take a rolling fourteen-day median of the open rate for each sending mailbox, computed on that mailbox alone, across whatever campaigns it ran. The median rather than the mean, because a single small campaign to an unusual segment should not move it. Fourteen days rather than seven, because a Monday-to-Friday sending window makes a seven-day window sensitive to which days had volume.
Then alert on a *relative* drop rather than an absolute threshold. A mailbox whose median sits at 55% and falls to 20% has had something happen. A mailbox whose median has always sat at 22% and stays there has not, and an absolute alert at 30% would page you about the second and stay silent about the first. The absolute value is a property of the mailbox's audience; only the discontinuity is a property of the mailbox.
Two more rules make the alert worth having. Require a minimum volume in the window before it can fire, or a quiet fortnight will trigger it. And check placement before you check copy: if a seed test says probes are landing in spam, the subject line is not the variable, and rewriting it will cost a week and change nothing.
What to put on the report instead
Four numbers, in the order they constrain each other.
Measured inbox placement. Not a delivery percentage. A seed test mails monitored mailboxes across providers and records where each probe landed, which is an observation rather than an inference. Be precise about its limits: Autocloz classifies a probe as inbox, spam or missing, and that is the entire vocabulary — there is no Promotions verdict, and there cannot be one, because probes are read over IMAP and Gmail's category tabs are not IMAP folders. The tests run when an operator triggers them; nothing schedules them for you.
Reply rate on a clean denominator. Human replies only, divided by contacts delivered to. Out-of-office responses, auto-acknowledgements and bounce notifications routinely add a point or more to an uncorrected figure. Autocloz classifies every inbound reply by intent — positive, negative, out-of-office, bounce or unsubscribe — across all five channels in one unified reply queue, which is what makes stripping the non-human ones a filter rather than a triage job.
Positive-reply share. The fraction of replies expressing interest. A healthy reply rate with a poor positive share means you are reaching people who will answer and never buy, and that is a targeting problem rather than a copy problem.
Meetings from contacts. The product of everything above, and the only one anybody outside the team cares about.
The benchmark calculator turns a funnel's raw counts into those rates with qualitative bands attached, which is the right shape for this data — a band you can argue with rather than a single number that implies more precision than exists.
What no open rate can tell you, and what Autocloz does not do
An open rate cannot tell you whether a person read your message, whether they read it more than once, whether they forwarded it, or whether they were interested. It cannot be compared with another organisation's open rate, another tool's open rate, or your own from a different segment. And it cannot be corrected, because the error term is a property of the audience rather than a constant.
Carrying a tracking pixel also has a cost worth counting on the other side of the ledger. A message from an unknown sender whose only remote resource is a one-pixel image on a shared tracking host is a recognisable pattern, and a shared tracking domain carries the history of everyone else using it. Turning open tracking off, or at minimum moving it to a subdomain of your own authenticated sending domain, is a defensible trade — you are otherwise paying real reputation for a number you have just been told not to trust. How a tracking pixel records a request rather than a reader sets out the mechanics in full.
Autocloz does not publish a proprietary open-rate benchmark and does not have a dataset that would justify one. It measures placement by probe rather than by inference, classifies replies by intent, and reports per mailbox and per sequence rather than pooling them. It cannot see inside any receiver's classifier, cannot tell a human open from a scanner's fetch with certainty, and cannot make an open rate comparable across audiences — nothing can. If you are choosing between a general-purpose CRM that reports this way and a dedicated sequencing tool, the comparison with Lemlist is the useful version of that decision.
Frequently asked
Is there an authoritative benchmark for cold email open rates?
No standards body or mailbox provider publishes one. Google, Yahoo and Microsoft publish authentication requirements and a spam-complaint threshold; none of them publishes an open-rate figure, and none of them could, since an open is recorded by the sender's own tracking pixel rather than by the receiver. Every number in circulation comes from a sending platform aggregating its own customers' campaigns, with the population and the methodology usually undisclosed.
Why do two identical campaigns report completely different open rates?
Because a recorded open depends on the recipient's mail client, not on the message. Apple Mail with Mail Privacy Protection fetches remote content on delivery whether or not anyone reads it, corporate security gateways fetch it to scan it, Outlook desktop blocks remote images by default so genuine reads register nothing, and Gmail proxies and caches images so repeat reads undercount. Two lists with different client mixes report different open rates for identical copy landing in identical folders.
What happened in September 2021 that changed open-rate data?
Apple shipped Mail Privacy Protection with iOS 15 and iPadOS 15. Apple's own description is that Mail downloads remote content in the background by default, regardless of whether you engage with the email, and that it hides the recipient's IP address. Any benchmark dataset that spans that release is averaging two structurally different measurements, which is why open-rate figures from before and after it are not comparable even when they come from the same source.
Can I still use open rate for A/B testing subject lines?
Not reliably. If a large share of your recipients register an open on delivery regardless of what the subject line said, a subject-line test measured on opens is partly measuring a constant. Test subject lines on reply rate instead, accept that this needs more volume per variant to reach a difference you can trust, and handle the mechanical checks such as length and preview-text truncation with a static check rather than a live test.
What should replace open rate as a top-line metric?
Measured inbox placement, and a reply rate with a clean denominator. Placement is observed by mailing monitored mailboxes and recording which folder each probe reached, so it is an observation rather than an inference from a pixel. Reply rate is computed from replies you can classify as human, divided by contacts you actually delivered to. Both can be wrong, but both can be audited, and an open rate cannot.