Email deliverability is the practical ability of your messages to reach the intended recipient experience, especially the inbox, rather than being rejected, deferred, routed to spam, or otherwise filtered away from normal visibility.
Email deliverability- What is email deliverability? Delivery is not inbox placement
- The 4 gates between send and inbox
- Gate 1
- Gate 2
- Gate 3
- Gate 4
- Why deliverability is a system, not one score
- Deliverability changes can be local even when the campaign was global
- Common failure patterns
- What to monitor continuously
- Worked example
- What Email deliverability requires
- A practical Email deliverability rollout
- What passes and what does not
- Common mistakes
- Comparison
- Questions we get asked
- OnVoard's take
What is email deliverability? Delivery is not inbox placement
It is broader than delivery. In provider event data, a message can be recorded as delivered when the recipient's mail server accepted it. Amazon SES, for example, defines a Delivery event as successful delivery to the recipient's mail server. That event does not tell you whether the receiving system later placed the message in the inbox, spam folder, or another category. So the useful mental model is not "sent versus delivered." It is a chain of separate gates.
The 4 gates between send and inbox
Four gates, four different kinds of evidence
"Delivered" usually means gate 2. Deliverability is about gate 3.
- Send1Sender platform accepts and queuesA successful API call
- Delivery2Recipient server accepts or rejectsSMTP response
- Placement3Mailbox filters the accepted messageInbox, spam, or another tab
- AttentionRecipient sees and actsOr ignores, deletes, or complains
Gate 1: your sending system accepts the message
Your application or ESP accepts the send request. A successful API call proves only that the sending system accepted work. It does not prove that the destination server accepted the message.
Gate 2: the receiving server accepts or rejects it
During SMTP delivery, the destination can accept the message, temporarily defer it, or permanently reject it. A temporary failure may be retried. A permanent failure ends that delivery attempt unless something changes.
Gate 3: the mailbox provider filters the accepted message
After acceptance, a provider can place a message in the inbox, spam, another tab or folder, or apply other internal filtering. The sender usually cannot observe this decision perfectly for each message.
Gate 4: the recipient actually sees and values it
Inbox placement still does not guarantee attention. A recipient may ignore, delete, unsubscribe, or mark the message as spam. Those actions can affect future filtering and reputation.
A healthy program therefore needs to distinguish send acceptance, SMTP delivery, inbox placement, and recipient response instead of collapsing all 4 into one "delivery rate."
Why deliverability is a system, not one score
There is no universal deliverability score shared by Gmail, Yahoo, Outlook, and every other mailbox provider. Each receiver observes its own traffic and applies its own policies, models, reputation signals, and user behavior. A useful diagnostic frame is: provider × identifier × stream × time
- Provider: Gmail can behave differently from Yahoo or Outlook for the same campaign
- Identifier: a problem may attach more strongly to a domain, DKIM identity, return-path domain, IP, URL, or some combination
- Stream: promotional newsletters, password resets, receipts, and support mail can have different audiences and behavior
- Time: reputation and filtering are dynamic. A sudden volume spike or complaint burst can create a short-term problem that looks different from a long-running structural issue
This is why "our deliverability is 98%" is often an incomplete diagnosis. A high server-acceptance rate can coexist with poor Gmail inbox placement, while a small group of invalid addresses can produce a visible bounce-rate problem without affecting all providers equally.
Deliverability changes can be local even when the campaign was global
Suppose the same campaign goes to Gmail, Yahoo, and Outlook recipients. Gmail performance falls sharply while Yahoo and Outlook remain stable.
That pattern should change your first hypothesis. A global template bug or broken sending domain is still possible, but a provider-local signal such as Gmail-specific reputation, recipient behavior, or filtering is now more plausible than a universal failure. The diagnostic unit should therefore include at least:
provider × sender domain × sending IP/pool × stream × time windowThen compare a control cohort that did not degrade. The contrast often teaches more than the absolute metric. If one provider, one IP pool, or one stream diverged while the others stayed healthy, investigate what is unique to that slice before changing the entire program.
This is the same reason “our overall bounce rate is fine” can hide a serious incident. Aggregate metrics can average a damaged cohort together with a much larger healthy one.
Common failure patterns
Narrow the failure before changing anything
- YesInspect status codes, authentication, and rate limitsA rejection or delivery problem
- No, messages are acceptedDoes a healthy control cohort (another provider, IP, or stream) stay normal on the same day?
- YesNarrow to the affected provider, identity, or stream before changing anything else
- No, broad declineCheck domain-wide authentication, content, and volume changes
- Yes
What to monitor continuously
A useful operational dashboard should separate at least:
- send attempts
- receiver acceptance / delivery events
- temporary failures
- permanent failures
- complaint events
- unsubscribe events
- authentication pass and alignment health
- provider-specific compliance signals
- domain and IP changes
- volume by provider and stream
- clicks, conversions, and revenue where relevant
Open rate can remain a directional engagement signal, but it should not be the only engagement input because privacy and automated fetching can distort it.
The most important habit is segmentation. A single blended number across every provider, domain, IP, stream, and day is useful for a board slide and weak for diagnosis.
Worked example
Imagine an ESP reports:
- 100,000 messages submitted
- 99,000 Delivery events
- 500 hard bounces
- 500 still pending or otherwise failed
- A team might conclude that deliverability is 99%
Now suppose Gmail recipients say the newsletter is consistently in spam and Gmail's own diagnostics show poor compliance or complaint signals. There is no contradiction. The ESP's Delivery event means the recipient mail server accepted the message. The provider's filtering decision occurs after or around that acceptance path and is not represented by the same metric.
The right response is to segment the evidence by provider and investigate Gmail-specific authentication, complaint, volume, and audience signals. A 99% receiver-acceptance rate does not answer the inbox-placement question.
What Email deliverability requires
Authentication and identity consistency
SPF, DKIM, and DMARC tell receivers whether sending identities are authorized and aligned. They are prerequisites for trustworthy identity, not automatic tickets to the inbox.
For large senders, provider rules make this operationally mandatory. Gmail and Yahoo both publish authentication requirements for bulk senders.
Complaints
A spam complaint is strong negative recipient feedback. But even the complaint-rate denominator differs by provider. Gmail's Postmaster spam rate focuses on messages delivered to engaged recipients' inboxes and then manually marked as spam. Yahoo says its rate is based on mail delivered to the inbox. That means complaint percentages from different dashboards are not necessarily interchangeable.
Bounces and delivery errors
Permanent failures can indicate invalid destinations or policy rejection. Temporary failures can indicate throttling, mailbox limits, transient server problems, or reputation pressure. The SMTP status and provider message often contain more diagnostic value than the label "hard" or "soft" alone.
Domain and IP reputation
Mailbox providers maintain their own views of sender reputation. Microsoft SNDS, for example, exposes data about individual sending IPs for Outlook.com. Gmail historically exposed domain and IP reputation ratings in Postmaster Tools, but Google now says those legacy reputation dashboards are being retired as it transitions Postmaster Tools to v2. That is a useful reminder not to build an operating model around one dashboard label.
Recipient engagement and list quality
Repeatedly sending mail that recipients do not want creates complaints, inactivity, and list-quality problems. Provider guidance consistently emphasizes permission, easy unsubscribe, and removing invalid recipients.
Engagement metrics themselves need interpretation. Opens can be distorted by Apple Mail Privacy Protection and other automated activity, so they should not be treated as direct proof that a human read a message.
A practical Email deliverability rollout
Seven steps, in order.
- 01A diagnostic sequence that avoids random fixesWhen a merchant says "email is going to spam," do not start by changing the subject line. Narrow the failure first.
- 021. Identify the affected provider and time windowAsk whether the issue is Gmail-only, Yahoo-only, Outlook-only, or broad. Compare the same stream before and after the change.
- 032. Separate rejection from placementIf the receiving server is rejecting or deferring messages, inspect SMTP errors and sending-provider events. If messages are accepted but recipients report spam placement, the investigation shifts toward filtering, reputation, complaints, authentication, and audience quality.
- 043. Identify the affected sender identitiesRecord: - visible From domain; - DKIM signing domain; - SPF / return-path domain; - sending IP or pool; - campaign or stream identifier. A domain problem and an IP problem can look similar at the user level but require different remediation.
- 054. Check provider-native evidenceUse the receiver's own tools where available. Gmail Postmaster data is scoped to personal Gmail traffic. Microsoft SNDS is scoped to Outlook.com IP reputation. Yahoo's Sender Hub and Complaint Feedback Loop provide Yahoo-specific signals. Do not treat any one tool as a universal reputation oracle.
- 065. Compare complaints, bounces, authentication, and volume togetherOne metric rarely explains a deliverability incident by itself. A complaint spike after a sudden volume increase means something different from a quiet domain with clean complaint data but repeated authentication failures.
- 076. Change the causal input, then re-measureIf the problem is a stale segment, stop sending to it. If a provider is throttling a new IP, reduce and stabilize volume. If DMARC alignment is broken, fix the sender identity. If a specific acquisition source creates complaints, isolate that source. Avoid cosmetic changes that do not address the evidence.
What passes and what does not
- Deliverability investigations become much faster when the evidence is matched to the layer of the problem
- A sender can therefore have perfect authentication and poor placement, or 99% SMTP acceptance and a major spam-folder problem. Those are not contradictions; they are measurements of different gates
Common mistakes
- A sender jumps from 10,000 messages a day to 200,000 after importing an old list. Authentication remains perfect, but complaints, invalid addresses, and unfamiliar volume appear together. The solution is not another DKIM key. It is audience and volume control
- On a shared pool, other senders can contribute to IP-level reputation. Gmail explicitly notes that activity from senders on a shared IP can affect other senders. The merchant still owns list quality and domain behavior, but the infrastructure dimension is not fully isolated
- A large marketing campaign can create complaint and throttling pressure around the same identities used for password resets or receipts. Separating streams gives operators better diagnostics and can reduce blast radius when a marketing stream deteriorates
- Gmail warns that its displayed spam rate can look low when Gmail is already routing many messages to spam, because fewer messages reach recipients' inboxes where they can be manually marked as spam. A low complaint rate is therefore not proof that reputation is healthy
Comparison
Evidence has different resolution: do not ask one dashboard to answer every question
| Option | Evidence | What it can answer well | What it cannot prove |
|---|---|---|---|
| SMTP response / bounce | SMTP response / bounce | Whether a receiving server accepted or rejected the delivery attempt and why it reported doing so | Inbox versus spam after acceptance |
| Authentication result | Authentication result | Whether SPF, DKIM, and DMARC identities passed/aligned | Desirability or inbox placement |
| Gmail Postmaster Tools | Gmail Postmaster Tools | Gmail-specific complaint/authentication/delivery signals for eligible traffic | The state of Yahoo, Outlook, or every individual recipient |
| Microsoft SNDS | Microsoft SNDS | Outlook.com-oriented data for specific sending IPs | Domain reputation at every provider |
| Complaint feedback loop | Complaint feedback loop | Which reported complaints the provider exposes | A universal complaint denominator or every negative user action |
| Open/click telemetry | Open/click telemetry | A downstream engagement signal with privacy/bot noise | Reliable proof of inbox placement or human reading |
Questions we get asked
Is email delivery the same as email deliverability?
No. Delivery commonly means the receiving mail server accepted the message. Deliverability is the broader problem of reaching the intended mailbox experience, including avoiding rejection, spam placement, and reputation problems.
Can an ESP know with certainty that every delivered email reached the inbox?
No. A sender can observe delivery responses and provider-specific signals, but mailbox providers control final placement and do not expose a perfect per-message inbox-placement truth to senders.
Does passing SPF, DKIM, and DMARC guarantee the inbox?
No. Authentication establishes authorized identity. RFC 9989 explicitly notes that a DMARC pass does not guarantee that delivery to the inbox is safe or desirable, and providers use additional anti-abuse and reputation signals.
What is the first thing to check when deliverability drops?
Define the scope: provider, sending identity, stream, and time window. Then determine whether the failure is SMTP rejection/deferment or post-acceptance placement. That prevents unrelated fixes from obscuring the real signal.
Should I optimize open rate to improve deliverability?
Treat opens carefully. They can help with directional engagement analysis, but privacy systems and automated fetching can make them an unreliable proxy for human attention. Complaints, bounces, clicks, conversions, consent, and provider-native diagnostics usually provide a stronger combined picture.
OnVoard's take
Deliverability should be operated like incident diagnosis, not like a single marketing KPI. Start with where the message failed in the chain, then slice by provider, identity, stream, and time.
That framing prevents 2 expensive mistakes: calling receiver acceptance "inbox delivery," and responding to a provider-specific reputation problem with generic content tweaks. The job is to find the failing gate and the identity that owns it.
Every app on every plan. Connect your store and switch on the flows in an evening.