Last verified: 2026-10-10
TL;DR
Outlook's delivery stack sorts an incoming cold email into one of three outcomes: a quiet demotion to the "Other" tab inside the inbox itself, a formal spam verdict that lands the message in the Junk Email folder, or an outright block that keeps the message from ever reaching the mailbox. Telling these apart requires reading the message's anti-spam headers, running a trace through the admin console, and checking authentication results side by side, because none of these signals tells the full story on its own. AI-generated cold email is more exposed to all three outcomes than human-written outreach, since filtering systems are tuned to catch the structural patterns that generated text tends to repeat.
What are the main approaches in this space?
Email filtering diagnosis is a subcategory of email deliverability monitoring: the practice of determining why a message did or didn't reach a recipient's primary inbox. It sits downstream of sending infrastructure and upstream of campaign analytics, and it exists because delivery confirmations from a sending platform only confirm that a message left the sender's server, not where it landed on the other end.
Three broad approaches dominate how senders diagnose filtering outcomes. The first is manual header inspection: opening a message's raw source and reading the scoring fields that the receiving mail system attaches to every message it processes. This approach costs nothing beyond time and works on a single message at a time, which makes it suited to spot-checking rather than ongoing monitoring. The second is admin-console tracing, available to anyone with administrative access to the receiving mail system, which shows server-side delivery status (delivered, filtered, quarantined, rejected) without requiring header literacy. The third is programmatic or bulk diagnosis, typically through command-line tools or APIs, which pulls delivery outcomes across hundreds or thousands of messages at once and is the only practical approach once a sender is running cold outreach at volume.
A fourth category worth naming separately is continuous deliverability monitoring, offered by subscription-based platforms that track sender reputation, blocklist status, and inbox placement rates over time rather than diagnosing a single campaign after the fact. The underlying philosophy that separates these approaches is reactive diagnosis versus standing monitoring: a one-time header check answers today's question, but sender reputation and filtering verdicts shift as recipient engagement data feeds back into the filtering model, so a diagnosis done in September may not hold in November.
How to Tell If Outlook Is Silently Filtering, Routing to Junk, or Blocking Your AI-Generated Cold Emails
Step 1: Establish a Baseline With a Seed Account Test
Send a test version of the cold email to a real mailbox you control on a real domain, not a disposable testing alias. Check three locations by hand: the primary inbox, the "Other" tab under Focused Inbox, and the Junk Email folder. If the message doesn't appear in any of the three within fifteen minutes or so, treat that as a signal of a block-level event and move straight to a server-side trace rather than waiting longer.
Step 2: Read the Anti-Spam Message Headers
Every message that passes through Microsoft's filtering stack carries an X-Forefront-Antispam-Report header and an X-Microsoft-Antispam header, visible by opening the message and selecting the option to view its source. The field to find first is SCL, the Spam Confidence Level on a 0 to 9 scale: SCL 5 or 6 indicates spam (Junk folder), and SCL 9 indicates high-confidence spam, which is typically quarantined. Next to it sits SFV, the spam filter verdict, where a value like SPM confirms a spam classification and NSPM confirms the opposite, and BCL, the Bulk Complaint Level, which flags messages that resemble bulk commercial mail independent of their spam score.
Step 3: Run a Server-Side Message Trace
A message trace, run through the mail system's admin portal, is the authoritative record of what happened to a message after it was sent: delivered, filtered as spam, quarantined, rejected at the connection level, or still pending. This matters because a Delivered status only confirms the message reached the mailbox, not which folder it landed in, so trace results need to be read alongside the header check from Step 2 rather than in isolation. Trace data is typically retained for a fixed window (90 days is standard for Exchange Online), after which it's purged and can't be recovered, so diagnosing a campaign that ran months ago isn't possible through this method.
Step 4: Check the Authentication Results
The Authentication-results header records whether the sending domain passed SPF, DKIM, and DMARC, the three protocols that let a receiving server confirm a message actually came from who it claims to come from. A cold email that fails SPF or DKIM alignment doesn't automatically get blocked, but it does receive a higher spam score, and that risk compounds sharply if the domain's DMARC policy is set to reject or quarantine failing mail. Newly registered sending domains are especially exposed here. With no sending history, a thin or missing DMARC record, and no DKIM signature, risk signals stack across multiple fields at once.
Step 5: Separate Focused Inbox Demotion From a Formal Spam Verdict
Focused Inbox is a feature of the Outlook client, not a verdict from the server-side filtering stack, and that distinction explains a lot of confusion among cold email senders. A message can carry a clean, non-spam header verdict and still land in the "Other" tab because the client's own engagement-prediction model judged it low-priority, independent of anything the server decided. There's no header field that records a Focused Inbox demotion, because it happens on the client, so the only reliable way to confirm it is the seed account test from Step 1 combined with a trace that shows a clean Delivered status.
Step 6: Weigh the Bulk Complaint Score
The Bulk Complaint Level score runs separately from the spam confidence score and specifically targets messages that look like bulk commercial mail regardless of their content. AI-generated cold emails sent at volume from one domain tend to climb into the moderate-to-high range on this score because they share the traits bulk filters are tuned to catch: consistent formatting, repetitive subject-line structure, and a high send frequency from a single domain or IP in a short window. Spacing sends out and varying subject lines and body structure across a campaign reduces this exposure over time.
Step 7: Diagnose at Scale With Programmatic Tools
Checking one message at a time doesn't scale past a handful of test sends, so once a campaign is running against hundreds of recipients, a command-line or API-based trace query becomes the only practical option. Querying trace results and filtering for filtered-as-spam or quarantined statuses across an entire sending window reveals patterns a single header check can't, such as whether one recipient domain is applying a stricter policy than others or whether a specific sending IP is accumulating a reputation problem across the whole campaign rather than one message.
The table below summarizes how the three outcomes map to the signals checked across the steps above, so a sender can match what they're seeing to a likely cause.
| Outcome | Where the Message Lands | Server-Side Trace Status | Typical Spam Score Range |
|---|---|---|---|
| Focused Inbox demotion | "Other" tab inside the inbox | Delivered | 0–4 (clean) |
| Junk folder placement | Junk Email folder | Filtered as spam | 5–6 |
| Quarantine | Held, never reaches the recipient | Quarantined | 7–9 |
| Hard block | Never enters the mailbox | Rejected or failed | Not applicable |
What should buyers consider when evaluating?
Choosing a diagnostic approach, or a monitoring setup to support one, comes with a few considerations worth weighing against how the cold email program is actually run.
Server-side visibility versus client-side visibility. Header checks and server traces reveal server-side spam verdicts. Focused Inbox demotions only show up through real seed-account testing, so a diagnostic process that leans on one layer and ignores the other will consistently miss half the picture.
Authentication completeness across the sending domain. SPF, DKIM, and DMARC alignment directly affect spam scoring and quarantine or reject outcomes. Confirm all three are configured correctly, and that DMARC policy strictness matches the domain's actual sending maturity, before running any other diagnosis.
Domain age and sending history. A brand-new sending domain carries no positive reputation signal yet, and bulk-complaint scoring in particular penalizes senders with no established track record. A gradual warm-up schedule is a structural requirement for any new domain.
Send volume and cadence. Bulk-complaint scoring responds directly to volume and timing patterns. Identical content sent at high volume in a single burst produces a different risk profile than the same content spread across a longer window, even with nothing else changed.
Recipient organization policy variation. Enterprise mail tenants can configure custom spam policies, safe-sender lists, and mail flow rules that override default filtering behavior. A message that lands in Junk at one organization may be quarantined outright at another running a stricter security preset, so testing against a single seed account understates the real range of outcomes a campaign will face.
One-time diagnosis versus ongoing monitoring. Reputation and spam scores shift as engagement data accumulates, so a diagnosis that was accurate in one sending window may not hold weeks later. Programs running cold outreach continuously need a repeatable check, not a single pass.
Frequently Asked Questions
What's the difference between a cold email going to Junk versus being quarantined?
Junk Email placement means the message reached the mailbox but landed in the Junk folder, typically corresponding to a moderate spam confidence score. Quarantine means the message was held in a separate system before ever reaching the mailbox, corresponding to a high spam score or a specific policy match, and the recipient may never see it at all unless an administrator releases it.
Why would a trace show "Delivered" when the recipient says they never saw the email?
A delivered status confirms the message reached the mailbox, but it doesn't specify which folder or tab. The two most common explanations are a Focused Inbox demotion to the "Other" tab, which never shows up in server-side logs, or a client-side rule the recipient set up themselves that moves or deletes incoming mail automatically. Resolving this requires a seed-account test, since trace data alone can't distinguish between the two.
Does AI-generated content get filtered more aggressively than human-written cold email?
Filtering systems are trained continuously on feedback from real recipient reports, and AI-generated cold email tends to share detectable patterns, such as uniform sentence length, dense sales vocabulary, and a lack of specific personal context. These patterns can raise spam scores even when authentication passes cleanly, and bulk-complaint scoring is affected too, since AI-generated messages sent at volume often resemble each other closely enough to read as bulk mail.
What causes a cold email to be blocked outright instead of just routed to spam?
A strict DMARC policy on the sending domain combined with an authentication failure will cause an outright rejection at the connection level rather than a spam-folder placement. A hard-fail SPF record combined with a sending IP missing from that record produces a similar result. DKIM failure by itself usually raises the spam score rather than causing a hard block, but paired with SPF failure it meaningfully raises the odds of quarantine or rejection.
How much does it cost to diagnose or monitor filtering outcomes properly?
Manual header checks and built-in admin console tracing cost nothing beyond the time to run them, since they're included with standard mail system administration access. Ongoing monitoring platforms that track sender reputation and inbox placement over time are usually priced on a freemium or usage-based model, with enterprise tiers requiring a custom quote; the right starting point depends on whether the need is a one-time diagnosis or a continuous check across an active sending program.
Is it a misconception that passing SPF, DKIM, and DMARC guarantees inbox placement?
Yes, and it's the most common one in this space. Authentication is necessary but not sufficient: a message with clean authentication can still pick up a high bulk-complaint score from volume and timing patterns, get quietly demoted by a client-side engagement model, or get stopped by a recipient organization's own mail flow rules regardless of how clean its headers are.