Password Reset and Login Link Email Reliability
Email authentication failures now silently lock users out of accounts with no warning.

A password reset email controls access to an account, not just information or persuasion, so it belongs in a different risk category than almost anything else a company sends. A marketing email that misses its window costs a click. A password reset email that fails to arrive locks a real person out of their account, in real time, with no workaround available until the message shows up.
The temptation is to treat password reset as a template problem: write good copy, pick a clean design, ship it. That framing misses where the actual risk lives. The request has to travel through an application layer, an email service provider, the recipient's mail server, and a spam filter, all before it ever reaches an inbox, and each of those stops can silently swallow the message. The sections that follow work through that path stop by stop, because each one fails in its own distinct way, and each failure looks identical from the outside: a user staring at an empty inbox, locked out, with no idea whether the problem is theirs or yours.
How the Email Authentication Environment Changed
Mailbox providers used to treat weak authentication as a reputation penalty: a message without proper SPF or DKIM might land in spam, but it usually still arrived somewhere. That forgiveness is gone. Yahoo started enforcing stricter requirements in February 2024, and by April 2025 it had escalated to active blocks. Gmail moved from educational warnings into active SMTP-level rejection in its November 2025 enforcement phase. The two largest mailbox providers reject noncompliant mail outright now, instead of quietly downgrading it, and a message blocked at one provider no longer has a fallback at the other. For a password reset specifically, an SMTP rejection does not produce a slow delivery or a spam-folder placement a user can go find. Instead, the message never existed as far as the sender's logs are concerned, and the user is locked out while nobody on the engineering side has any signal that something went wrong. WP Mail SMTP's deliverability guide draws a line between email delivery and email deliverability, and that gap is what decides whether a locked-out user ever gets back into their account.
SPF, DKIM, and DMARC: what each one does and where each one breaks

Authentication misconfiguration is among the leading causes of dropped transactional email, and each of the three protocols in the modern authentication stack fails in its own way. Published domain analysis has found that only a minority of top domains carry a valid DMARC record, which says less about awareness of DMARC's existence than about how easy it is to configure these three systems incorrectly without any immediate symptom.
SPF authorizes which mail servers are allowed to send on behalf of a domain, declared as a DNS TXT record. Only one SPF record is permitted per domain, and a domain with two competing SPF records fails outright, because the receiving server has no way to decide which one to trust. The error that produces this failure is mundane: a team adds a new ESP's SPF include when switching providers or adding a new sending tool, and never goes back to remove or merge the old one, leaving two SPF records live at once.
DKIM attaches a cryptographic signature to the outgoing message, verified by the receiving server against a public key published in the sending domain's DNS. The ESP generates the DKIM keys, but they only work once added to the sending domain's own DNS records, and DKIM signing then has to be explicitly turned on inside the ESP before it actually applies to outgoing mail. Skipping either step leaves messages unsigned even though the sender believes DKIM is configured.
DMARC sits above both and checks alignment: the domain authenticated by SPF or DKIM has to match the domain shown in the message's visible From header. DMARC publishes a policy, none, quarantine, or reject, along with a reporting address, and that policy tells receiving servers what to do with mail that fails alignment. Knock's guide recommends a specific rollout order: start at p=none with reporting turned on, confirm every legitimate sending source is accounted for in the reports, then move to quarantine, then finally to reject. Jumping straight to p=reject before a full inventory of every legitimate sender is complete risks blocking real mail.
Diagnosing problems in this stack does not require guesswork. Running the sending domain through SPF, DKIM, and DMARC record checkers surfaces conflicts directly. Sending a test message and reading the authentication results in the raw headers shows exactly which check passed and which failed. ESP logs often carry an explicit fail signal for any of the three, which is usually the fastest way to find the actual break.
Domain reputation and stream isolation: why mixing transactional and marketing email poisons both
Passing SPF, DKIM, and DMARC does not guarantee inbox placement, because mailbox providers also score the reputation of the sending domain itself. That score gets evaluated per sending identity: bounce rates, spam complaint rates, and recipient engagement all feed into the trust an ISP assigns to a given domain or subdomain.
The failure mode that catches most teams is contamination between unrelated mail streams. If marketing email and transactional email go out from the same domain or IP, a spike in spam complaints from an old, stale marketing list can depress inbox placement for password reset mail on that same domain, even though the reset mail did nothing wrong. The ISP has no way to separate the two kinds of intent; it only sees a domain generating complaints. The fix is structural rather than behavioral: give transactional mail its own subdomain, something like auth.yourdomain.com, so a marketing list problem has no path to touch account-access email. A platform built to manage both transactional and marketing sending under one roof can enforce that subdomain separation as a built-in structural property of how it routes mail, rather than leaving it to a developer to remember to configure correctly on their own.
Watching for reputation damage before it costs deliverability means checking a small number of concrete signals: the reputation report in Google Postmaster Tools, MX blacklist checks, and hard bounce rate trends inside the ESP's own dashboard. Recovery from a damaged domain reputation is slow once it happens, so isolation is worth building up front, before a drop in inbox placement forces reputation repair.
One smaller but telling habit compounds this problem: sending reset email from a no-reply@ address. Knock's guide points out that no-reply addresses signal low legitimacy to ISPs, since spammers use no-reply precisely because they don't want replies, raising the odds of spam filtering, and it takes away a user's ability to reply when something breaks, so more of those users hit the spam-report button. Sending from a monitored address like security@ keeps that channel open and avoids a signal ISPs already associate with low-quality senders.
Token security and the failure modes that appear after delivery
An email that reaches the inbox perfectly can still fail the user if the token inside it is broken. Delivery success and reset success are two separate outcomes, and the token lifecycle is where the second one can fall apart even after the first has gone right.
The OWASP Forgot Password Cheat Sheet requires reset tokens to be single use, to expire after a short, defined period, and to be invalidated server-side the moment the token is redeemed. A token that stays valid after use, or that lives too long before expiring, gives an attacker a usable window that has nothing to do with whether the email itself was delivered correctly.
Logging practices around these tokens carry their own risk. You should never let the reset token or the full reset URL show up in application logs. A keyed digest or an internal request ID gives support staff enough to trace a missing message, but it never exposes a token that could still be redeemed.
The reset endpoint itself needs to behave identically whether the submitted email address belongs to a real account or not: the same outward response, roughly the same response time, and rate limiting on repeated requests. Without that consistency, the endpoint becomes a way to enumerate which email addresses have accounts on the system, which is a security failure independent of whatever happens to the reset email itself.
The content of the message matters for the same reason. A guide from the DEV Community uses a health marketplace example to make the point concretely: the email should confirm that a recovery request was made and give the user the action to take, without including account details that would expose sensitive information if the mailbox itself were shared or already compromised.
Idempotency, retry logic, and the duplicate-send problem
A send that times out gets retried, and that retry is where a second failure mode appears, one that has nothing to do with authentication or reputation. If no idempotency key is attached to the reset request, a retried send produces two separate emails carrying two different tokens. Whichever token gets used first invalidates the other, so the user who clicks the second link, perhaps because it arrived later or they simply picked the wrong one, hits a dead end. The system registers this as a successful send. The user experiences it as a broken link. That gap between what the application logs and what the user actually experiences is the exact shape of the silent failure that makes agentic email sending riskier than it looks on paper, a theme the next section develops further.
An idempotency key scoped to the specific reset request is the right fix: any repeated send attempt that falls inside the key's window returns the same response without dispatching a second email. Retry logic also needs a single clear owner. If the ESP already retries automatically on SMTP failures, adding an application-level retry on top of it is how duplicate sends happen, because now two systems are independently deciding the first attempt failed and acting on that belief at the same time.
If you store the ESP's returned message ID against the reset request row after a successful send, the system has an anchor. That ID is what lets a delivery webhook event get correlated back to the specific reset attempt it belongs to, which is the only reliable way to know, after the fact, which attempt actually reached the recipient's mail server.
Autonomous Agents and These Failure Modes
Every failure mode covered so far, misconfigured authentication, reputation contamination from a shared stream, token races, duplicate sends from uncoordinated retries, gets harder to catch and more damaging once an agent is doing the sending without a human reviewing each dispatch. Conventional transactional email APIs were built on the assumption that a human operator would notice something odd before it spread. An agent that pulls a recipient list from a database and dispatches messages at machine speed has no equivalent check built in unless the sending infrastructure itself provides one.
Catch-all and disposable addresses show this clearly. Some mail servers accept any incoming message whether or not the target mailbox exists, and they send back a success signal even when no real inbox receives it. An agent reading that success signal has no way to know the send was never really seen, and a batch of these accumulates as delayed bounces or silent drops that damage the sending domain's reputation well before any monitoring dashboard shows a problem. A pre-send validation layer, checking MX records, confirming the inbox exists, confirming the address is actively receiving mail, catches this before the send happens rather than after reputation has already taken the hit.
An agent needs machine-readable errors because, unlike a human operator, it cannot interpret a vague bounce message on its own. A person can read a vague bounce message and work out roughly what went wrong. An agent needs a structured error with a specific, codified reason and a remediation action it can either act on directly or surface to a developer. That structure is what lets an agent operate safely at the speed it's capable of, instead of accumulating failures that nobody notices until the domain's reputation has already been damaged.
AgentiSend addresses this by exposing sending through an MCP server as a structured tool call, with annotations that mark sends as destructive operations requiring confirmation. That design gives an agent the same kind of governance surface a human operator already has, a point where a risky action gets flagged before it happens. Every layer this article has walked through, authentication, reputation isolation, token handling, retry discipline, exists because password reset email has no tolerance for silent failure. Building sending infrastructure around that fact, rather than around the assumption that someone is watching every message go out, is what makes handing the keys to an agent a reasonable decision instead of a risk a developer has to accept blind.
