How to run a clean before-and-after email deliverability retest
Run a clean retest by changing one failure class, preserving every other material variable and comparing the predicted component verdict—not only the total score.
- Define one predicted verdict change before editing anything.
- Keep the sender, ESP, message and seed set constant unless one is the treatment.
- Compare the target component and regressions, not only the total score.
Primary sources for this guide
- Folderly Flash scoring and privacy methodology
- Gmail email sender guidelines
- Gmail Postmaster Tools API
- RFC 7208 - Sender Policy Framework
- RFC 6376 - DKIM Signatures
Define the claim first
Write one falsifiable sentence: changing one configuration or message element should change one observable verdict while the rest of the test remains materially the same. 'Publishing the correct DKIM key should turn the DKIM verdict from fail to pass for the same selector' is testable. 'This should improve deliverability' is not.
Capture the baseline
Save enough context to reproduce the first result.
- Test timestamp and result ID.
- Visible From and Return-Path domains.
- ESP or sending system.
- DKIM signing domain and selector.
- Subject, body version, links and message type.
- Relevant DNS values and configured provider seed set.
- Authentication, content, blacklist, hygiene and placement verdicts.
Create a control table
Write the baseline and retest values for the sending domain, ESP account, message body, tracking and destination links, seed set, send time and intended treatment. Mark each non-treatment variable as same, materially equivalent or changed. If a supposedly controlled value changed, record it before reading the result. This simple table makes an inconclusive comparison visible before a stakeholder turns it into a success story.
Change one failure class
One logical intervention may require several implementation steps. Enabling DKIM can require a key, DNS record and ESP setting, but those steps still address one failure class. Do not combine that work with a new provider, rewritten body, different domain and changed list; the result would have too many possible causes.
Wait until the change is observable
DNS caches retain old values until TTLs expire, ESP changes require a fresh message, and placement observations can settle after the deterministic setup grade. Record the change time, confirm it through an independent lookup and send the retest only when the intended state is visible. Repeated edits during propagation destroy the comparison.
Predefine invalidation conditions
Decide what would make the pair unusable: a different sending pool, missing seed providers, a sitewide redirect change, an ESP migration, an algorithm or policy event, or a message version that cannot be reconstructed. Do not choose these conditions after seeing the score. A retest that meets an invalidation rule can still reveal a problem, but it should not be presented as clean before-and-after proof.
Compare components
Ask whether the predicted verdict changed, whether any passing verdict regressed, whether the observed message identity matches the intended configuration and whether the same provider set reported. A flat total can still contain a successful repair; a higher total can still hide a new blocker.
Use stronger evidence for reputation
Authentication can be verified on a received message. Reputation recovery usually cannot. Review comparable sends over time, provider diagnostics, complaints, hard bounces and recipient quality before declaring a domain or IP recovered. One seed test is one observation.
Schedule the observation window
A configuration verdict can be read as soon as the new message and DNS state are observable. Reputation and engagement conclusions need repeated comparable sends and a defined review window. Set the checkpoints before the remediation—such as immediate setup verification followed by 7-, 28- or 60-day trend reviews appropriate to the mailstream—so the team does not stop measuring on the first flattering day.
Write a decision log
For each checkpoint, record the observation, inference, uncertainty and next decision as separate fields. The observation might be that DKIM now passes for the intended selector. The inference is that the new key fixed that authentication failure. The uncertainty is whether reputation and placement will respond over time. The decision may be to resume a limited send while monitoring provider diagnostics. Keeping those statements separate lets another operator challenge the interpretation without disputing the received-message evidence, and it prevents the final report from upgrading a bounded fix into a guarantee.
End with a decision
Classify the retest as resolved, partially resolved, not resolved or inconclusive. Keep the baseline, change log, retest and conclusion together. If the tests are not comparable, say so and run a cleaner experiment rather than manufacturing confidence from the score. Name the owner, next checkpoint and rollback condition so the protocol ends in an operating commitment instead of a dashboard observation.
Teams that preserve comparable test pairs will build a remediation evidence base; teams that chase totals will repeat the same incident with different explanations.
Limitations
A before-and-after pair supports only the predefined claim when material controls remain comparable. One seed observation cannot prove reputation recovery, represent a whole recipient population or guarantee future inbox placement.