Most brands running A/B tests on subject lines are not actually learning anything. A test on a list of 800 people, decided after two hours, comparing two subject lines that differ only by an emoji, produces a result that looks decisive and means almost nothing. The framework below fixes the parts that usually go wrong.
List size determines whether a test is worth running
Statistical significance requires enough recipients per variant to separate a real effect from random noise. As a practical floor, testing on segments below a few thousand recipients per variant rarely produces a difference large enough to trust. Below that size, differences you see are as likely to be noise as signal, and acting on them trains you to chase false patterns.
If your list is small, that is not a reason to skip testing entirely, it is a reason to test bigger, more structural changes (send time, sender name, offer framing) rather than small wording tweaks that need a larger sample to detect reliably.
Test one variable at a time
Changing the subject line wording, the emoji, and the preview text simultaneously means you cannot attribute the result to any single change. Isolate the variable:
- Curiosity vs. clarity. "You won't believe this" vs. "20% off your next order." One of these usually wins consistently for a given list; find out which.
- Personalization. First name in the subject line vs. none. Effect size varies a lot by brand and audience.
- Length. Short, punchy subject lines vs. longer, more descriptive ones. Mobile preview truncation matters here.
- Urgency framing. "Ends tonight" vs. no urgency language. Works well when true, damages trust fast when it isn't.
Let the test run long enough
Deciding a winner after one or two hours captures only the earliest openers, who are not representative of your full list's behavior. Most platforms let you set a test window before the winning variant auto-sends to the remainder; give it enough time to include recipients across different time zones and check-in habits, generally at least several hours, ideally closer to a full day for anything but time-sensitive sends.
Open rate is the wrong metric to optimize alone
A subject line that inflates opens through curiosity-gap tactics but does not match the email's actual content will show a great open rate and a weak click and conversion rate, because the mismatch between promise and content erodes trust. Track click rate and, where possible, conversion or revenue per recipient as the real scoreboard. Open rate is a useful early signal, not the goal.
This matters more since Apple's Mail Privacy Protection began pre-loading images for a large share of inboxes, which inflates reported opens independent of whether a human actually read the email. Open rate differences between two subject lines can partly reflect this noise rather than a real behavioral difference.
Build a testing calendar instead of testing randomly
Testing one variable in isolation, on one campaign, tells you something about that campaign. A running calendar, testing the same variable type (say, urgency framing) across several campaigns over a few months, tells you something durable about your audience that you can apply as a default going forward. That compounding knowledge is the actual point of testing, not the win on any single send.
Setting this up on TheMarketer
We run structured tests as part of the standing campaign calendar in every account, sized correctly for that list and tracked to revenue per recipient rather than open rate alone, so the learnings actually change how future campaigns get written.