People ask me how to A/B test emails all the time, usually right after a subject line flopped, and the honest answer is most tests fail not because the idea was bad but because the test itself was never big enough or clean enough to tell you anything.
An A/B test splits your list into two random groups, sends each group one version of an email that differs by exactly one thing, and compares the result. That's the whole mechanism, and the value comes entirely from that one word: random. If the split isn't random, or if the two groups differ in some other way, one gets sent an hour later, one is a different segment, you're not measuring the variable you think you're measuring, you're measuring noise.
Most email platforms handle the random split for you, which is good, because doing it by hand is where people mess up. What they don't handle is picking a variable worth testing in the first place. Testing button color when your open rate is the real problem is testing the wrong thing, so before you set anything up, decide what number you're actually trying to move, whether that's opens, clicks, or replies, and test toward that.
Subject lines move the needle more than almost anything else, because they decide whether the email gets opened at all, and everything downstream depends on that first decision. If you've never run a split test, start there. Write two genuinely different subject lines, not two versions of the same joke, and see which one people actually open.
Once you've got a handle on subject lines, preview text is next, since it's the second thing a subscriber sees before opening and most people leave it to whatever the first line of the email happens to be. After that, test the call to action itself: a text link inside a paragraph versus a standalone button, or the wording of the button. Save layout and design tests for last. They matter, but they move the smallest amount of the outcome, and small effects need bigger samples to detect, which is a problem I get into next.
This is where most tests quietly fail. If you send a subject line test to 200 people, 100 per variant, and one gets 22 opens and the other gets 25, that's not a winner, that's random variation you'd see again if you reran the exact same email twice. Small samples produce big swings that look meaningful and aren't.
There's no single magic number, but as a rough guide, you want at least a few thousand recipients per variant before a gap in open rate or click rate starts to mean something reliably, and smaller gaps need larger samples to trust. If your list is smaller than that, don't force a single-send test. Run the same test structure across several campaigns and look at the pattern across all of them rather than betting everything on one send.
The most common one is changing more than one thing at a time and then crediting the win to whichever variable you had a hunch about. If you change the subject line and the send time in the same test, you don't know which one did the work.
The second is calling the test too early. Opens and clicks trickle in for hours after a send, sometimes a full day for B2B lists that check email on a work schedule, so a variant that's ahead after twenty minutes can easily flip by the next morning. Wait for the test to actually finish before you declare a winner.
The third is ignoring novelty. If you always send from the same subject line style and suddenly try something different, part of the lift is just the pattern break, not the new approach being inherently better. That's fine, but don't assume the same trick works forever once it stops being new.
Roll the winning variant out to the rest of the list if you tested on a portion, and write down what won and why you think it worked, because six months from now you'll be tempted to retest the exact same question. Keep a simple swipe file of subject lines and CTAs that outperformed, sorted by what kind of email they were attached to.
One thing worth saying plainly: if your open rates are low across every variant in a test, no subject line is going to fix that, because the problem usually isn't the words, it's where the email is landing. I've had clients run beautiful subject line tests for months without moving their real numbers because the emails were sitting in Gmail's Promotions tab or getting delayed into spam, and no amount of copy testing solves a placement problem. If that sounds familiar, it's worth checking your Inbox Scorecard before you spend more time optimizing words nobody's seeing. For more on the send-time side of testing, see best time to send marketing emails, and if you haven't segmented your list yet, list segmentation will make every test you run afterward more accurate.
Long enough to reach your full sample size, not just the first hour or two. Opens and clicks keep coming in well after a send, especially for B2B audiences, so calling a winner early is one of the most common ways tests go wrong.
There's no universal number, but as a rough guide you want at least a few thousand recipients per variant before a difference in open or click rate is reliable. Smaller lists should look at patterns across several sends rather than trusting one test.
You can, but then you're running a multivariate test, not a simple A/B test, and you need a much larger sample to separate which change caused the result. If you're just getting started, change one thing at a time.
No, sending two versions to a randomly split list doesn't affect deliverability on its own. What affects deliverability is sender reputation, authentication, and engagement, none of which a normal split test touches.
I test subject lines for a living too, but half the deliverability audits I run turn up a placement issue no amount of copywriting would have fixed. If your opens have been flat no matter what you try, get a free <a href="/audit">deliverability audit</a> before you run another test.