Everyone talks about "testing their emails." Few teams do it right. A poorly designed A/B test leads to false conclusions that worsen results rather than improve them.
The basic principle: one element at a time
This is rule number one—and the one that’s most often broken. Testing a different object AND a different piece of text in the same test doesn’t tell you anything: there’s no way to know which of the two elements made the difference.
Each test focuses on a single element. Same object, different content. Or same content, different object. Never both at the same time.
What to Test First
Start with the subject line. It determines the open rate, which in turn determines everything else. Two formats to test: question vs. statement (“Are you already using automation?” vs. “How [Client] doubled its appointments”). Then, highly personalized vs. neutral (“Following your SDR hire” vs. “A quick question”).
Next, the CTA. Open-ended vs. closed-ended questions. “Available to discuss this?” vs. “15 minutes on Tuesday or Thursday?” Specifying a specific time in the CTA often increases the response rate by 20 to 40%.
Message length. Short version (60 to 80 words) vs. medium version (120 to 150 words). In some segments, the short message outperforms. In others, a little more context is reassuring.
The opening line. LinkedIn post vs. client results vs. direct question. This is often the test that yields the greatest differences in performance.
The minimum sample size
This is the most common mistake: drawing conclusions based on 30 or 50 test runs. You need at least 200 test runs per variant to obtain statistically meaningful data. With fewer than that, you’re making decisions based on noise.
If your list is too small to reach this threshold, test across several consecutive campaigns and aggregate the data.
How to Interpret the Results
Look only at the final response rate, not just the open rate. An email that generates a 60% open rate but a 2% response rate is less effective than one with a 40% open rate and a 10% response rate.
Wait 7 to 10 days after the last message in the sequence before drawing conclusions. Late responses (Day 14, Day 21) skew the results if the analysis is conducted too early.
To analyze these results using the right metrics, B2B lead generation KPIs outline the metrics to track at each stage. And for the tests to be valid, cold email deliverability must be consistent between the two variants; otherwise, you’re testing deliverability, not the content.
.png)


