← All articlesMeasurement

How to run an email campaign test that answers a question

5 min read

An email campaign test should help you make a decision. “Does a sizing guide help shoppers choose?” is more useful than “Which email looks better?” Define the change, the eligible audience, and the result you will compare before sending. Whether you organize the comparison manually or use available testing tools, avoid changing several things and crediting the result to just one.

Isolate the change

Keep the audience rules, offer, destination, and send conditions comparable. Change one meaningful element, such as the headline angle or the presence of sizing guidance. If everything changes together, you may identify a better package but not understand which part mattered.

Use the campaign and audience controls available in your Sendvio account to organize the comparison. If a workflow requires manually prepared groups, document how they were assigned and keep them non-overlapping. Do not assume an automatic testing feature or winner-selection behavior that you have not verified.

Write a test brief that can be checked afterward

A useful brief might say: “For eligible subscribers considering the same cover collection, adding a concise size explanation will increase qualifying purchases per assigned recipient over the chosen observation period.” Name the change, audience, expected outcome, and time window. Keep the offer and product availability comparable.

Where the available process supports it, randomly assign eligible recipients to non-overlapping groups. Splitting alphabetically or sending one version to the most recent subscribers can introduce systematic differences. If important customer groups need balanced representation, plan that deliberately rather than repairing the explanation after results arrive.

Decide whether you are testing one element or a complete package. Both can be useful, but they answer different questions. If you change subject, layout, discount, and destination together, the result may favour one package without telling you which individual change caused it. Label the test honestly.

Decide what success means beforehand

Choose a primary outcome connected to the question, such as orders per eligible recipient. Keep unsubscribes, complaints, and relevant costs as guardrails. An open-rate increase alone may not represent more human attention or better commercial results.

Set the observation period before launch and allow for the normal buying cycle. Repeatedly stopping as soon as one version looks ahead can produce misleading conclusions. Small samples may remain inconclusive even when one percentage is visibly larger.

Plan how much evidence the decision needs

A change from two purchases to three is a 50% relative increase, but it is also a difference of one purchase. That is weak evidence for a permanent rule. Record absolute counts and percentage-point differences alongside relative lift so the presentation does not exaggerate a small result.

For an important decision, estimate sample needs from the baseline rate and the smallest effect worth acting on, using an appropriate statistical method or qualified help. There is no single sample size that makes every email test reliable. Rare outcomes and small expected differences usually require more evidence than common outcomes and large changes.

Set the observation period and stopping approach before launch. Repeatedly checking and stopping at the first favourable result can distort the conclusion unless the analysis was designed for that process. If the campaign must end for a commercial reason, record the limitation instead of calling the partial result definitive.

Read the result in counts, rates, and commercial terms

Suppose two comparable, randomly assigned groups each contain 5,000 eligible recipients. Version A produces 100 unique purchasers and version B produces 115 during the agreed window. The observed purchaser rates are 2% and 2.3%. That is a difference of 0.3 percentage points, or a 15% relative increase from the 2% baseline. These are illustrative results, not evidence about a real Sendvio campaign.

The larger percentage does not settle the decision. An appropriate statistical analysis still needs to assess uncertainty, and the business needs to decide whether the possible improvement is worthwhile. Check whether discounts, stock, delivery, or tracking differed unexpectedly. Keep the original assigned groups in the analysis rather than excluding people after the fact because they did not open or click.

Then inspect the guardrails. If the second version creates more purchases but also a costly misunderstanding about included items, the result is not simply a success. Review contribution and relevant negative feedback over the planned period. Do not choose a new primary metric after seeing which number makes the preferred version look best.

Write the conclusion in proportion to the evidence: what changed, what was observed, how uncertain it is, and what action follows. If the result remains inconclusive, a simpler version may still be a reasonable operational choice, but describe that as a practical decision rather than proof that its conversion rate is better.

Record the lesson, including uncertainty

Save the versions, audience rules, sample sizes, timing, and result. If the evidence is weak, say so and avoid promoting a tentative difference into a permanent rule.

Use the result to choose the next practical improvement. A test that shows little difference can still save time by suggesting the team should focus elsewhere. The point is a better decision process, not a calendar full of experiments whose results nobody uses.

Check practical importance as well as statistical evidence. A tiny improvement that requires expensive production may not be worth maintaining. A version that increases orders but also creates misleading expectations or more costly returns may fail the broader decision even if its primary metric improves.

Close the test with one of three honest outcomes: adopt with the evidence described, keep exploring because the result is uncertain, or reject because the change does not justify its cost or harm. Preserve the brief and counts. That record makes future experiments cumulative instead of turning each campaign into a new opportunity to discover a “winner” from noise.

Put it into practice

Explore campaign tools