Klaviyo A/B testing: campaign and flow method
Short answer. Write one hypothesis, change one meaningful variable, randomly split the same eligible audience, choose the primary metric and evaluation window before launch, and keep unsubscribe, complaint, and margin as guardrails. Treat Klaviyo's winner label as evidence for that test context, not a permanent rule for every audience.
A/B testing is useful when it changes a decision. It is not useful when variant A is a red button with a new subject line, different offer, later send time, and shorter email while variant B changes none of those things.
Klaviyo supports testing across campaign, flow, and form use cases. The exact options and winner metrics vary by asset, so define the business question before opening the builder.
The experiment brief
Complete this table first:
| Field | Example |
|---|---|
| Decision | Which value proposition should lead the launch? |
| Hypothesis | Product-specific utility will increase click rate versus broad brand language |
| Audience | Consented active prospects interested in category |
| Variable | Hero proposition only |
| Control | Existing category-led proposition |
| Primary metric | Unique click rate |
| Guardrails | Unsubscribe, complaint, conversion, margin |
| Evaluation window | 48 hours after send |
| Action threshold | Adopt only if result is meaningful and no guardrail deteriorates |
If the team cannot fill this table, it is not ready to test.
What to test first
Prioritize variables close to a customer decision:
- Audience and lifecycle state.
- Offer or value proposition.
- Timing after a behavior.
- Message count and sequence.
- Product or category relevance.
- CTA and creative hierarchy.
- Subject line and preview text.
Subject lines are easy to test, but they optimize an imperfect open signal unless clicks or orders select the winner. Apple Mail Privacy Protection can generate machine opens, so an open-rate winner may not represent more human attention.
Choose the winning metric from the hypothesis
| Hypothesis | Primary metric | Why |
|---|---|---|
| Subject makes message easier to recognize | Open rate with privacy caveat, plus click guardrail | Closest observable inbox response |
| Hero proposition creates interest | Unique click rate | Measures action after message view |
| Offer increases purchase behavior | Placed order rate or net revenue per recipient | Matches commercial outcome |
| New flow timing improves conversion | Order rate over full buying window | Captures delayed behavior |
| Shorter form improves acquisition | Confirmed subscription rate | Avoids optimizing raw submit only |
| Product recommendation improves value | Net revenue or contribution per recipient | Accounts for order value and margin |
One practical constraint: the campaign builder only offers three winner metrics, open rate, click rate and placed order rate, the last one on eligible accounts only. Revenue or contribution per recipient is not in that selector, so pick the closest available metric for the automated winner and compute the commercial figure yourself from the results report.
The open-rate caveat is not only statistical on a French list: check what the CNIL changes for open tracking before building a protocol on that metric alone.
Do not choose the metric after seeing the charts. That turns normal noise into a story.
Statistical significance in Klaviyo
Klaviyo's current campaign significance documentation says a campaign result is labeled statistically significant when each variation has at least 50 recipients and the win probability reaches at least 90% for the selected winner metric.
That is a platform decision rule, not a universal sample-size calculator. Fifty recipients per variant is rarely enough to detect a small difference in a low-rate outcome such as orders. Required sample size depends on baseline rate, minimum effect worth detecting, number of variants, split, and desired error tolerance.
Use a pre-test sample-size calculation for high-stakes decisions. If the list is too small, run fewer variants, test a larger behavioral difference, or collect evidence across comparable sends without pretending each send is identical.
Campaign A/B test process
1. Select one eligible audience
Use the same inclusions, exclusions, consent, frequency, and send window. Randomization should be the only reason otherwise comparable profiles see different versions.
2. Create the control
The control is the current best-known version, not an intentionally weak message. A weak control exaggerates lift and wastes audience.
3. Create one meaningful variant
Change the variable described in the brief. If testing proposition, keep sender, subject, preview, layout, products, offer, audience, and send time constant where possible.
4. Set split and winner behavior
Choose how much of the audience enters the test and whether Klaviyo automatically sends a winner to the remainder. Automatic winner rollout can be useful for time-sensitive campaigns, but it also means the winner is selected on an early window. For durable learning, a full random split may provide cleaner comparison.
5. Freeze the analysis plan
Record metric, attribution settings, end time, guardrails, and excluded anomalies before the send.
6. Analyze the full chain
Review delivery, opens, clicks, onsite behavior, orders, revenue per recipient, unsubscribes, complaints, returns, and margin where relevant. A click winner with weaker conversion may have created curiosity without purchase intent.
Flow A/B testing
Flows introduce additional complexity because profiles enter over time and the business environment changes.
Test one of:
- Message content.
- Time delay.
- Number of messages.
- Branch strategy.
- Offer.
- Channel sequence.
Keep trigger, flow filters, conversion exits, Smart Sending, consent, and re-entry rules consistent across variants.
Flow email tests behave differently from campaign tests in two ways. The winner metrics are limited to open rate and click rate, and Klaviyo recommends click rate because Apple Mail Privacy Protection can inflate opens. The winner is selected as soon as the metric reaches statistical significance or at a date limit you set, whichever comes first, with manual selection still available.
Use a long enough evaluation window
A post-purchase test may need the next expected order cycle. A cart test may resolve within days. Define the observation period from customer behavior, not dashboard impatience.
Avoid contamination
Campaigns and other flows can contact the same profiles during a test. Record frequency rules and major concurrent promotions. If one variant is exposed to a different campaign mix, interpretation becomes harder.
Test incrementality, not only variants
A content A/B test answers which message performed better. It does not answer whether sending either message was better than sending none. Add a randomized holdout when the decision is whether the flow creates incremental behavior.
Form A/B testing
Optimize for valuable consent, not only raw submit rate. A form variant can increase email capture while reducing confirmed opt-in, first purchase, or long-term engagement.
Track:
- Unique form views.
- Submit rate.
- Confirmed subscription rate.
- Email and SMS consent separately.
- Welcome click and first-order rate.
- Unsubscribe and complaint rate.
- 90-day active or customer rate.
Keep targeting and traffic comparable. A mobile-only variant and desktop-heavy control are not a valid creative test.
Klaviyo's own form test reports unique views, unique submits, submit rate, orders, average order value and revenue alongside a win probability, so the platform covers the first half of that list and your own analysis covers the consent and retention half. Note that form variations cannot be edited once a test is live: the test has to be ended first, so settle the creative before launch.
What not to combine in one test
- Subject line and sender name.
- Offer and audience.
- Email body and landing page.
- Send time and segmentation.
- Flow delay and message count.
- Form creative and traffic source.
If the goal is to compare two complete strategies, label it as a package test. You can learn which package wins, but not which component caused the difference.
Build a test log
| Field | Purpose |
|---|---|
| Test ID and dates | Reproduce the experiment |
| Asset and audience | Preserve context |
| Hypothesis | Prevent retrospective storytelling |
| Variant difference | Confirm one variable changed |
| Primary metric | Define winner |
| Guardrails | Protect customer and economics |
| Sample and split | Evaluate power and balance |
| Result and uncertainty | Record evidence honestly |
| Decision | Adopt, retest, reject, or no change |
| Follow-up | Turn result into next question |
Add screenshots or exported results because live dashboards and attribution settings can change.
Interpret common outcomes
Statistically significant and commercially useful
Adopt in the tested context, monitor after rollout, and test whether the principle transfers to a comparable audience.
Significant but tiny effect
Estimate annual incremental value and implementation cost. A detectable effect may still be too small to matter.
Promising but uncertain
Retest the same hypothesis on a comparable campaign. Do not change the variant definition to chase a win.
No meaningful difference
Keep the simpler or lower-risk version. A null result prevents unnecessary work.
Guardrail deterioration
Reject or redesign even when the primary metric wins. More clicks with more complaints or lower margin is not a clean improvement.
Common Klaviyo testing mistakes
- Testing many variables at once.
- Selecting the winner metric after launch.
- Using opens without Apple privacy context.
- Stopping when one version briefly leads.
- Declaring a rule from one small test.
- Comparing segments that were not randomized.
- Ignoring attribution settings and concurrent campaigns.
- Testing content but never testing whether the flow is incremental.
Testing methodology matters less if the underlying Klaviyo setup has gaps, so consider a free 30-minute Klaviyo audit before you scale the winning variant.
FAQ
How many recipients does a Klaviyo A/B test need?
Klaviyo requires at least 50 recipients per campaign variation for its significance label, but a useful sample can be much larger. Calculate it from baseline, minimum detectable effect, and the selected metric.
What should I test first in Klaviyo?
Start with a meaningful customer decision, such as audience, offer, timing, or proposition. Test subject lines when inbox recognition is the actual question.
Should the winner be based on opens or clicks?
Choose the metric that matches the hypothesis. Clicks are usually more dependable for engagement because Apple privacy behavior affects opens. Use orders or revenue when purchase is the question.
Can I A/B test a Klaviyo flow?
Yes. Open the email inside the flow builder and select Create A/B test in the email details panel. Variations are weighted equally by default and the percentages can be changed. Keep trigger, filters, and evaluation window consistent, and distinguish variant performance from incrementality.
What if the result is not significant?
Do not force a winner. Keep the control or simpler version, record the result, and decide whether a larger or more distinct test is worth running.
Turn experiments into accumulated knowledge
Testing compounds only when hypotheses, context, results, and decisions are preserved. Deliver helps ecommerce teams build that experimentation system across campaigns and flows, or you can talk to our Klaviyo agency directly. Book a Klaviyo and CRM diagnostic.
Provenance and verification
Checked on 2026-08-10 against the vendor pages listed in sources: campaign winner metrics (open rate, click rate, placed order rate), flow winner metrics limited to open and click rate, significance labels (50 recipients minimum per variation, 90% win probability, 1,800 recipient tier), sign-up form A/B test entry point and reported metrics, and Apple Mail Privacy Protection open inflation. Claims the documentation does not settle were removed rather than rephrased.
- Sources checked on
- Reviewed by
- Claude (local CLI) counter-check against official Klaviyo and CNIL documentation
- AI assistance
- Yes
- Sources
-
- help.klaviyo.com/hc/en-us/articles/115005228148
- help.klaviyo.com/hc/en-us/articles/360052793012
- help.klaviyo.com/hc/en-us/articles/360052794352
- help.klaviyo.com/hc/en-us/articles/6960371049115
- help.klaviyo.com/hc/en-us/articles/360045462071
- help.klaviyo.com/hc/en-us/articles/4416791883163
- help.klaviyo.com/hc/en-us/articles/360045012632
- www.cnil.fr/fr/recommandation-pixel-suivi-courriels
Want to apply this to your stack?
Spend 30 minutes with Charlotte to review your CRM setup, size the opportunity and leave with a practical action plan.
Book a 30-minute call →