CRO and Experimentation
What I Look for Before Launching an A/B Test
A practical checklist for deciding whether a test has a real hypothesis, reliable measurement, enough traffic, and a business question worth answering.
Most tests fail before they launch
A test does not fail because the variation lost. It fails because it was never capable of giving a clear answer. Underpowered, badly measured, or built around a change too small to matter. So most of my work on experimentation happens before anything goes live, in deciding whether the test deserves to run at all. A losing test that teaches me something is fine. An inconclusive test that burned three weeks of traffic is the one I try to prevent.
A hypothesis, not a hunch
"Let us try a green button" is a hunch. A hypothesis names the problem, the change, and the expected effect: because visitors are not sure the product is legitimate, adding reviews near the buy button will raise add to cart rate. That structure forces me to have a reason, and it tells me in advance what result would confirm or kill the idea. If I cannot write the sentence, I am not ready to test.
Enough traffic to be conclusive
This is where most tests quietly die. If a page gets a few hundred visits a week and converts at two percent, a small lift can take months to reach significance, and the business will not wait that long. Before launching I do a rough sanity check on the current conversion rate, the traffic, and the size of the change I could realistically expect. As a rule of thumb I want each variation to reach at least a few hundred conversions before I trust the result, and I would rather run one well-powered test than three that never resolve. If the traffic is not there, I do not run an underpowered test and pretend. I test a bigger change, test higher up where traffic is greater, or make the improvement as a judgment call and move on.
One primary metric decided in advance
I pick the single metric the test is meant to move and I write it down before launch, along with the decision each outcome triggers. If I wait to see the data and then go looking for any metric that improved, I will always find one, and I will fool myself. Guard metrics matter too. A checkout change that lifts conversion but quietly raises refunds is not a win, so I watch the downstream number, not just the one I am optimizing.
A change big enough to matter
Testing a headline word against another word usually produces noise, because the effect is smaller than the natural variation in the data. I focus tests on things large enough to move behavior: the offer, the page structure, what appears above the fold, the number of steps in a flow. If a variation is barely different from the control, the most likely result is no result.
Clean measurement and honest QA
Before launch I confirm the tracking actually fires, the two variations are truly identical except for the change, the split is random, and the experience is not broken on mobile. I have seen "winning" tests that were really just a broken tracking tag. Five minutes of QA protects weeks of traffic.
Decide the decision before you look
The last thing I do before launch is agree what happens at each outcome. If it wins, we ship it. If it loses, we keep the control and note why. If it is flat, we stop and spend the traffic on a bigger question. Deciding this in advance is what keeps a test from turning into an argument about interpretation after the fact.
Related
More From the Journal
Planning a Small Business Campaign With a Limited Budget
How I narrow the audience, channel, offer, and measurement plan when the budget is too small to test everything at once.
Building an E-commerce Funnel From Awareness to Retention
How I connect acquisition, merchandising, landing experience, conversion, and lifecycle into one growth system instead of four disconnected campaigns.
Open to relevant marketing opportunities
If this way of thinking fits your team, review my experience or start a conversation.
View Resume and Contact