The way to settle the "should we make this change or not" debate with data instead of opinion is A/B testing. But most teams either test the wrong thing or misread the right thing. In this article we cover the fundamentals of experiment-driven optimization, from building a hypothesis to sample size to reading results correctly.
What exactly is an A/B test?
An A/B test is a method of comparing two (or more) different versions of the same page or flow at the same time, showing one version to part of your visitors and another version to the rest. The goal is to statistically and reliably determine which version performs better on the metric you've defined (conversion rate, add-to-cart, click-through).
The critical phrase here is "statistically reliable." Publishing two versions back-to-back in different weeks and comparing the results is not an A/B test; the difference between those two weeks could stem from seasonality, a campaign, or pure chance. In a real A/B test, visitors are randomly split into two groups within the same time window, so external factors are distributed equally across both groups.
Hypothesis first, then the test
A good A/B test doesn't start with a random curiosity like "I wonder what happens if we change the button color" — it starts with a clear hypothesis. A solid hypothesis contains three parts: the problem you're observing in the current state, the change you're proposing, and the reasoning for why that change should work.
Example: "According to our analytics data, 40% of users abandon their cart when they see the shipping fee on the checkout page (problem). If we show the shipping fee earlier, on the cart page (change), the feeling of surprise cost will decrease, so the checkout completion rate will increase (reasoning)." A hypothesis built this way teaches you something regardless of the test outcome; random experiments, even when the result is positive, can't explain "why it worked." Pulling the data that feeds your hypothesis from the right place matters too — we covered how to see exactly where you're losing users in the funnel in our setting up and reading GA4 correctly article.
What should you test, and what shouldn't you?
Trying to test everything wastes your resources. Focus on high-impact areas:
- High-traffic pages: The homepage, best-selling product pages, and the checkout flow; getting a meaningful result on a low-traffic page can take months,
- Decision points: Call-to-action copy and color, price display, shipping-fee transparency, product image order,
- Trust elements: Placement of reviews and ratings, secure-payment badges, return-policy emphasis.
Conversely, low-impact details like shifting the brand logo a few pixels usually aren't worth testing; such changes can be left to design preference. Reserve your resources for changes likely to actually move the needle. For sensitive areas like price display, reading our pricing psychology article before setting up a test makes it easier to see which variations are worth trying.
The table below summarizes example variations and metrics to measure for four commonly tested areas:
| Test type | Example variation | Metric to measure |
|---|---|---|
| Headline test | Feature-focused headline vs. outcome-focused headline | Scroll rate, form starts |
| Image test | Studio shot vs. in-use lifestyle shot | Product page conversion rate |
| CTA test | "Add to Cart" vs. "Get It Now" | Click-through rate, add-to-cart |
| Price test | Full price display vs. installment-emphasized display | Checkout completion rate |
One variable at a time
The most common mistake is changing more than one thing in the same test: altering the headline, the image, and the button text all at once, then squeezing the result into a single "version B is better" statement. Even if the result is positive, you'll never know which change caused the difference, and the lesson learned can't be carried over to the next page.
Testing a single variable may look like a slower path to results, but over time it accumulates far more, and far more actionable, learning. If you want to try multiple changes at once, use a method designed for that, like multivariate testing; it requires much more traffic and is generally impractical for small-to-medium stores.
"An A/B test doesn't give you the right answer; it shows you how well you asked the right question."
Sample size and duration: don't stop early
Stopping a test after a few days because "version B seems to be ahead right now" is the second most common mistake. Natural fluctuations in traffic can create a misleading early lead in the first days of a test; over time that gap can shrink or even reverse.
Before launching a test, roughly estimate how long it will take: your current conversion rate, the improvement you expect, and your daily traffic determine the sample size you need. Online sample size calculators give you this estimate in minutes. As a general rule, don't draw conclusions before running the test for at least one or two full weekly cycles (to capture the weekday-weekend behavior difference), and don't stop before reaching the statistical significance threshold (usually a 95% confidence level).
If getting a meaningful result on low-traffic pages is difficult, test an earlier step in the funnel instead of the purchase itself — micro-conversions like add-to-cart or form starts; sample size accumulates faster on these steps, so you can reach a decision sooner.
To sum up, a healthy A/B test proceeds in this order:
- Write a clear, data-driven hypothesis (problem, change, reasoning),
- Identify a single variable and choose the metric to measure,
- Calculate the required sample size and estimated duration,
- Run the test uninterrupted for at least one to two full weekly cycles,
- Read the result by segment, not just on the overall average.
What to watch for when interpreting results
The test is over, and version B won; so what now? First, check whether the result is consistent across segments: the version that wins on mobile might be losing on desktop. Checking the result across device, traffic source, and new-vs-returning-visitor breakdowns surfaces contradictions that the overall average hides.
Also, "statistically significant" and "meaningfully large" are different things. A 0.2% improvement can be statistically significant without making a practical difference to your business. Before rolling out the winning version, calculate what that difference actually translates to in revenue or operational terms; you can find other ways to increase conversion rate beyond A/B testing in our 7 proven methods to increase conversion rate article.
Pre-test checklist
- Have you written a clear hypothesis (problem, change, reasoning)?
- Are you testing a single variable?
- Have you calculated the required sample size and estimated test duration?
- Are you planning to run the test for at least one full weekly cycle?
- Will you check the result by segment (device, channel)?
- Will you evaluate the practical/commercial impact alongside statistical significance?
A/B testing is one of the most practical tools for reconciling intuition with data; done right, it moves debates out of the meeting room and into real user behavior. Because product page, cart, and checkout flow components in stores built on Şimşek Software infrastructure are flexibly configurable, testing one version against another means changing a few settings in the panel rather than writing code.