Why conversion alone is not enough
Two stores can both convert at 3% and need different amounts of traffic. One may earn similar profit on most purchases; the other may have a few large baskets, different shipping subsidies, and costly returns. That variation changes how clearly a profit difference stands out.
Conversion and profit per buyer can also move in opposite directions. For example, 2.8% conversion with $44 buyer profit produces $1.232 per exposure. Compared with 3% and $40, that is about 2.7% more profit per exposure, even though conversion fell. The app evaluates the combined outcome.
What the estimates assume
These are approximate comparisons of average profit between two independent groups, split equally, with one planned analysis, a two-sided 5% significance level and 80% power. Power describes how often the design would detect the assumed difference if it were real. It is not the probability that a particular result is correct.
A nonbuyer contributes zero in this model. Buying shoppers have a positive average contribution and the selected variation. Include all their orders, shipping contributions and refunds over the same measurement window. Negative individual contributions are possible; a zero or negative average needs an absolute-profit target instead of this relative-lift model.
The variation relative to average buyer profit stays the same in both groups. Choosing “both” increases conversion and buyer profit by the same factor: for a 10% total profit lift, each rises by about 4.88%. Real changes can alter both averages and variation differently, so use the scenarios to explore sensitivity, not as a store-specific forecast.
Calculation and sources
With conversion probability p, average buyer profit m, and buyer-profit standard deviation s, profit per exposure has mean p × m and variance p × s² + p × (1 − p) × m². This includes the zero-profit nonbuyers.
Exposures per side ≈ (1.96 + 0.8416)² × (V₁ + V₂) / Δ², where V₁ and V₂ are the two exposure-profit variances and Δ is the intended difference in their means. Totals are rounded up to the next hundred. Scaling every profit amount by the same currency factor leaves the estimate unchanged.
The calculation uses the standard error for two independent means documented by NIST and a large-sample normal power approximation, as described in statsmodels power analysis. Heavy tails, rare large losses, clustering, or changing trading conditions can make this approximation unreliable.
Minimum evidence for a decision
At least seven full days, 1,000 exposures per side, and 200 confidently matched orders in total. Confidence and traffic-quality checks must also pass. These minimums are separate from the estimates above. Do not repeatedly check ordinary significance and stop at the first favorable result.
Make a smaller-store test worthwhile
Choose a change large enough to matter to your margins and shoppers. Verify the offer first, keep other promotions stable, and decide in advance what you will do if the result stays uncertain. Small stores can benefit from the offer and investigate patterns without treating an early lead as proven profit.
More traffic is not the only consideration. Seasonality, campaign changes, and the app's allowed duration limit how long one comparison stays useful. These estimates may exceed what your store can collect during one test.