A/B Test Sample Size Calculator for App Store & Google Play Screenshots
Before you trust an App Store Product Page Optimization or Google Play Store Listing Experiment result, check whether you actually collected enough visitors. Enter your numbers below for an instant answer, or use the worked-example table for common baselines. Free, no signup.
Calculator
Fixed at 95% two-sided significance and 80% power — the two most common defaults. See the FAQ below for why those, and why this needs at least a week regardless of the number above.
Worked examples
| Baseline rate | Relative MDE | Treatment rate | Per variant | Total |
|---|---|---|---|---|
10% |
5% | 10.5% | 57,760 | 115,520 |
10% |
10% | 11% | 14,749 | 29,498 |
10% |
20% | 12% | 3,839 | 7,678 |
10% |
30% | 13% | 1,772 | 3,544 |
20% |
5% | 21% | 25,580 | 51,160 |
20% |
10% | 22% | 6,507 | 13,014 |
20% |
20% | 24% | 1,680 | 3,360 |
20% |
30% | 26% | 769 | 1,538 |
30% |
5% | 31.5% | 14,853 | 29,706 |
30% |
10% | 33% | 3,760 | 7,520 |
30% |
20% | 36% | 961 | 1,922 |
30% |
30% | 39% | 435 | 870 |
40% |
5% | 42% | 9,490 | 18,980 |
40% |
10% | 44% | 2,387 | 4,774 |
40% |
20% | 48% | 601 | 1,202 |
40% |
30% | 52% | 267 | 534 |
50% |
5% | 52.5% | 6,272 | 12,544 |
50% |
10% | 55% | 1,562 | 3,124 |
50% |
20% | 60% | 385 | 770 |
50% |
30% | 65% | 167 | 334 |
Why sample size math matters for App Store & Play Store tests
Both Apple's Product Page Optimization and Google Play's Store Listing Experiments will happily show you a "winner" the moment one variant is numerically ahead — neither platform stops you from reading a result before it's statistically meaningful. A test that ends after 200 visitors per variant on a 25% baseline is very likely reading noise, not a real difference. Run the numbers first, then trust the result.
This complements the sizing question with the visual one: see what to actually test first if you haven't started yet.
Download the worked-example dataset
Embed this calculator on your site
Includes a canonical link back to this page. Theme query: ?theme=light|dark|auto.
Localize the screenshots you're about to test.
Shotlingo localizes your App Store screenshots into 40+ languages so your winning variant ships everywhere, not just your home market. Free tier, no credit card.
Start free What to test first →FAQ
What is "minimum detectable effect" (MDE)?
The smallest relative improvement in conversion rate you actually want the test to be able to notice. A 10% MDE on a 25% baseline means detecting a move to roughly 27.5%. Chasing a smaller MDE needs a much bigger sample — sample size grows roughly with the inverse square of the effect you are trying to detect.
Why does a lower baseline conversion rate need a bigger sample?
Sample size math is driven by the variance of a proportion, p(1−p), which peaks at 50% and shrinks toward the extremes — but a low baseline also means a fixed relative lift (say 10%) is a smaller absolute gap in percentage points, and it is the absolute gap that has to clear the noise. In practice, low-baseline product pages need more visitors per variant to detect the same relative lift than high-baseline ones do.
What confidence level and power does this calculator use?
95% two-sided significance and 80% power — the two most common defaults across public A/B testing calculators. Two-sided means the test can catch the treatment doing either better or worse than control, which is the safer default when you have not already ruled out the treatment underperforming.
How long should I run an App Store or Google Play screenshot test?
At least one full week regardless of how fast you hit the sample size below — daily conversion rate swings by day of week (weekday vs. weekend app-browsing behavior), and a 3-day test skews toward whichever days it happened to cover. Do not stop early just because a mid-week check crosses significance; that inflates your false-positive rate.
Does this apply to Apple's Product Page Optimization or Google Play's Store Listing Experiments?
Yes — the math is platform-agnostic. Both platforms' native experiment tools handle the random traffic split and report the observed conversion rates for you; this calculator answers the question neither tool answers up front, which is how many product page visitors you need before the result is trustworthy rather than noise.
What if my daily traffic is too low to hit the sample size in a reasonable time?
Two honest options: test a bigger change (a larger expected lift needs a smaller sample), or accept a larger MDE and only expect the test to catch big wins, not small ones. Underpowering a test to make it finish faster just means you will "detect" noise as often as you detect real effects.
Methodology & changelog
- 2026-09-12: Initial publication. Standard unpooled two-proportion z-test, 95% two-sided significance / 80% power fixed defaults.