"We surveyed 1,000 Americans" is a phrase you've seen hundreds of times in news articles. But why 1,000? Why not 100, or 10,000? The answer comes from a set of mathematical formulas that balance the cost of collecting data against the precision you need from your results. Understanding sample size determination means understanding exactly what you're trading off—and why the size of the total population matters much less than most people assume.
The Four Inputs to Sample Size
Every sample size calculation requires four inputs:
- Confidence level (α): How often, across many repeated surveys, would your interval capture the true population value? 95% is standard; 90% gives a smaller sample (less certainty) and 99% requires a larger one (more certainty).
- Margin of error (E): How precise do you need the result? A ±3% margin is standard for political polling; a ±1% margin requires roughly nine times as many respondents.
- Population proportion (p): Your best guess at the true proportion who hold a particular opinion or characteristic. When uncertain, use p = 0.5 (50%), which produces the largest—most conservative—sample size.
- Population size (N): The total population you're sampling from. For large populations, this barely matters; for small populations (a single company's 200 employees), it can significantly reduce the required sample.
The Formula and How to Read It
The standard sample size formula for large populations is: n = (Z² × p × (1−p)) / E², where Z is the z-score corresponding to your confidence level (1.96 for 95%, 1.645 for 90%, 2.576 for 99%), p is the estimated proportion, and E is the margin of error as a decimal.
Example: You want to estimate what proportion of a city's residents support a policy proposal, with 95% confidence and a ±3% margin of error. Using p = 0.5 (most conservative): n = (1.96² × 0.5 × 0.5) / 0.03² = (3.8416 × 0.25) / 0.0009 = 0.9604 / 0.0009 ≈ 1,068 respondents. This is why political polls so often use samples of around 1,000—that's the mathematically derived minimum for 95% confidence with ±3% margin. Use our Sample Size Calculator to compute this for any inputs.
Why Large Populations Need Surprisingly Modest Samples
Here is the most counterintuitive result in survey statistics: whether you're sampling from a city of 100,000 or a country of 330 million, the required sample size for a given precision is nearly the same. The formula above has no population size in it at all.
This seems wrong, but the logic is sound. The precision of a sample depends on the absolute sample size, not the fraction of the population sampled. Sampling 1,000 people from a city of 100,000 (1% sample) gives the same statistical precision as sampling 1,000 people from a nation of 300 million (0.0003% sample). What matters is that 1,000 independent, randomly selected responses are enough to produce a stable estimate.
The caveat is the finite population correction (FPC), which does apply when your sample is a large fraction of the total population—typically over 5%. The corrected formula is: n_adjusted = n / (1 + (n−1)/N). If a school has 200 students and your formula says you need 132, the FPC adjusts this to about 82—you don't need to survey 66% of the school, because the small population itself constrains the variance. For populations of tens of thousands or more, the FPC correction is negligible.
Choosing a Confidence Level: 90% vs. 95% vs. 99%
The confidence level is the probability that your confidence interval contains the true population value across repeated sampling. A 95% confidence level means that if you ran the same survey 100 times, about 95 of the resulting intervals would contain the true proportion.
- 90% confidence (Z = 1.645): Smaller sample needed. Acceptable for informal research, preliminary studies, or when data collection is very costly. The tradeoff is a higher error rate—1 in 10 intervals misses the truth.
- 95% confidence (Z = 1.96): The industry standard for political polling, market research, and most academic studies. Balances precision against feasibility.
- 99% confidence (Z = 2.576): Required for high-stakes decisions—clinical trials, product safety, regulatory submissions. Requires roughly 75% more respondents than a 95% study for the same margin of error.
Once you have your interval, our Confidence Interval Calculator can compute the actual bounds around your observed result.
Sensitivity: How Each Input Affects Required Sample Size
Understanding sensitivity helps you make smart tradeoffs:
- Halving the margin of error quadruples the sample size. Going from ±4% to ±2% precision doesn't double the sample—it multiplies it by 4. Margin of error appears squared in the denominator, so it's the highest-leverage input.
- Increasing confidence level from 95% to 99% increases sample size by about 75%. (2.576/1.96)² ≈ 1.73.
- Using p = 0.5 vs. p = 0.1 cuts the sample nearly in half. p(1−p) is maximized at 0.5 (value: 0.25) and falls to 0.09 at p = 0.1. If you have reason to believe the true proportion is far from 50%, you can use that estimate to reduce the required sample.
Real-World Applications
Election Polling
A national poll with 1,068 respondents produces a ±3% margin at 95% confidence. When a race is projected at 48% vs. 52%, the 4-point gap exceeds the 3% margin on each side—but only barely. This is why pollsters describe races within the margin as "too close to call": the statistical uncertainty spans the entire gap between the candidates.
A/B Testing in Product Development
When testing two versions of a webpage, you need enough users in each group to detect a meaningful difference in conversion rates. If your baseline conversion is 5% and you want to detect a 1 percentage-point improvement (to 6%), you'll need thousands of users per variant—because small proportions have high relative variance. Underpowered A/B tests fail to detect real improvements and waste development time. Use our Probability Calculator to think through the likelihood of various outcomes.
Quality Control Sampling
Manufacturers use sample size formulas to determine how many units to inspect from a production run. A factory producing 50,000 bolts per day doesn't test every bolt—it tests a calculated sample. If the acceptable defect rate is 1%, sampling 600 units at 95% confidence provides enough statistical power to detect whether the true defect rate exceeds the threshold.
Sample size mathematics turns the seemingly subjective question of "how much data is enough" into a precise engineering problem. Once you understand the four inputs and how they interact, you can design surveys, experiments, and quality audits that deliver the precision you need at the minimum possible cost.
