How to Use Statistical Calculators Correctly (Sample Size, Power & Effect Size)
Calculators are only as good as their inputs. This guide covers the three that matter most — sample size, power, and effect size — and the assumptions behind each.
Compute sample size before you collect, not after
An a priori power analysis tells you how many participants you need to detect an effect of a given size with acceptable confidence. It takes four ingredients, any three of which determine the fourth: the effect size you expect, the significance level (usually α = .05), the power you want (usually .80), and the sample size. To plan a study you fix effect size, α and power, and solve for N.
Rough anchors for 80% power at α = .05: detecting a medium correlation needs about 84 participants; a medium between-groups difference (two groups) about 128 total; mediation typically 150–250; complex models more. These are floors, not targets — see the inflation factors below.
Where does the expected effect size come from?
This is the input people guess at. Three legitimate sources, in order of preference: a meta-analysis or prior studies of the same effect; a pilot study; or, last resort, a field convention (small/medium/large). Never reverse-engineer the effect size from the sample you can afford — that guarantees an under-powered study dressed up as a planned one.
Match the effect size to the design
- Cohen's d — standardised mean difference, for t-tests.
- η² or partial η² — variance explained, for ANOVA.
- r — correlation strength, for associations.
- f² — for regression and multiple predictors.
- Odds ratio — for logistic regression.
Using the wrong family (a d where you needed an f²) produces a confident but irrelevant sample-size estimate.
Inflate for the real world
The textbook N assumes perfect data. Reality erodes it. Inflate your target for expected dropout (longitudinal designs can lose 30–50%), careless responding and failed attention checks (5–15%), and incomplete cases. If a power analysis says 200 and you expect 20% unusable responses, recruit 250.
Reliability: compute it, then read it
Cronbach's α is the common internal-consistency index; ≥ .70 is the conventional floor, .80+ is comfortable. But α rises mechanically with more items, so a high α on a 20-item scale is less impressive than on a 4-item one. For confirmatory models, also report composite reliability and AVE. A low α usually means a mis-scored reverse item, a multidimensional scale treated as one, or genuinely poor items.
The mistakes that make calculator output meaningless
- Post-hoc "observed" power computed from your own non-significant result — it is circular and uninformative.
- Plugging in a hoped-for large effect to justify a small sample.
- Ignoring the design — using independent-groups formulas for paired or clustered data.
- Forgetting multiple testing — many comparisons inflate false positives; adjust α or plan for it.
Final thought
The hard part of a calculation is choosing what to calculate and with what assumptions; the arithmetic is trivial. Use a decision navigator to pin the right test and effect-size family first, feed it honest inputs, and inflate for the real world. The Calculators suite does the arithmetic — your judgement supplies the inputs that make it true.
Plan your study's numbers
Use the research decision navigator to find the right test, then the matching calculator for sample size, power, effect size or reliability — 100 calculators, free.