Search Russell Research

Find a solution, case, or article

Enter at least two characters to search the site.

    Press Esc to close. Use Ctrl/⌘ K to search from anywhere.
    Research planning tools

    Detectable difference and statistical power

    Could your study detect the difference that matters?

    Compare the sample you can recruit with the difference you need to detect. Plan separate audiences or test cells using percentages or average scores.

    Start with the comparison you intend to make.

    Use assumptions about the true population difference before collecting data. The opening example uses 400 usable people per group and a 50% baseline. Add your own meaningful difference to assess whether that sample supports your decision.

    View result

    Start with your own inputs, or try a worked example.

    01 What will you compare?

    Define the same measure in both groups, such as definitely would buy, top-two purchase intent, or average appeal.

    Check the research question and analysis plan

    Included in your download. Inputs are processed in your browser.

    02 Describe the difference that matters

    50% is an editable example, not always the most conservative assumption.

    The two-sided test can detect either direction. This choice sets the meaningful difference and the curve.

    The smallest true difference that could change your decision. Enter your threshold; the calculator will not choose it for you.

    03 Set the usable sample and detection target

    Count people with valid answers to this measure in each group. Overall survey completes, invitations, and repeated answers are different quantities.

    Chance of detecting the stated true difference under the model. 80% is a convention; 90% is shown alongside it.

    False-positive threshold under no true difference. Applies to the declared family if using Bonferroni.

    Test direction, multiple comparisons, and weighting

    1 = one primary comparison. More than 1 applies Bonferroni. Four cells mean six all-pairs comparisons, or three alternatives versus one control.

    An entered design effect adjusts the effective base for sensitivity only. It cannot replace a clustered design, recover representativeness, or verify the actual survey weights.

    04 Check uncertain assumptions

    Model-based planning estimate

    Smallest increase at 80% power

    9.84percentage points
    Group A400usable people
    Group B400usable people

    Group A baseline: 50%. Two-sided test; 5% significance threshold per comparison; 80% target power.

    Detectable increase9.8450% → 59.84%
    Detectable decrease9.8450% → 40.16%

    Enter a meaningful difference to compare this sample with your business requirement and calculate the sample needed.

    Two-proportion normal approximation, with pooled null and separate alternative variances; no continuity correction. Method and scope.

    This estimates detection of an audience difference. It does not establish what caused the difference. Selection, coverage, and measurement bias are not quantified.

    Compare detection targets

    Detectable increase at your entered sample sizes, in percentage points.

    80% power9.84percentage points
    90% power11.36percentage points

    80% and 90% are alternatives, not universal requirements. Higher power usually needs more usable data.

    Detection becomes more likely as the true difference grows.

    Power for a true increase, using 400 people in A and 400 in B. The curve uses the same method as your result.

    Power by true differenceThe horizontal axis is the true difference in percentage points. The vertical axis is the probability of a statistically significant result. Use the controls below for exact values.0%50%80%80%90%90%100%03.77.411.114.8
    7.382 percentage points55.3% power

    The slider explores assumptions; it does not change your meaningful difference. An 80% result means about 80 detections in 100 comparable studies under this model.

    View curve values as a table
    True increase and estimated power
    Difference (percentage points)Power
    05%
    1.8458.2%
    3.69118.1%
    5.53634.8%
    7.38255.3%
    9.22774.7%
    11.07388.5%
    12.91895.9%
    14.76498.9%
    Where would 100 more respondents help?
    Detectable increase (percentage points)
    AllocationA + BDifference
    Current400 + 4009.84
    Add 100 to A500 + 4009.35
    Add 100 to B400 + 5009.34

    These scenarios change available bases only. Recruitment costs and group availability also matter.

    Methods and your plan

    The download records your assumptions, rounded usable bases, achieved power, and scope. These are planning calculations, not observed results.

    Interactive calculator loading.

    Keep three different questions separate.

    What difference matters?

    A five-point improvement might change an investment decision even when the planned study has little chance of detecting it. Choose the business threshold from the decision and evidence, then assess the sample against it.

    What can this sample detect?

    A minimum detectable difference is tied to a stated probability of detection. At 80% power, the model still allows about one in five comparable studies to miss that true difference. It is not a hard significance boundary.

    What did the study find?

    Observed results need an appropriate test, estimates, and uncertainty intervals. Power calculated from the same observed difference does not resolve an inconclusive result. This tool plans future comparisons.

    Match the calculation to the study.

    A subgroup's usable base may be much smaller than the full survey. Repeated respondents, rare outcomes, and multiple comparisons also change what a study can establish.

    Discuss your sample and comparisons
    Why is a margin of error different?

    With 400 independent observations and percentages near 50%, a simple 95% margin of error is about ±4.9 points for one percentage. For the difference between two groups of 400, it is about ±6.9 points. Detecting an increase from 50% with 80% power needs about 9.84 points under this planner's model.

    These answer different questions: precision around an estimate and the chance of detecting a specified true difference. They are not interchangeable sample-size requirements. Plan precision for a single percentage.

    What if the same people answer twice?

    The responses are paired. For means, variability in the within-person differences matters. For percentages, the shares changing in each direction matter. Counting each person in two independent groups gives the wrong model.

    Partial overlap also needs the covariance between groups. The calculator withholds independent-group results when you select these designs. Mean-comparison assumptions.

    Does 80% power prove an improvement exceeds the business threshold?

    No. This planner tests against zero. Power to detect a true five-point difference does not mean power to establish that the difference exceeds five points. That needs a different null boundary and a plausible true effect beyond it. Likewise, showing similarity requires equivalence bounds; a nonsignificant difference does not establish equivalence.

    How should we plan several audiences or concepts?

    Specify the comparisons that will support the decision. Four cells produce six possible pairs, or three comparisons if each alternative is tested only against one control. The Bonferroni option allocates the family significance threshold across that declared count.

    The power result applies to the pair entered here. It is not the probability of detecting every effect, and this pair's total is not the whole study budget when comparisons share a control. Evaluate the actual contrasts and reuse shared groups in the recruitment plan. Plan concept and message allocation. Multiple-testing methods.

    Can a larger sample correct an unrepresentative one?

    It can reduce random uncertainty under a model, but it does not correct coverage, selection, measurement, or missing-data bias. Random assignment helps evaluate treatments within the study; it does not automatically make participants representative of the wider market. Report the recruitment and analysis assumptions alongside any precision claims. AAPOR disclosure guidance.

    Methods and research

    Named calculations. Explicit assumptions.

    Percentages: a large-sample two-proportion calculation with a pooled variance under no difference and separate variances under the assumed alternative. Both rejection tails are included for a two-sided test. There is no continuity correction. Expected yes/no counts below 10 trigger a review note; they need a suitable finite-sample test or simulation.

    Means: a Welch-Satterthwaite noncentral-t power approximation using each group's assumed standard deviation. It includes a t critical value and uncertainty in the variance estimate, rather than substituting a normal cutoff. It remains an approximation to Welch's test; skew, outliers, and complex survey designs need further assessment.

    Required sample sizes are solved numerically, rounded to usable people, and checked again for achieved power. Searches require an effective base of at least 10 and stop at 1,000,000 usable people per group. A failed search reports that limit. Optional design effects reduce the effective bases for approximate sensitivity; they do not validate weighting or replace cluster-specific degrees of freedom.

    The distribution calculations are checked against SciPy and Statsmodels under matching settings. The sources below were reviewed September 29, 2026.

    If you already have respondent weights, the weighting diagnostic shows their dispersion and concentration. Its weight-only effect is not automatically a valid design effect for this power calculation.

    1. Lakens (2022). Sample Size Justification. Collabra: Psychology.

      Distinguishes power-based planning, sensitivity analysis, and the smallest effect that matters. A budget alone does not define a meaningful difference.

    2. Statsmodels. Power for two independent proportions.

      Documents the pooled variance under the null and separate variances under the alternative used by this planner. Matching numerical settings are essential when comparing tools.

    3. NCSS. Two-Sample T-Tests Allowing Unequal Variance.

      Explains the unequal-variance mean comparison and its distributional assumptions. Our named noncentral-t planning approximation is not a claim to reproduce every PASS option or an exact Welch power integral.

    4. SciPy. Noncentral Student’s t distribution.

      Defines the distribution used in the mean-power approximation. Independent SciPy calculations check the browser engine's rejection probabilities.

    5. Cook and colleagues (2018). DELTA² guidance on choosing the target difference. Trials.

      Supports justifying the target difference and testing sensitivity to uncertain baseline rates and variability. Its planning principles are adapted here; clinical thresholds are not imported into market research.

    6. R documentation. Adjust P-values for Multiple Comparisons.

      Distinguishes Bonferroni from other multiple-testing methods. This planner divides the declared family threshold by the number of planned comparisons.

    7. AAPOR. Transparency Initiative.

      Precision claims need disclosed methods and assumptions. Numerical power does not measure selection bias or establish that an opt-in sample represents a population.