Concept & message test planner
What should each person see—and what can the study tell you?
Compare test designs, respondent counts, exposure coverage, and survey time. Start with the decision, then build a study specification you can review with your team.
Plan the comparison before writing the questionnaire.
Use up to 12 candidates and eight disjoint, equally sized audience groups. Defaults are hypothetical examples. Independent binary and mean comparisons support a power estimate; repeated exposures, choice models, interactions, and historical benchmarks need further statistical planning.
Illustrative starting points. Replace the assumptions with your own.
Incomplete-block sequential. 402 usable people. Coverage only; statistical adequacy unresolved.
Suggested design · planning guidance
Incomplete-block sequential
A subset limits repeated work while covering the full list across people. Balance exposures and positions, then inspect pair coverage.
402usable people · coverage plan only
Coverage target only, rounded to complete exposure/position cycles.
- Arms in this study
- 6 candidates
- Exposures per person
- 3
- Evaluations per arm, overall
- 201
- People in each audience
- 402
- Evaluations per arm per audience
- 201
- Completed interviews before quality exclusions
- 447
Statistical adequacy is unresolved
You selected exposure coverage. No statistical sample requirement has been calculated.
- Multiple exposures or direct choices need a repeated-response or choice model. This release does not calculate their power.
Time from your assumptions
7.2–10.8 minutes per person
Viewing 1.5 · ratings 1.2 · open ends 1.5 · other 3 minutes.
Both scenarios fit your entered time allowance. These are arithmetic scenarios; pilot actual timing. Control arms use the candidate path for budgeting.
Compare an alternative
Monadic: 1,200 people at 1 exposures each for the same exposure target; 4.4–6.6 minutes.
Power is not transferred to this alternative. One stimulus per person protects isolated evaluation.
Inputs stay in this browser. Export includes assumptions, sources, and unresolved requirements.
Interpretation and review
Concept ratings can guide development and comparison. Stated purchase consideration is not a sales forecast.
- Evaluations from the same person are dependent. Fit a model that accounts for respondents and exposure order; do not treat the evaluation total as independent people.
Large samples do not remove selection, coverage, or measurement bias.
Calculated allocation template
See where each candidate appears.
Every complete cycle balances exposures and positions within each audience. The template uses 67 cycles per audience. Its pairings are measured, not optimized.
Pairs appear together 74–85 times overall. Equal marginal exposure does not guarantee equal pair coverage or statistical efficiency.
The CSV records how many people to assign to each ordered row. Randomize those assignments within audience in the survey platform, retain the within-row order, and replace losses within affected cells or blocks.
| Arm | Evaluations | Position 1 | Position 2 | Position 3 |
|---|---|---|---|---|
| A | 201 | 67 | 67 | 67 |
| B | 201 | 67 | 67 | 67 |
| C | 201 | 67 | 67 | 67 |
| D | 201 | 67 | 67 | 67 |
| E | 201 | 67 | 67 | 67 |
| F | 201 | 67 | 67 | 67 |
Pair co-exposure and audience counts
Each audience has 201 evaluations per arm, 67 in each position. Pair counts can differ between audiences; review the exported rows for an audience-specific design. Diagonal cells are not comparisons.
| Arm | A | B | C | D | E | F |
|---|---|---|---|---|---|---|
| A | — | 85 | 79 | 74 | 79 | 85 |
| B | 85 | — | 85 | 79 | 74 | 79 |
| C | 79 | 85 | — | 85 | 79 | 74 |
| D | 74 | 79 | 85 | — | 85 | 79 |
| E | 79 | 74 | 79 | 85 | — | 85 |
| F | 85 | 79 | 74 | 79 | 85 | — |
Draft questionnaire
Measure the decision first.
- Consent, eligibility, and only necessary pre-exposure measures.
- Assigned stimulus and the primary outcome: Purchase consideration (top two boxes).
- Choose diagnostics such as relevance, comprehension, credibility, uniqueness, and value when price is shown. Keep the entered 4 ratings and 1 open ends per stimulus consistent with the final questionnaire.
- Finish with classification questions needed for analysis.
Draft analysis and pilot
Check the plan before fieldwork.
Use a repeated-response model appropriate to the outcome, including respondent dependence and exposure order. Review stimulus effects, missing ratings, and pair coverage before selecting contrasts.
Specify the primary contrasts, exclusions, missing-response rules, and multiplicity treatment before reviewing results. The planner assumes complete usable ratings; partial records can disrupt balance.
Pretest comprehension and device rendering. Pilot the full survey to check duration, dropout, assignment, and position counts. The allocation download is a draft specification, not a validated production randomizer.
Discuss the study designChoosing a design
Isolated ratings, relative preference, and exposure effects answer different questions.
A design suggestion is a starting point for review. Stimulus quality, audience selection, the primary outcome, and implementation determine what a study can support.
Concept and product testingBrand positioning and messagingWhen does one stimulus per person help?
Monadic cells keep an evaluation separate from other candidates in the study. They are useful when first impressions, complex stimuli, or consistent benchmark conditions matter. Randomly allocate participants within each audience.
A concurrent control is another study arm. A historical benchmark also has measurement conditions and potentially sampling uncertainty; the planner does not treat it as a known constant.
What changes when people evaluate several candidates?
Sequential monadic testing collects a separate evaluation after each exposure. An incomplete block shows only a subset of the list to each person. Both can reduce the people required for exposure coverage.
Responses from one person remain dependent. Order, fatigue, and comparison with preceding candidates can matter. This release balances positions and displays pair coverage but does not calculate repeated-response power or optimize a block design.
Does the largest score establish a winner?
No. Specify the outcome, relevant comparisons, and a meaningful gap before fielding. For supported independent comparisons, the planner applies a two-sided test against zero with Bonferroni correction for the stated family.
The displayed power applies to a specified pair under common assumptions. It is not the probability of ranking every candidate correctly, detecting every effect, or proving a gain beyond a business threshold. Direct choice measures relative preference in the tested set and uses a different model.
How does a message-effect experiment differ from message ratings?
To estimate an exposure effect, randomly assign people to a candidate message or a no/current-message control. Measure the same outcome in every arm. A question such as correct understanding of an offer can apply to both groups.
Asking participants whether a message seems persuasive is a diagnostic rating. It does not by itself measure a change caused by exposure or predict real-market behavior.
Can the allocation CSV go straight into fieldwork?
It is a draft quota template. Each row gives an ordered set of arms and the number of usable participants to assign to it. Expand or implement those counts in your survey platform, randomize assignment within audience, and retain the specified within-row order.
Validate the programming and pilot losses, missing ratings, and timing. Exposure and position totals balance exactly for complete usable records; pair totals can differ. The template is not an optimized experimental design or a tested survey randomizer.
Methods and sources
Separate coverage, statistical assumptions, and judgment.
Coverage uses complete cyclic allocations. Binary power uses a normal approximation; mean power uses the shared Welch–Satterthwaite noncentral-t approximation. Timing and design suggestions are disclosed planning assumptions. Sources reviewed September 29, 2026.
- AAPOR: Best Practices for Survey Research
Guidance on objectives, question order, burden, pretesting, and piloting. It supplies no universal timing or per-concept sample rule.
- NIST: Bonferroni’s method
Family error control for prespecified comparisons. This planner divides alpha across the entered comparison family.
- Statsmodels: Power for two independent proportions
Documents the pooled-null and unpooled-alternative normal approximation used for binary outcomes.
- NIST: Two-sample t-test
Independent means and Welch–Satterthwaite degrees of freedom. Mean power reuses the site’s noncentral-t approximation.
- Kim & Cappella (2019): Reliable, valid and efficient evaluation of media messages
An original message-evaluation protocol that accounts for messages and respondents. Its health-communication evidence does not define a universal commercial message score.