Why are conventional power calculations not appropriate for pilot studies, and what alternative method is suggested for estimating sample size in pilot work?
Conventional power calculations are not appropriate for pilot studies because pilot studies do not aim to test hypotheses about intervention outcomes. Instead, experts recommend using confidence intervals around expected feasibility estimates, such as attrition rates, to determine the sample size needed to establish whether a full-scale trial is feasible.
Sample size estimation in pilot work differs from that in a full trial because pilot studies are not designed to test hypotheses about intervention outcomes. Conventional power analysis relies on an effect size estimate for the primary outcome, but pilot effect size estimates are unreliable and can lead to both Type I and Type II errors if used to project the sample size for a full trial. For example, an inflated pilot effect size could result in an underpowered full trial, whereas an unduly small estimate could cause a promising intervention to be abandoned. Because the purpose of a pilot is to assess feasibility, experts such as Arnold, Hertzog, and Thabane have suggested using confidence intervals to estimate the sample size needed to establish feasibility. For instance, if a full-scale trial is deemed feasible only if attrition at 3 months is 20% or less, and the expected attrition is 12%, a 95% confidence interval approach would require a total sample of 64 participants so that the upper confidence limit does not exceed 20%. Relaxing the standard to a 90% confidence interval would reduce the needed sample to 46. Hertzog has also recommended at least 30 to 40 participants per group when pilot funding is being sought. Thus, pilot sample size is based on precision around feasibility criteria rather than on power for detecting outcome effects.
Key points
- Pilot studies are not intended to test hypotheses about intervention outcomes, so conventional power analysis is not appropriate.
- Using pilot effect size estimates for power calculations can lead to underpowered full trials or abandonment of promising interventions.
- Experts recommend using confidence intervals around projected feasibility estimates, such as attrition rates, to determine pilot sample size.
- If the expected attrition is 12% and the feasibility criterion is attrition no greater than 20%, a 95% CI approach suggests N = 64 while a 90% CI suggests N = 46.
- Hertzog recommends at least 30 to 40 participants per group for funded pilot work.
Related questions
Nursing research: generating and assessing evidence for nursing practice
Beck, Cheryl Tatano Polit, Denise F.
Tenth edition · Wolters Kluwer