Fundamentals | 10 min read | Beginner
How to Choose p₀ and p₁ in a Simon Two-Stage Design
Learn how to choose p₀ and p₁ for a Simon two-stage phase II trial, understand the effect size, the four-way trade-off with α and β, and common mistakes to avoid.
What is p₀?
Two numbers determine almost every aspect of a Simon two-stage design: sample size, stopping probability, and statistical power. Those numbers are p₀ and p₁.
For example, if the current therapy works in 20% of patients, investigators might only consider a new treatment worthwhile if it achieves 40%. In that case, p₀ = 0.20 and p₁ = 0.40.
Once p₀ and p₁ are chosen, α and β determine how confidently the design can distinguish between them. Together, these four parameters define the trial's operating characteristics.
p₀ is the response rate that would be considered clinically unacceptable. If the true response rate is at or below p₀, the treatment is not sufficiently effective to justify further development.
Think of p₀ as the benchmark for an ineffective treatment.
A note on choosing p₀: This value should be informed by historical data and the current standard of care. It represents the threshold below which you would not want to proceed with further development—not necessarily the lowest response rate you might observe.
For example, if existing therapies produce responses in about 20% of patients, investigators may decide that a new treatment achieving only a 20% response rate is not worth pursuing. In that case: p₀ = 0.20.
What is p₁?
p₁ is the minimum response rate that would make the treatment clinically promising. If the true response rate reaches or exceeds p₁, the treatment warrants further investigation in a larger confirmatory trial.
Continuing the previous example, suppose clinicians agree that the new treatment should achieve at least a 40% response rate to justify further development. Then: p₁ = 0.40.
A note on choosing p₁: p₁ should represent the smallest response rate that would be considered clinically meaningful. It should be ambitious enough to justify further development while remaining realistic based on existing evidence and clinical expectations. The goal is not to choose the highest possible response rate, but one that would genuinely change clinical decision-making if achieved.
Where Do These Numbers Come From?
The values should not be chosen because they produce a convenient sample size. Instead, they should be based on clinical judgment and existing evidence, including:
Historical response rates from previous studies.
Outcomes achieved by the current standard of care.
What clinicians consider to be a meaningful improvement for patients.
Regulatory guidance or protocol requirements.
Selecting p₀ and p₁ is typically a collaborative discussion between clinicians, statisticians, and other study stakeholders.
The Effect Size Matters
The difference between p₁ and p₀ is the effect size that your study aims to detect.
For example:
p₀ = 20%, p₁ = 40% → Effect size = 20 percentage points.
p₀ = 30%, p₁ = 35% → Effect size = 5 percentage points.
A smaller effect size is harder to distinguish statistically. Just as in comparative clinical trials, detecting a smaller effect requires more patients.
Conversely, a larger effect size is easier to detect and generally results in a smaller required sample size.
The Four-Way Trade-Off
Many investigators focus only on p₀ and p₁, but Simon's two-stage design is determined by four parameters:
p₀ – Unacceptable response rate.
p₁ – Promising response rate.
α – Probability of falsely declaring an ineffective treatment promising.
β (or equivalently, power = 1 − β) – Probability of failing to identify a truly promising treatment.
These four parameters work together. For a fixed p₀ and p₁: lowering α (being more conservative) generally increases the sample size. Increasing power (lowering β) also increases the sample size.
Likewise, for fixed α and β: bringing p₀ and p₁ closer together increases the sample size. Moving them farther apart generally decreases the sample size.
If the resulting design is too large, the question is not simply: 'Can we reduce the sample size?' Instead, ask: 'Is there another combination of p₀, p₁, α, and β that remains both clinically justifiable and statistically acceptable?'
Choosing p₀ and p₁ Is Often an Iterative Process
Although p₀ and p₁ should be driven by clinical considerations, they are rarely finalized in a single discussion.
A typical workflow is:
1. Choose clinically meaningful values for p₀ and p₁.
2. Select acceptable values for α and power.
3. Generate a Simon two-stage design.
4. Review the resulting maximum sample size, expected sample size, probability of early termination, and decision boundaries.
5. Decide whether the design is operationally feasible.
6. If necessary, revisit the clinical and statistical assumptions and generate a new design.
For example, suppose the initial assumptions are: p₀ = 0.20, p₁ = 0.40, α = 0.05, Power = 80%. The resulting design may require more patients than the study can realistically recruit.
The team then discusses several possibilities: Is a larger treatment effect required before considering the treatment promising? Can a slightly higher α be justified? Can slightly lower power be accepted? Do the original clinical assumptions still best reflect the study objectives?
After each adjustment, a new Simon two-stage design is generated and evaluated.
The goal is not to obtain the smallest possible sample size. Rather, it is to identify a design that remains clinically meaningful while being practical to conduct.
This iterative process is a healthy and expected part of designing a high-quality phase II trial.
In practice, many teams evaluate several candidate combinations of p₀, p₁, α, and β before selecting the final design. Interactive Simon two-stage design software can make this exploration much easier by allowing investigators to compare multiple designs side by side.
Common Mistakes
Some common pitfalls include:
Choosing p₀ solely to reduce the sample size.
Selecting an unrealistic p₁ simply to obtain a more favorable design.
Ignoring clinical input when defining the target response rates.
Treating p₀, p₁, α, and β as independent decisions rather than a set of interconnected design parameters.
Remember that Simon's two-stage design should answer a clinically meaningful question—not simply produce the smallest possible study.
Key Takeaways
p₀ is the highest response rate that would still be considered clinically unacceptable.
p₁ is the lowest response rate that would justify further clinical development.
The difference (p₁ − p₀) is the effect size. Smaller effect sizes require larger studies.
p₀, p₁, α, and β should always be chosen together, balancing clinical relevance, statistical rigor, and operational feasibility.
Choosing these parameters is usually an iterative process rather than a one-time decision.
Frequently Asked Questions
What values should I use for p₀ and p₁?
p₀ should reflect the response rate of the current standard of care or a rate that would be considered clinically unacceptable. p₁ should be the smallest response rate that would make the treatment clinically promising. These values should be based on clinical judgment and historical evidence, not chosen to produce a convenient sample size.
How does the effect size affect my design?
The effect size is the difference between p₁ and p₀. A smaller effect size is harder to detect and requires a larger sample size. A larger effect size is easier to detect and generally results in a smaller required sample size.
Can I change p₀ and p₁ after seeing the design?
Yes, choosing these parameters is often an iterative process. You may start with clinically meaningful values, generate a design, review the sample size, and then revisit your assumptions if the design is not operationally feasible. The goal is to find a combination that remains clinically meaningful while being practical to conduct.
What is the four-way trade-off?
Simon's two-stage design is determined by four interconnected parameters: p₀ (unacceptable response rate), p₁ (promising response rate), α (Type I error), and β (Type II error). Changing any one affects the others. For fixed p₀ and p₁, lowering α or increasing power increases sample size. For fixed α and β, bringing p₀ and p₁ closer together increases sample size.
What is a common mistake when choosing p₀ and p₁?
A common mistake is choosing p₀ solely to reduce the sample size, or selecting an unrealistic p₁ simply to obtain a more favorable design. These parameters should be driven by clinical considerations, not statistical convenience.
How do I know if my p₀ and p₁ are realistic?
Base them on historical response rates from previous studies, outcomes achieved by the current standard of care, what clinicians consider a meaningful improvement, and any relevant regulatory guidance. This is typically a collaborative discussion between clinicians, statisticians, and other study stakeholders.