Sample Size | 8 min read | Intermediate
Expected Sample Size in Two-Stage Trials
Learn how expected sample size is calculated in Simon's two-stage design, how it relates to PET, and why it can be lower than the maximum sample size.
Maximum sample size is only one part of planning
A Simon two-stage design does not have just one meaningful sample-size number. The maximum sample size, n, is the largest number of patients that can be enrolled if the trial continues through both stages.
This number is important for budgeting, recruitment planning, site capacity, and worst-case timeline estimates.
But a two-stage trial may stop early for futility after stage 1. When that happens, the stage 2 patients are never enrolled. The actual number of patients enrolled therefore varies from trial to trial.
Expected sample size provides another way to describe this behavior: it gives the average number of patients enrolled under a specified response probability.
What is expected sample size?
Expected sample size is the average number of patients that would be enrolled if the same two-stage design were repeated many times under a particular response probability, p.
It is important to distinguish this from the maximum sample size. Maximum sample size describes the upper bound on enrollment. Expected sample size describes average enrollment under a specified scenario.
For example, a design with a maximum sample size of 40 does not necessarily enroll 40 patients in every trial. Some trials may stop after stage 1, while others continue to stage 2 and enroll all 40 patients.
The expected sample size summarizes these different possible outcomes as an average.
How expected sample size is calculated
Let n1 be the number of patients enrolled in stage 1, and let n be the maximum total sample size. The number of patients enrolled in stage 2 is therefore n − n1.
Let PET(p) denote the probability that the trial stops for futility after stage 1 when the true response probability is p.
The expected sample size at p is:
The calculation follows directly from the two-stage structure. Every trial enrolls n1 patients in stage 1. Only the trials that continue to stage 2 enroll the additional n − n1 patients.
The probability of continuing to stage 2 is 1 − PET(p). Therefore, the expected number of stage 2 patients enrolled is (n − n1) × [1 − PET(p)].
The relationship between PET and expected sample size
PET is the probability that the trial stops after stage 1 for futility. A higher PET means that the design is more likely to stop early when the treatment is not promising.
PET therefore has a direct effect on expected sample size. For a fixed two-stage design, a higher PET means that fewer trials continue to stage 2, reducing the expected number of patients enrolled.
Conversely, when PET is low, most trials continue to stage 2 and the expected sample size is closer to the maximum sample size.
For a detailed explanation of how PET is defined, calculated, and interpreted, see the article on probability of early termination.
A simple example
Consider a Simon two-stage design with 20 patients in stage 1 and a maximum total sample size of 40 patients.
The stage 2 sample size is therefore 20 patients.
Suppose that at p0, the probability of early termination is 0.60. The probability of continuing to stage 2 is therefore 0.40.
The expected sample size at p0 is:
The expected sample size at p₀ is 20 + 20 × (1 − 0.60) = 28.
So this design has a maximum sample size of 40 patients but an expected sample size of 28 patients at p0.
This does not mean that the trial can be planned to enroll only 28 patients. The trial still has a maximum enrollment of 40 patients. Some trials will stop after 20 patients, while trials that continue to stage 2 will enroll all 40 patients.
The value of 28 represents the average enrollment across repeated trials under the assumed response probability p0.
Expected sample size depends on the response probability
Expected sample size is not a single fixed number for a Simon two-stage design. It depends on the assumed response probability, p.
At p0, the treatment is assumed to have the null or uninteresting response rate. If the treatment is ineffective, early stopping for futility is an important possibility. Expected sample size at p0 therefore helps describe how much enrollment the design typically requires under the null scenario.
At higher response probabilities, the trial is generally more likely to continue to stage 2. As a result, expected sample size tends to increase toward the maximum sample size.
This means that expected sample size can be viewed as a function of the assumed response probability rather than simply as another fixed sample-size value.
For the same design, expected sample size may therefore be quite different under p0, p1, and response probabilities between them.
Expected sample size at p0 and p1
The values p0 and p1 have different roles in Simon's design. p0 represents the response probability considered uninteresting, while p1 represents the response probability considered promising.
Expected sample size at p0 is particularly useful for assessing the enrollment burden when the treatment does not achieve a clinically meaningful response rate.
Expected sample size at p1 describes enrollment when the treatment performs at the targeted promising response rate. Because the probability of continuing to stage 2 is typically higher at p1, expected sample size at p1 will often be closer to the maximum sample size.
Looking at expected sample size at both p0 and p1 can therefore provide a more complete picture of how the design behaves under unfavorable and favorable scenarios.
Maximum vs expected sample size
Maximum sample size and expected sample size answer different planning questions.
Maximum sample size asks: how many patients could the trial enroll if it proceeds through both stages?
Expected sample size asks: how many patients would the trial enroll on average under a specified response probability?
The distinction matters because an expected sample size should not be interpreted as a replacement for the maximum sample size in trial planning.
The maximum sample size remains relevant for understanding the largest possible enrollment requirement. Expected sample size provides additional information about the average enrollment burden under a particular scenario.
Using expected sample size to compare designs
When comparing Simon two-stage designs, expected sample size should be considered together with maximum sample size, PET, type I error, and power. No single number captures the complete operating profile of a design.
For example, one design may have a smaller maximum sample size, while another may have a higher probability of early termination and therefore a lower expected sample size at p0.
A design with a smaller maximum sample size is not automatically preferable. It may have a larger expected sample size under p0 if its probability of early termination is lower.
Conversely, a design with a somewhat larger maximum sample size may provide greater efficiency under the null scenario if it stops early more often.
Which design is preferable depends on the priorities of the study, including recruitment feasibility, maximum enrollment, expected enrollment when the treatment is ineffective, and the desired statistical operating characteristics.
This is why Simon design selection is better treated as a comparison of operating characteristics rather than a search for a single 'best' sample size.
BioStatHub's Designs workspace is intended to make these trade-offs visible: review candidate designs in the Overview, then compare them in the Compare view.
Expected sample size and the power function
Expected sample size and power describe different aspects of a two-stage design.
Expected sample size describes how many patients the design is expected to enroll under a specified response probability. Power describes the probability that the design will successfully reject the null hypothesis when the true response probability is a specified value.
Together, these operating characteristics help show the balance between statistical performance and enrollment burden.
The power function provides a broader view of how the probability of success changes across possible response probabilities.
How to interpret expected sample size in practice
When reviewing a candidate Simon design, start with the maximum sample size to understand the largest possible enrollment requirement.
Then consider PET at p0 to understand how often the trial is expected to stop early when the treatment is ineffective.
Finally, consider expected sample size at p0 to quantify the average enrollment burden under that scenario.
These quantities answer complementary questions: maximum sample size describes the upper bound, PET describes the chance of early stopping, and expected sample size summarizes average enrollment.
Considering them together provides a more informative view of the design than looking at maximum sample size alone.
Key takeaways
Maximum sample size is the number of patients enrolled if the trial continues through both stages.
Expected sample size is the average number of patients enrolled under a specified response probability.
For a Simon two-stage design, expected sample size can be calculated as n1 + (n − n1) × [1 − PET(p)].
PET and expected sample size are directly related: a higher probability of early termination leads to a lower expected sample size for a fixed design.
Expected sample size at p0 is particularly useful for understanding enrollment efficiency when the treatment is ineffective.
Expected sample size should be considered together with maximum sample size, PET, type I error, and power when comparing candidate designs.
Frequently Asked Questions
What is expected sample size in a Simon two-stage design?
Expected sample size is the average number of patients the trial is expected to enroll under a specified response probability. It accounts for the possibility that the trial stops early for futility after stage 1.
How is expected sample size calculated in a Simon two-stage design?
For a Simon two-stage design, expected sample size at response probability p is n1 + (n − n1) × [1 − PET(p)], where n1 is the stage 1 sample size, n is the maximum sample size, and PET(p) is the probability of stopping after stage 1.
Why can expected sample size be lower than the maximum sample size?
Because some trials stop after stage 1 for futility. Those trials do not enroll stage 2 patients, so the average enrollment across repeated trials is lower than the maximum sample size.
Is expected sample size the same as the planned sample size?
No. Maximum sample size is the number of patients enrolled if the trial continues through both stages. Expected sample size is an average that depends on the assumed response probability.
How are PET and expected sample size related?
They are directly related. A higher probability of early termination means fewer trials continue to stage 2, which reduces expected sample size.
Should expected sample size be evaluated at p0 or p1?
Expected sample size can be evaluated at any response probability. Evaluating it at p0 is particularly useful because p0 represents the null or uninteresting response rate and shows how much enrollment the design typically uses when the treatment is ineffective.