Design Selection | 9 min read | Intermediate
Probability of Early Termination (PET) Explained: Why It Matters in Simon's Two-Stage Designs
Learn what probability of early termination (PET) is, how it is calculated, how to interpret it, and why it should be considered alongside expected sample size and power when selecting a Simon two-stage design. PET quantifies one of the key advantages of Simon's design: the ability to stop a trial early for futility when the treatment appears insufficiently promising.
What Is Probability of Early Termination?
Probability of Early Termination (PET) is the probability that a trial stops after Stage 1 because the observed results are insufficient to justify continuing to Stage 2.
Suppose a Simon design specifies a Stage 1 sample size of 15 patients and an early stopping boundary of 2 responses. After treating the first 15 patients, if 0, 1, or 2 responses are observed, the trial stops early for futility. If 3 or more responses are observed, the trial proceeds to Stage 2.
In the standard Simon two-stage design, early stopping is only for futility. Unlike some adaptive or group sequential designs, Simon's design does not allow a trial to stop early because the treatment appears highly effective. Promising treatments must still complete Stage 2 before efficacy can be concluded.
PET measures how likely that early stopping event is.
Why Was PET Introduced?
Traditional single-stage Phase II trials require enrolling every planned patient before determining whether a treatment is promising.
Richard Simon introduced the two-stage design to avoid unnecessarily exposing patients to ineffective treatments while preserving the desired statistical properties of the trial.
When an experimental therapy performs poorly during Stage 1, there is little value in continuing recruitment. Early termination allows investigators to protect patients from ineffective therapies, reduce trial costs, reach conclusions sooner, and redirect resources toward more promising treatments.
These advantages explain why Simon's two-stage design remains one of the most widely used designs in oncology and other early-phase clinical research.
PET Depends on the True Response Rate
A common misconception is that PET is a fixed characteristic of a design. It is not. PET depends on the true underlying response probability, usually denoted by p.
When researchers summarize a Simon design, they typically report PET at two clinically important response rates: PET under p0, the probability of stopping early when the treatment is truly ineffective, and PET under p1, the probability of stopping early when the treatment truly achieves the desired response rate.
These two values tell very different stories. A high PET under p0 is desirable because ineffective treatments are abandoned quickly, saving patients and resources. A low PET under p1 is desirable because promising treatments are unlikely to be stopped before sufficient evidence has been collected.
Good Simon designs balance these competing objectives while maintaining the desired Type I error rate and statistical power.
How Is PET Calculated?
PET is the probability of observing at most r₁ responses in Stage 1, resulting in early termination for futility.
If n₁ is the Stage 1 sample size, r₁ is the early stopping boundary, and p is the true response probability, then:
In other words, PET is the cumulative distribution function (CDF) of a binomial random variable evaluated at the Stage 1 stopping boundary.
Fortunately, modern statistical software computes PET automatically, so researchers rarely need to perform these calculations by hand.
A Numerical Example
Suppose a Simon design specifies a Stage 1 sample size of 15, an early stopping boundary of 2, and a null response rate of 20%. The trial stops if 2 or fewer responses are observed among the first 15 patients.
The PET under the null is therefore P(X ≤ 2), where X follows a Binomial(15, 0.20) distribution, which equals approximately 40%.
This means that if the treatment truly has only a 20% response rate, approximately 4 out of every 10 trials will stop after enrolling just the first 15 patients.
The remaining trials continue to Stage 2 because random variation occasionally produces more than two responses, even when the treatment is ineffective.
This does not mean those treatments are effective; it simply reflects the random variation inherent in a binomial outcome.
PET Is a Curve, Not Just a Number
Although PET is often reported only under p0 and p1, it is fundamentally a continuous function of the true response probability.
As the true response rate increases, early stopping becomes less likely, more trials continue to Stage 2, and PET decreases steadily. Conversely, when the true response rate decreases, PET increases because poor Stage 1 results become more common.
Viewing PET as a curve provides much richer insight than comparing only its values at p0 and p1.
PET and Expected Sample Size
PET has a direct influence on the Expected Sample Size (ESS). The higher the probability of stopping after Stage 1, the fewer patients are expected to be enrolled when the treatment is ineffective.
This relationship explains why PET and ESS are closely connected.
PET in Optimal and Minimax Designs
PET also helps explain one of the fundamental differences between Simon's two most commonly used design criteria.
Optimal designs minimize the expected sample size under the null hypothesis. Consequently, they often achieve relatively high PET values because many ineffective treatments stop after Stage 1.
Minimax designs, on the other hand, minimize the maximum total sample size. Achieving this objective often requires accepting a lower PET, allowing more studies to continue into Stage 2 before reaching a final decision.
Neither design is universally superior. An optimal design may be preferable when minimizing patient exposure to ineffective treatments is the primary concern, whereas a minimax design may be preferred when limiting the maximum trial size is more important.
Does Early Stopping Affect Type I Error?
A common question is whether allowing early stopping changes the significance level of the trial. The answer is no.
Although Simon's design includes an interim futility analysis after Stage 1, it also specifies a final rejection boundary after Stage 2. The Stage 1 and Stage 2 decision rules are jointly selected so that the overall Type I error rate remains at the desired level while preserving the target statistical power.
Early stopping therefore improves efficiency without compromising the operating characteristics specified during the design stage.
How BioStatHub Helps
BioStatHub automatically calculates the Probability of Early Termination (PET) for every candidate Simon design and presents it alongside the other key operating characteristics.
Rather than evaluating PET in isolation, you can compare probability of early termination, expected sample size, maximum sample size, statistical power, and type I error together.
Viewing these metrics together makes it easier to understand the trade-offs between competing designs and select the Simon two-stage design that best aligns with the objectives of your Phase II trial.
Key Takeaways
Probability of Early Termination is one of the defining features of Simon's two-stage design. It measures how frequently an ineffective treatment is expected to stop after the first stage, reducing patient exposure, conserving resources, and improving trial efficiency.
Although PET is often summarized by its values under the null and alternative response rates, it is fundamentally a function of the true response probability. Interpreting PET together with expected sample size, power, maximum sample size, and type I error provides a much more complete understanding of a design's operating characteristics.
Frequently Asked Questions
Is PET a fixed number for a given design?
No. PET depends on the true underlying response probability, p. It is usually reported at two clinically important values: under the null response rate p0 and under the target response rate p1.
Does early stopping change the type I error?
No. The stage 1 futility rule and the stage 2 final rejection boundary are selected together so that the overall type I error rate remains at the desired level while preserving the target power.
Why is PET higher under p0 than under p1?
Under p0 the treatment is truly ineffective, so poor stage 1 results are more common and early stopping is more likely. Under p1 the treatment truly achieves the target response rate, so promising stage 1 results are more common and early stopping is less likely.
Should I maximize PET?
Not necessarily. An overly aggressive stopping rule may terminate promising treatments too frequently and reduce statistical power. PET should be balanced with expected sample size, maximum sample size, power, and type I error.