Skip to main content
Back to Blog
Industry

The $2.6B problem: where late-stage trial attrition actually comes from

The often-cited figure for average drug development cost traces to a Tufts Center for the Study of Drug Development analysis that has been updated and debated periodically since its first version. Whether the exact number is $2.6 billion or somewhat higher or lower in any given calculation, the underlying structure of the cost is not in serious dispute: the single largest driver of total development cost is the capital tied up in programs that fail after Phase II, including the cost of capital during the years those programs ran before failing.

What gets less attention is where, within that attrition pattern, the failures concentrate by failure mode. Not all Phase II or Phase III failures are the same kind of failure. The post-mortems differ, and understanding those differences is necessary for identifying which part of the development process a technology like virtual patient modeling can actually address.

Breaking down Phase II failure by proximate cause

Academic analyses of clinical trial failures have categorized Phase II attrition into broad cause categories for decades. The usual categories are: lack of efficacy, unacceptable safety profile, strategic deprioritization, and operational issues. Lack of efficacy consistently accounts for the plurality of failures, typically in the range of 50 to 60% depending on the therapeutic area and time period analyzed.

But "lack of efficacy" as a failure category is not a root cause. It is a readout. A trial fails to show efficacy because one or more of the following conditions held: the compound did not hit its target at the doses administered, the target engagement did not produce the expected downstream pharmacological effect, the pharmacological effect did not translate to the clinical endpoint measured, or the enrolled patient population diluted the effect below the threshold the trial was powered to detect.

The first three are classical pharmacology failures. The fourth is a population design failure. They are distinct, and only one of them is potentially addressable by better pre-trial characterization of the patient population.

The population design failure category and its frequency

Estimating how often a Phase II efficacy failure was actually a population design failure is methodologically difficult. You cannot run the same trial in a better-enriched population and compare results except in rare cases where the same compound was subsequently tested in a molecularly defined subgroup after an initial negative result. Those cases exist and are published, but they are not representative of all population design failures because the programs that look unpromising after a failed broad-population trial rarely get the resources to run a follow-up enriched trial.

Working from the published cases where enriched follow-up trials have been conducted, and from the theoretical framework of signal dilution in heterogeneous populations, the fraction of efficacy failures attributable to population design rather than compound pharmacology is likely to be meaningfully large in oncology, CNS, and metabolic disease indications. These are disease areas where patient populations are molecularly and clinically heterogeneous, where the plausible responder subgroup is a minority of the eligible population under broad criteria, and where enrichment handles are often available but underused at the time of Phase II design.

We think the actual number is higher than it appears in published analyses, for the simple reason that programs that fail Phase II are rarely subjected to rigorous post-mortem investigation that would distinguish population from pharmacology failures. The compound is shelved, the team moves to the next program, and the failure is recorded in the database as "lack of efficacy" without further decomposition.

Why the population failure mode is underdiagnosed

Drug development organizations are structured around compounds, not populations. A development team owns a development candidate from IND through Phase III. The expertise that evaluates that candidate, and the institutional memory of the development decision, is compound-specific. The translational medicine team understands the compound's mechanism, the PK/PD model, and the biomarker strategy. They may not have deep expertise in the population-level epidemiology of the target indication, and they typically do not have access to population simulation tools that would let them examine the conditional structure of the eligible population before committing to eligibility criteria.

The result is that eligibility criteria in Phase II protocols are often set by anchoring on prior trials in the same indication, adjusting for what the compound-specific PK/PD data suggests about the relevant patient characteristics, and accepting the aggregate response rate in the reference population as the appropriate planning assumption. None of these steps explicitly model what happens to the effective responder fraction when the eligibility criteria interact with the joint distribution of patient characteristics in the reference population.

This is not negligence. It is a rational response to working with the information and tools that are actually available. The problem is that the information and tools available have had a systematic blind spot: they characterize the aggregate population well and the conditional structure of that population poorly.

What happens to the cost structure if population failures are addressable

The $2.6 billion figure is driven primarily by the cost of failure late in development. Early failures are relatively inexpensive in absolute terms. The expensive failures are the ones that reach Phase III with sufficient evidence to justify the investment, enrol hundreds of patients across many clinical sites over multiple years, and then read out negative. Those are the failures where hundreds of millions of dollars in direct trial costs plus the opportunity cost of the clinical team's time and the compound's patent life are all consumed before the negative result is known.

If a meaningful fraction of those late-stage failures started as population design problems, and if those population design problems are detectable before Phase II enrolment begins rather than only after Phase III readout, then the intervention window exists. Not to guarantee success, but to change the information available at the Phase I to Phase II transition, where the decisions that determine population composition are still open.

That is the economic case for pre-IND and pre-Phase II population modeling. We are not claiming it eliminates late-stage attrition. There will always be compound failures, safety failures, and strategic deprioritizations that no amount of population modeling can prevent. The question is whether shifting the probability distribution on population design failures, specifically by surfacing enrichment scenarios and signal dilution risks before protocols are locked, is worth the cost of the analysis. For development programs in indications where population heterogeneity is high and the consequence of a failed Phase III is 9 to 12 digits, the expected value calculation is not complicated.

Explore the platform

The science described here is the basis for the platform's inference architecture. If you are working on a Phase I-II decision, request early access.

Request access Learn how it works