Skip to main content
Back to Blog
Platform

Predicting trial failure by subgroup before enrolment: what virtual patient modeling makes possible

In the majority of oncology Phase II readouts that fail to meet primary endpoints, the biology was not the problem. The compound cleared its mechanistic hypothesis in a subset of enrolled patients. The trial failed because the enrolled population diluted the signal below threshold. We started Valinor Discovery because we kept observing this pattern in work with translational teams, and because we believed the information needed to avoid it was recoverable before enrolment began.

The conventional framing of Phase II attrition is compound-centric: the target was wrong, the biomarker hypothesis was wrong, the dose was wrong. That framing is accurate for a fraction of failures. But it directs resources toward compound-level explanation when the actual failure mode was population design. Those are different problems with different remedies.

The disconnect between aggregate statistics and enrolled populations

Every Phase II protocol is built on population assumptions. The development team uses published response rates, epidemiological prevalence data, and PK/PD modelling to set eligibility criteria and power calculations. That process is rational. The problem is that published population statistics are aggregated across the same heterogeneity that will later produce the signal dilution.

Published registry data for a disease population tells you what percentage of patients have a given comorbidity, the distribution of disease severity scores, and aggregate prior treatment patterns. What it does not tell you is the joint conditional distribution of those variables with genetic variants, imaging-derived biomarkers, and molecular subtype. Those conditional relationships are what determine which subgroup will respond and which will not.

When eligibility criteria are set from marginal statistics, the resulting enrolled population reflects the marginal distribution of the reference cohort, not the conditional distribution of likely responders. Responders and non-responders get enrolled in proportion to their prevalence in the reference population, not in proportion to their likely response to the compound. The protocol looks adequately powered on paper. In practice it is enrolling a mixed population whose signal density is lower than the power assumptions require.

What virtual patient cohort simulation produces

The output of a virtual patient simulation is a set of synthetic patient profiles whose joint covariate distribution matches the reference population. Not individual patient records, but probabilistic profiles derived from the multimodal data available for the indication: structured EHR cohort statistics, genomic variant frequency data, imaging-derived biomarker distributions, and natural history priors from the literature.

The development team can apply any candidate eligibility definition to that virtual cohort and observe what the filtered population looks like across dimensions not directly controlled by the eligibility criteria. If there is a subgroup within the filtered population whose covariate profile is likely to produce a different response pattern from the aggregate, the simulation surfaces that as a distribution of enrichment scenarios rather than a single point estimate.

This is different from subgroup analysis after trial readout. It is not data-mining a completed dataset. It is forward-looking inference about the population that a given protocol design will enroll, before any patients are enrolled. The distinction matters because the decisions that the information changes are only available before the trial begins.

What a pre-enrolment simulation changed in a real program

A Boston-area team working on a fibrosis-related indication was at the Phase I to Phase II decision point in late 2025. Their aggregate real-world evidence data showed a 59% responder rate in the reference population, drawn from a published registry with over 2,000 patients. Their Phase II design was built on that assumption with standard power parameters.

When we ran a virtual patient simulation against their eligibility criteria, the 59% aggregate resolved into a bimodal distribution. One patient cluster, defined by a variant frequency pattern in a receptor gene combined with an early-onset comorbidity profile, showed a modeled response distribution approximately 38 percentage points above the complementary cluster. The split was roughly 45/55 in the eligible population. Standard inclusion/exclusion criteria would have enrolled both clusters indistinguishably.

The team had two actionable options once they saw the simulation output: add a genomic stratification criterion to enrich toward the responsive cluster, or power the study to account for the dilution. Both options were available at protocol design time. Neither was visible from the aggregate 59% figure alone. The simulation did not predict whether the compound would succeed. It told the team that their protocol, as written, was enrolling a mixed population whose signal density was lower than their power assumptions required.

Where the method's limits sit

We're not saying virtual patient simulation eliminates Phase II risk. The simulation characterizes the population the trial will enroll; it does not validate the compound. A favorable enrichment scenario in the simulation does not mean the compound will perform well in the enriched subgroup. The pharmacology still needs to hold, and the simulation has no insight into whether it will.

The method also depends on the quality and coverage of the reference population data. For common conditions where structured registry data and EHR cohort statistics are available at scale, the simulation can reach useful specificity. For rare diseases where the reference cohort is measured in the hundreds of patients, the uncertainty bounds on simulated enrichment scenarios are genuinely wide, and the output needs to be interpreted as a range of plausible scenarios rather than a prediction.

There is also a category of Phase II failure that population design cannot address: the compound fails its mechanism in all patients, or is unsafe at the doses required for efficacy. Simulation cannot rescue a compound with a broken hypothesis. It can only reduce the category of failures that result from enrolling the wrong population around a compound that would have worked in the right one.

The decision window this method operates in

The practical value of running a virtual patient simulation is highest at the Phase I to Phase II design transition, when eligibility criteria are still open, adaptive design elements can still be built into the protocol, and sample size assumptions can still be revisited. An enrichment strategy added at protocol amendment after Phase II initiation requires regulatory interaction and, in some jurisdictions, partial site requalification. An adaptive design element not included at original protocol submission is difficult to add mid-trial under ICH E9(R1) without substantial amendment.

The decisions that population simulation most directly informs, eligibility criterion design and enrichment strategy, are both decisions that are costly to reverse once a trial is running. That is the window where pre-enrolment modeling adds its value. Valinor Discovery operates in that window deliberately, not because post-enrolment analysis has no value, but because it is the window where acting on the information is still free.

Explore the platform

The science described here is the basis for the platform's inference architecture. If you are working on a Phase I-II decision, request early access.

Request access Learn how it works