There is a gap between what translational teams often receive from population modeling analyses and what they can actually use to change a protocol. The gap is not about whether a subgroup signal exists. By the time a program is approaching Phase II design, translational teams generally have some indication, from biomarker data, from preclinical models, from early Phase I PK/PD observations, that not all patients will respond equally. The question they need answered is not "is there a subgroup?" but "how large is it, can I identify it with eligibility criteria I can actually implement, and what does my trial look like if I try to enrich for it?"
Standard statistical approaches to subgroup analysis, whether retrospective post-hoc analysis of Phase I data or exploratory analysis of published trials in the indication, produce effect estimates for pre-specified or identified subgroups. What they do not produce is a probability distribution over enrichment scenarios that a team can use for protocol planning. That is what Valinor Discovery's simulation output is designed to provide, and in this post I want to walk through what that output actually looks like and what decisions it supports.
What a simulation output report contains
The primary deliverable from a virtual patient simulation engagement is a set of enrichment scenario distributions, not a single prediction. For a given candidate eligibility criteria set, the report shows the distribution of the simulated enrolled population across the patient characteristics most relevant to the compound's mechanism, with uncertainty quantified explicitly.
The core table covers three sets of scenarios: the estimated composition of the enrolled population under the current eligibility criteria as written, the estimated composition under candidate enrichment modifications to those criteria (adding a stratifier, tightening a disease severity threshold, adding a molecular exclusion criterion), and the estimated composition if no enrichment is applied.
For each scenario, the output includes the estimated fraction of enrolled patients in the modeled responsive and non-responsive subgroups, with 80% and 95% credible intervals derived from the posterior over the reference population's joint distribution. It includes an estimated identifiability score for the enrichment stratifier under consideration, reflecting how well the stratifier's distribution in the eligible population predicts subgroup membership. And it includes an estimated protocol impact summary: how the power assumptions would need to change if the current eligibility criteria are expected to enroll a specific mix of responsive and non-responsive patients.
This last element is the one that most directly connects the simulation output to a decision the protocol team can act on. "Your protocol as designed is likely to enroll a population where the modeled responsive fraction is 38% to 52% under the current eligibility criteria. Adding a genomic stratification criterion shifts that estimate to 61% to 74%, at the cost of reducing the eligible population by approximately 28%." That is the form of the finding that a translational team can bring to a protocol design discussion.
What the output cannot tell you
We are direct with teams about the limits of what the simulation report represents. The enrichment scenario distributions are probabilistic estimates based on a reconstructed model of the reference population, not ground truth about what will happen in any specific trial. The identifiability scores for candidate stratifiers assume that the stratifier's availability in clinical practice matches its availability in the data sources used for reconstruction, which is not always true.
The simulation also cannot predict the magnitude of the treatment effect in the responsive subgroup. That is a pharmacology question that the population model has no direct information to answer. The simulation can tell you how the responsive fraction in the enrolled cohort changes under different eligibility designs. It cannot tell you whether the compound will produce a 30% response rate or a 60% response rate in that responsive fraction. The compound's mechanistic profile, Phase I safety and PK/PD data, and the team's own translational biology judgment are the inputs to that question, and they sit outside the population model.
This boundary is important for how teams use the simulation output in portfolio decision-making. The output is most useful as an input to protocol design, specifically the eligibility criteria and enrichment strategy. It is not useful as an input to go/no-go decisions about the compound's pharmacology, and it should not be presented or interpreted as such. We push back when teams want to use simulation-based enrichment scenario estimates as a proxy for efficacy confidence, because they are measuring a different thing.
The operational workflow for incorporating simulation into protocol design
The workflow that works best in translational teams is one where the simulation is run early in the Phase I to Phase II transition process, when the first draft of the Phase II eligibility criteria is being written, rather than after the Phase II protocol is substantially complete. At that point, the simulation output can directly inform the first real version of the eligibility criteria rather than arriving as a challenge to criteria that have already been through several rounds of internal review.
Practically, this means providing the simulation team with the compound's mechanism of action summary, the proposed disease indication and initial eligibility criteria, and whatever translational biomarker hypotheses have been developed in Phase I. The simulation will generate the reference population model in parallel with the Phase I team's ongoing work, and the scenario output will be ready for the protocol design meeting where Phase II eligibility criteria are first discussed substantively.
For programs that are further along and where the Phase II protocol is already substantially drafted, the simulation output can still be valuable, specifically for evaluating whether an adaptive enrichment design element should be built into the protocol structure. An adaptive design that allows for pre-specified mid-trial enrichment based on an interim analysis requires that the enrichment criteria be defined in the original protocol submission, under ICH E9(R1). Having a simulation-based enrichment scenario analysis available at the time of original protocol writing means the adaptive element can be designed on an informed basis rather than as a precautionary hedge.
Integrating the output into a protocol package
From a regulatory documentation standpoint, the simulation output is best treated as a supporting analysis for the protocol rationale section, specifically the section that justifies the eligibility criteria and enrichment strategy. The simulation report provides the analytical basis for statements about why a specific patient characteristic was included as a stratifier, what the expected enrolled population composition is under the proposed criteria, and why the power assumptions reflect the anticipated responder fraction.
We do not claim that including simulation-based rationale in the protocol package changes the regulatory review outcome for eligibility criteria. Regulators evaluate eligibility criteria on clinical and scientific grounds, and the adequacy of those grounds is a judgment about the biology and the clinical context, not about the modeling methodology. What the simulation output provides is a more rigorous basis for the team's own confidence in the eligibility criteria they have chosen, and a more explicit acknowledgment of the uncertainty in their assumptions about enrolled population composition.
That internal rigor is where we think the simulation output's value is clearest. Most translational teams proceed to Phase II with power assumptions built on aggregate response rates and eligibility criteria drawn from prior precedent in the indication. The simulation output introduces a conditioning step: here is what the enrolled population actually looks like under your criteria, given everything the available data tells us about the reference population's heterogeneous structure. Whether that conditioning step changes the criteria, the power assumptions, or nothing at all depends on what the output shows. But the team is making the decision with better information than the aggregate statistics alone provide.