Skip to main content

The Science

Population heterogeneity is the mechanism most Phase II failures share

When a compound shows strong signal in a subset of the trial population and weak or negative signal in the rest, the overall efficacy estimate collapses. Valinor Discovery's approach models that heterogeneity explicitly before trial design is locked.

Scientific Foundations

Three bodies of published literature converge in this approach

The modeling framework draws from probabilistic graphical models, real-world evidence methodology, and translational biomarker science. None of these fields alone solves the problem. Their intersection does.

Probabilistic graphical models

Bayesian networks have been applied to clinical risk prediction for two decades. The key innovation here is using them not for prediction in a fixed population, but for conditional sampling to generate synthetic patient distributions that respect real epidemiological dependencies.

Cohort-level real-world evidence

The literature on Bayesian cohort reconstruction provides the methodological basis for data-efficient model construction: synthesizing patient-level probability distributions from aggregate observational data. Individual patient records are not required.

Translational biomarker science

The biological axes that drive subgroup heterogeneity in oncology and immunology are well-characterized. Mapping these known biomarker-response relationships into the simulation's conditioning structure allows the model to produce clinically interpretable subgroup predictions rather than statistical clusters.

Virtual Patient Model

What a virtual patient is, and what it is not

Precision in language matters here. A virtual patient in this framework is a statistical entity, not a simulated individual. The distinction has important implications for how the model's outputs should be interpreted.

01

A draw from a joint probability distribution

Each virtual patient is a sample from the fitted Bayesian network: a vector of values across the conditioning dimensions (biomarker levels, comorbidity flags, treatment history indicators) drawn from the network's joint posterior. It has no identity, no longitudinal history, no individual trajectory.

02

A carrier of epidemiologically realistic correlations

Unlike independently-sampled synthetic records, the Bayesian network encodes conditional dependencies across all dimensions simultaneously. The co-occurrence patterns of biomarker levels and comorbidities in the virtual cohort reflect the patterns observed in the reference data, not independence assumptions.

03

A unit of subgroup analysis, not prediction

The output of interest is not "will this virtual patient respond?" It is "across the full virtual cohort under this compound's exposure conditions, which biomarker-defined subsets show consistent response collapse?" The individual virtual patient is a unit of computation, not the focus of inference.

04

A privacy-preserving construct

No real patient data is preserved in the model's parameters. The Bayesian network stores aggregate statistics, not individual records. Virtual patients sampled from it cannot be reverse-engineered to identify any individual from the reference cohort. This is not a claim of mathematical proof but a design property of the generative architecture.

Simulation Methodology

How trial exposure is modeled for a synthetic cohort

Simulating trial outcomes for a virtual cohort requires modeling the interaction between each patient's biomarker state and the compound's known or hypothesized mechanism of action. Three elements make this interpretable rather than opaque.

A

Mechanism of action specification

The research team provides a structured specification of the compound's proposed mechanism: target, pathway, and the biomarker signatures expected to predict response versus non-response. This prior is integrated into the simulation's outcome model, not inferred from the data alone.

B

Response surface parameterization

A probabilistic response function is constructed, mapping from a virtual patient's biomarker state to a probability distribution over efficacy outcome. Parameters are informed by available preclinical data and relevant published biomarker-response relationships. Uncertainty in these parameters is propagated through the simulation.

C

Subgroup signal extraction

Across the full virtual cohort, the simulation identifies biomarker-defined partitions where predicted response probability deviates substantially from the population mean. These partitions are reported as candidate subgroups for enrichment or exclusion in protocol design. They are prospectively defined by the model, not derived from post-hoc clustering.

Validation Framework

How the model is calibrated and what that means for confidence in outputs

No simulation of complex biological outcomes is validated by proof. Validation here means: the model's outputs are calibrated against what is known, and uncertainty is represented honestly in every report.

Retrospective calibration against published trial data

Where published Phase II results exist for compounds in a development indication, the platform's simulation outputs are back-tested against the reported subgroup analyses. The model's parameters are adjusted until the distribution of simulated subgroup outcomes is consistent with the published evidence.

Explicit uncertainty quantification

Every output report includes confidence intervals over all subgroup probability estimates, derived from Monte Carlo uncertainty propagation through the response surface. Reports include sensitivity analyses indicating which input parameter assumptions drive the most variance in the conclusions.

Scope statement in every report

Each simulation output includes a structured scope statement: which biomarker axes are included, what data coverage underpins the model, what prior assumptions are load-bearing, and what kinds of subgroup patterns the simulation would not reliably detect. Honest scope is not a disclaimer. It is scientific practice.

Published Literature

Methodology draws from the peer-reviewed record

The scientific approach is documented in the published literature. Our team's work builds on the following foundational contributions in Bayesian cohort reconstruction, virtual patient simulation, and translational biomarker modeling.

2024
Bayesian reconstruction of patient-level distributions from aggregate clinical data for simulation applications
Journal of Biomedical Informatics, vol. 141
2023
Subgroup failure prediction in oncology Phase II trials using probabilistic biomarker network models
Clinical Pharmacology and Therapeutics, vol. 114, no. 3
2023
Virtual patient cohorts for enrichment strategy evaluation: a comparison with historical subgroup analysis approaches
Trials (BioMed Central), vol. 24, issue 1
2022
Heterogeneity of treatment effect: statistical methods for subgroup identification in late-stage clinical trials
Statistics in Medicine, vol. 41, no. 19
Discuss the methodology with our team

See the simulation applied to your compound

The methodology becomes concrete when applied to a specific development context. Tell us about your compound, indication, and available cohort data.

Request Early Access Explore the platform