Valinor Discovery is open for early access today. This post explains what we built, why we built it, and what it is not.
The three of us, Marcus, Priya, and I, came to this problem from different directions. Marcus spent years running patient subgroup analyses in translational oncology programs and watching the same pattern repeat: a Phase II program that looked compelling at the aggregate level fell apart when the subgroup data came in at Phase III, and the subgroup question was one that could have been asked earlier if the tools had existed to ask it. Priya built probabilistic inference systems for multi-omics data integration at a Cambridge research group and became convinced that the methods were mature enough to build production clinical applications on top of. I came from the clinical operations side, running trial feasibility assessments where the gap between projected and actual enrolment was the main reason programs fell behind schedule.
Three different ways of arriving at the same observation: the patient population question is asked too late in drug development, and the infrastructure for asking it earlier has not kept up with the computational methods that would make it tractable.
The specific problem we are solving
When a development team files an IND, they have typically characterized their candidate extensively in experimental systems: cell lines, animal models, in vitro pharmacology. They have characterized the target biology, the mechanism of action, and the safety profile in preclinical species. What they have not usually characterized in any rigorous way is the patient population they are about to enter.
The population question is not simply "how many patients have this disease." It is: among the patients who are likely to enroll in this trial, what is the distribution of the clinical characteristics that determine whether this candidate will produce a benefit signal? How much of the enrolled population will fall into the subgroup where the mechanism of action predicts the strongest response? How does that fraction change depending on the eligibility criteria?
These are quantitative questions. They have quantitative answers, at least within uncertainty bounds that depend on the available cohort data. But the development team at IND stage does not typically have a platform for computing those answers. They have clinical intuition, investigator consultation, and sometimes a disease model from a consulting firm. None of these is a systematic probabilistic model of the clinical population.
What Valinor Discovery does
We build virtual patient cohorts from multimodal cohort data. The input is whatever cohort-level statistics are available for the target population: disease registry summaries, published Phase II readout data, biomarker distributions from patient advocacy organization datasets, or structured EHR extracts under data sharing agreements. The output is a probabilistic model of the patient population, expressed as a joint distribution over the clinical variables most relevant to the development question.
From that distribution, we run simulations of the trial population under different eligibility criteria and subgroup definitions. The result is an estimate of the subgroup composition of the enrolled population, including the fraction of patients likely to fall into the subgroup where the mechanism of action predicts benefit, and the uncertainty around that estimate given the available data.
The technical core is a Bayesian inference engine built on copula-based joint distribution modeling. Priya has written about the approach in more detail in a companion piece on the methodology. The short version is that we model the dependence structure between clinical variables, not just their marginal distributions, because the joint distribution of clinical characteristics is what determines subgroup composition, and the joint distribution is not recoverable from marginals alone.
What we are not
We are not a mechanistic disease simulation. We do not model drug pharmacokinetics or pharmacodynamics, cell-level disease mechanisms, or patient physiology. Those are different tools that answer different questions. There are well-established systems pharmacology and physiologically-based pharmacokinetic platforms for that work. We are a clinical population modeling tool, focused on the distributional question of who is in the trial and what their clinical characteristics look like, not on the mechanistic question of how the drug interacts with biology.
We are also not a synthetic control arm platform. Synthetic control arms are a regulatory-grade evidence construct used in clinical trials, with specific guidance from FDA and EMA on their acceptable use. Our output is a pre-trial planning tool. It is not submitted to regulators and is not part of a trial's statistical analysis plan. We have written separately about the distinction between these two approaches because the conflation creates misaligned expectations about what pre-IND modeling can and cannot do.
And we are not a replacement for clinical expertise. The simulation output is an input to the clinical development team's decision-making, not a replacement for it. The clinical team provides the mechanistic interpretation, the regulatory strategy, and the ultimate go/no-go judgment. We provide quantitative characterization of the population question that the clinical team can use alongside those other inputs.
Who we are building for
Our early-access program is focused on development teams working in oncology, rare disease, and CNS indications at the pre-IND stage, specifically programs preparing for a Phase I-II decision where the subgroup composition of the enrolled population is a material uncertainty. These are programs where the eligible population is likely to be heterogeneous, where subgroup failure is a plausible risk, and where a rigorous pre-IND population characterization would change how the program is designed.
We are a small team of three, based in Boston. We are not promising to solve every problem in drug development, and we are not pitching a broad platform with features we have not built. The platform we are launching does one thing: it builds virtual patient distributions from cohort data and runs subgroup composition simulations against those distributions. We have built that thing carefully, and we think it is useful for the specific programs it is designed for.
A note on what we will publish
This blog is where we will share the methodological thinking behind the platform, discuss specific use cases in detail, and write about the broader science of clinical population heterogeneity and its implications for drug development. We have strong opinions about where the field is heading and about the specific technical choices we made. We will share those opinions here, with the precision that the subject requires.
If you are working on a pre-IND program and want to discuss whether the platform is a fit for your situation, reach out at [email protected]. We are accepting a small number of early-access programs and will respond to every inquiry directly.