Real-world evidence has become a standard input to trial design. Regulatory guidance documents from FDA and EMA have both acknowledged its relevance for natural history characterization and control arm augmentation. But the form in which RWE actually arrives, aggregated cross-tabulations, published cohort summaries, registry-derived frequency tables, is not the form you need for virtual patient simulation.
Simulation requires joint conditional distributions. You need to know not just that 34% of the reference population has a particular disease severity score, but how that severity score distributes conditional on age, prior treatment, comorbidity burden, and whatever genomic stratifiers are relevant to the compound's mechanism. Published sources almost never report those conditional distributions directly. They report marginal statistics with selected bivariate breakdowns, and the conditional structure you need has to be recovered from what you can observe.
This is the problem Bayesian cohort reconstruction addresses. It is the core methodological challenge we work on at Valinor Discovery, and getting it right is what separates a virtual patient distribution that is useful for protocol decisions from one that just looks plausible on the surface.
Why marginal statistics systematically understate population complexity
Most clinical research publications report outcome associations adjusted for confounders rather than the joint distributions that generate those associations. A published hazard ratio adjusted for age, comorbidity, and ECOG status tells you about the treatment effect in the analyzed population. It does not tell you the joint frequency of the covariate combinations that together define subgroup membership in a trial.
Registry data is structurally similar. Disease registries typically report counts and percentages for individual variables and selected pairs. If you want the proportion of patients with both a specific comorbidity profile and a particular molecular subtype, you cannot compute that from standard registry tables unless the registry explicitly cross-tabulates those variables. In most disease areas, that level of cross-tabulation is not published, and often is not computed at all by the registry administrators.
This matters enormously for simulation. When we generate virtual patients from only marginal distributions, we are implicitly assuming independence between variables. That independence assumption is almost always false. Age and comorbidity burden are correlated. Molecular subtype and prior treatment history are correlated in many oncology indications. Simulating from marginals produces a virtual population that has the right average properties but the wrong internal structure, and that internal structure is exactly what subgroup prediction depends on.
The Bayesian network approach to distribution reconstruction
The reconstruction method we use treats the problem as a probabilistic inference task. We start with what is observed, the marginal statistics and bivariate relationships available in published sources, and treat the full joint distribution as a latent quantity to be inferred.
The mechanics are built around a Bayesian network structure that encodes the conditional dependence relationships between patient characteristics. The network topology is not learned purely from the aggregate statistics, because the data is too coarse for reliable structure learning. Instead, we specify a prior topology derived from biological and clinical domain knowledge: for a given indication, which variable pairs are likely to be conditionally dependent given a third variable, and in what direction? That prior structure constrains the space of joint distributions the inference can reach.
Given the network topology, we fit the conditional probability tables to the observed marginal constraints using a combination of expectation propagation and constraint-based optimization. The output is a parameterized Bayesian network from which we can sample joint patient profiles that satisfy all the observed marginal constraints while respecting the conditional independence structure encoded in the topology.
The uncertainty in this reconstruction is explicit and tracked. For variable pairs where the marginal data strongly constrains the joint distribution, the posterior over conditional probabilities is tight. For pairs where the marginal data leaves substantial ambiguity, the posterior is wide, and the resulting virtual patient samples carry a correspondingly wide uncertainty range in any downstream simulation.
Handling the multi-source integration problem
Real programs rarely have a single clean RWE source for the reference population. A typical project involves a primary disease registry with demographic and clinical variables, a genomic frequency database for variant distributions, published natural history studies that cover partially overlapping patient characteristics, and sometimes unpublished cohort statistics from collaborating clinical sites that have signed data-sharing agreements.
Each source covers different variables, at different levels of detail, with different patient selection criteria. Combining them naively, treating all sources as observations from the same population, would introduce systematic bias because the selection criteria differ. A patient enrolled in a disease registry is not statistically identical to a patient in a natural history study even if both have the same diagnosis.
The Bayesian framework handles this through hierarchical modeling of source-level selection. We model each source as a sample from the underlying patient population conditioned on that source's implicit selection criteria, and we propagate uncertainty about those selection effects into the reconstructed joint distribution. The result is a unified posterior over the reference population that accounts for the different windows each data source provides on that population.
What reconstruction quality actually looks like in practice
The validation question for a reconstructed joint distribution is whether it would predict the bivariate and multivariate statistics of held-out data from the same reference population. For most disease areas where published data is reasonably rich, we see reconstruction models recover held-out two-way cross-tabulations with mean absolute errors in the 2 to 5 percentage point range. That is the accuracy regime that makes the downstream simulation useful for protocol decisions without overstating its precision.
There are disease areas where we honestly cannot hit that standard. Rare disease indications with published cohort sizes below 300 patients, or where the only available data is from highly selected referral populations, produce reconstructions with much wider uncertainty bounds. In those cases we present the simulation output as a scenario distribution rather than a probability estimate, which is an honest representation of what the data actually supports.
The connection to enrichment strategy
The reason this methodological infrastructure matters for trial design is that enrichment strategy evaluation depends entirely on knowing how patient characteristics co-distribute in the eligible population. If you want to ask whether adding a genomic stratifier to your eligibility criteria changes the proportion of responsive patients in the enrolled cohort, you need to know how the stratifier's distribution overlaps with the distribution of the clinical characteristics already in the criteria. That overlap is a conditional distribution question, and it can only be answered from a joint model of the reference population.
We are not claiming that Bayesian cohort reconstruction produces ground truth about the reference population. It produces the best posterior estimate consistent with the available aggregate data and domain-knowledge priors, with explicit uncertainty quantification. That is considerably more informative than using marginal statistics as if they were independent, which is the implicit assumption in most current practice. The difference matters when the decision you are making is whether to add or remove an enrichment criterion from a protocol that will cost tens of millions of dollars to run.