Skip to main content
Back to Blog
Science

Trial population simulation vs synthetic control arms: different questions, different tools

Two terms are in wide circulation in early-stage drug development and they are frequently conflated in conversations with development teams: trial population simulation and synthetic control arms. We build the first kind of tool. We do not build the second. The distinction matters because conflating them creates misaligned expectations about what pre-IND modeling can do, and about what role regulatory agencies will accept for each.

This post is an attempt to untangle them clearly.

What each term refers to

A synthetic control arm is a statistical construct used in a clinical trial, typically at Phase II or pivotal stage, to replace or supplement a concurrent control arm with historical data. The synthetic controls are patients drawn from prior trial datasets or registries who receive standard-of-care treatment, and their outcomes are used to estimate the counterfactual: what would have happened to treated patients if they had received standard-of-care instead. Synthetic control arms are most discussed in rare disease contexts, where randomized concurrent controls are logistically or ethically difficult, and regulatory agencies including the FDA and EMA have issued guidance documents on their acceptable use in these settings.

A trial population simulation, as we use the term, is a pre-trial tool. It operates before IND, at the stage when a development team is making decisions about whether to advance a candidate, what the eligibility criteria should be, and what the subgroup composition of the enrolled population is likely to look like. The output is not a control arm and is not used in the trial's statistical analysis. It is a decision-support tool for the design and go/no-go stage.

Both involve generating virtual or synthetic patient representations. That is the source of the conflation. But the question they answer, the regulatory context they operate in, and the consequences of being wrong about them are entirely different.

The regulatory position on synthetic control arms

The FDA's guidance on the use of real-world data and real-world evidence in regulatory submissions, along with the EMA's guidance on adaptive designs, establish a framework for when synthetic control data may support regulatory submissions. The core requirement is that the synthetic controls be drawn from an identified, documented, and auditable data source, and that the match between synthetic controls and treated patients be performed using validated methods, typically propensity score matching or weighting, with prespecified statistical analysis plans.

The regulatory bar for synthetic control arms is high because the stakes are high: if the synthetic controls are not comparable to the treated patients in relevant prognostic characteristics, the trial's efficacy finding is biased and the regulatory submission is at risk. Regulatory agencies have accepted synthetic control data in specific settings, generally rare diseases with small populations and unmet medical need, but they have also rejected submissions where the synthetic control methodology was insufficiently rigorous.

None of this applies to pre-IND population simulation. Pre-IND simulation is a planning tool. It is not submitted to regulators, it is not part of the statistical analysis plan, and it does not need to meet the evidentiary standard of a regulatory submission. The appropriate standard for a planning tool is whether it produces useful decision-support output given the available data, not whether it meets the evidentiary standard for a regulatory submission.

Different failure modes

A synthetic control arm that fails produces a biased trial result and potentially a regulatory rejection. The failure is discovered after the trial is complete, when it is too late to redesign.

A trial population simulation that fails produces a subgroup composition estimate that diverges from the actual enrolled population. This is a planning error, not a scientific error. The consequences are: a mis-sized trial, an unexpected subgroup distribution at readout, or a go/no-go decision made on incorrect population assumptions. These are serious but recoverable, especially if the simulation flags its own uncertainty appropriately. A simulation that produces wide credible intervals when the input data is sparse is failing informatively: it is telling you that you don't know enough to commit to a tight subgroup estimate, and that is correct behavior.

This is not to say pre-IND simulation errors are inconsequential. A team that relies too heavily on a poorly constrained simulation to make a go/no-go decision is making a different kind of mistake than a team that uses a well-constrained simulation appropriately. The point is that the error modes are structurally different: synthetic control arm errors affect the validity of trial evidence; population simulation errors affect the quality of pre-trial decisions.

Where the tools interact

The two approaches are not entirely independent. A development team that uses pre-IND population simulation to understand the subgroup composition of their likely enrolled population will, if the trial eventually proceeds, have much better information for designing the synthetic control arm than a team that did not do pre-IND simulation.

Specifically: the population simulation output includes a characterization of the dependence structure between the clinical variables most relevant to the indication. When the team later constructs a synthetic control arm from historical data, the matching variables they choose should reflect the prognostically relevant characteristics, and the pre-IND simulation provides an evidence-based basis for identifying which variables those are. This is an indirect benefit: the simulation does not produce the synthetic control arm, but it produces the population characterization that informs it.

We are not claiming this as a primary use case for our platform. We are noting that teams sometimes ask whether the simulation output can be used directly to construct synthetic controls. The answer is no: the virtual patients in a trial population simulation are not drawn from an identified and auditable external data source in the sense that regulatory guidance requires. Using them as synthetic controls would not meet the evidentiary standard. The virtual patients are a planning tool, and substituting a planning tool output for a regulatory-grade evidence source would be a category error.

Complementary timing

The practical resolution is to treat the two tools as operating at different phases of the development lifecycle. Trial population simulation is a pre-IND tool: it operates during study design, three to twelve months before IND filing, and its output informs eligibility criteria, go/no-go decisions, and site selection feasibility. Synthetic control arm methodology is a late-Phase design decision: it is discussed with regulatory agencies in pre-Phase II meetings, designed as part of the statistical analysis plan, and executed using identified historical cohort data.

Teams that understand this temporal separation use both tools appropriately. Teams that conflate them tend to either over-rely on pre-IND simulation output in contexts where regulatory-grade synthetic controls would be required, or they dismiss pre-IND simulation as unnecessary because they plan to use synthetic controls at trial stage. Both of these are errors. The pre-IND tool and the trial-stage tool are answering different questions at different times.

A note on terminology in the field

The terminology in this space is not stable. Different vendors, consultants, and academic groups use the term "digital twin," "virtual patient," "in silico trial," "synthetic cohort," and "computational patient" to mean different things. Some of these terms overlap substantially with what we mean by trial population simulation; others are closer to mechanistic physiological modeling or systems pharmacology, which is a distinct approach with its own validation standards.

When evaluating any computational approach in early drug development, the question to ask first is: what is the question this tool is designed to answer, and at what stage of the development lifecycle? If the answer is "it replaces the clinical trial" or "it produces data suitable for regulatory submission without clinical trial evidence," be skeptical. If the answer is "it helps you understand your trial population before you commit to a design," that is a defensible claim and worth evaluating on its technical merits.

Explore the platform

The science described here is the basis for the platform's inference architecture. If you are working on a Phase I-II decision, request early access.

Request access Learn how it works