Skip to main content
Back to Blog
Use Cases

Site selection in rare disease programs: why simulation should come before site activation

Rare disease trial timelines are disproportionately lost at site activation, not at patient identification. The sequence typically goes like this: a CRO is contracted, sites are selected based on investigator reputation and historical enrolment rates in adjacent indications, site initiation visits are conducted, IRB submissions go in, and then, several months in, the sites report back that the eligible patient pool at their center is smaller than projected by a factor of two or three.

This is not a site management failure. It is a site selection failure made earlier, and it stems from confusing the raw patient volume at a center with the volume of patients who will meet the study's eligibility criteria simultaneously, which is a joint probability calculation, not an arithmetic one.

The eligibility bottleneck in rare disease

Rare disease trials routinely carry eligibility criteria that screen out the majority of the addressable population. Genetic confirmation criteria, disease severity thresholds, biomarker windows, and prior treatment exclusions each reduce the candidate pool. When four or five such criteria are applied simultaneously, the eligible fraction can be well under 10% of the diagnosed population at a given center.

This is not an argument against rigorous eligibility criteria. The criteria exist to protect the interpretability of the trial and the safety of participants. The argument is that the impact of those criteria on site-level feasibility cannot be estimated by looking at diagnosed-patient caseload alone. You need an estimate of how many patients at each site will pass all criteria simultaneously, and that estimate requires modeling the joint distribution of the relevant clinical characteristics in that site's patient population.

Site visit questionnaires capture this inadequately. The question "approximately how many patients with condition X have you treated in the past two years?" produces a total caseload figure, not an eligibility-adjusted one. Investigators rarely have ready access to the cross-tabulated sub-population statistics that would let them estimate eligibility-adjusted patient counts, and even when they do, the calculation is done informally. The formal modeling step is almost never performed during site selection.

What the simulation adds

Before a site is contracted, a feasibility simulation models the expected eligible patient pool at that site under the proposed eligibility criteria. The input is whatever cohort-level data is available for that site or for a comparable population: disease registry extracts, patient advocacy organization databases, published single-center series, or data from the CRO's own feasibility databases.

From this, we construct a virtual patient distribution representing that site's plausible patient population for the indication. We then apply the proposed eligibility criteria computationally to the virtual patients, estimating the fraction who would pass each criterion and, crucially, the fraction who would pass all criteria simultaneously.

The simultaneous pass rate is almost always lower than the product of the individual pass rates because the criteria are not independent. Genetic criteria and disease severity criteria are correlated. Biomarker eligibility and prior treatment exclusion are often correlated because patients who have progressed through prior treatment lines have different biomarker profiles than treatment-naive patients. Treating each criterion as independent and multiplying pass rates is a common and often significant error in feasibility estimation.

A worked example in a lysosomal storage disorder program

Consider a development team working on an enzyme replacement program for a lysosomal storage disorder, late 2025. The eligibility criteria include: confirmed genetic diagnosis via one of three specific mutations, disease severity score above a threshold, absence of active cardiac involvement, and no prior gene therapy exposure. Four criteria, each with a published or estimable marginal pass rate in the diagnosed population.

If you multiply the four pass rates, you get a feasibility estimate. If you model them jointly, you get a different and lower estimate, because in this indication the three specific mutations are over-represented in patients with milder disease severity. The genetic eligibility criterion and the severity threshold criterion are negatively correlated in this population. Multiplying marginal pass rates overestimates the eligible fraction for this reason alone.

The simulation output, in this scenario, would be a credible interval around the eligible patient fraction, typically expressed as a posterior distribution over the fraction of the center's diagnosed patient pool who would qualify. That output, produced before sites are contracted, allows the study team to identify which candidate sites have sufficient expected eligible volume to justify activation costs, and which do not.

Ranking sites by expected eligible volume

The output of site-level feasibility simulation is a probability distribution over expected eligible patient volume per site per year. Sites can be ranked by the expected value of this distribution, or by a more conservative quantile such as the 25th percentile, depending on the risk tolerance of the program.

The advantage of ranking by a lower quantile rather than expected value is that it prioritizes sites where the feasibility estimate is both high and confident. A site with an expected eligible volume of eight patients per year but a wide credible interval is less useful than a site with an expected eligible volume of six patients per year with a tight interval, because the wide interval reflects higher uncertainty about the local patient population characteristics.

This ranking can be produced for a candidate list of twenty or thirty sites in a matter of days, before any site visits are conducted. It does not replace site visits, but it makes them more targeted: the team spends site visit budget on the sites that the simulation identified as most likely to be feasible, rather than running visits uniformly across the candidate list.

Eligibility criteria sensitivity analysis

A secondary output of the simulation that development teams often find useful is a sensitivity analysis of each eligibility criterion on overall feasibility. The question is: which criterion is the binding constraint on eligible patient volume?

In some programs, one criterion is responsible for the majority of exclusions. If that criterion can be modified without compromising the scientific or safety rationale for it, relaxing it may substantially improve feasibility. The simulation quantifies the feasibility gain from criterion modification, allowing the clinical team to make an explicit trade-off between scientific rigor and enrolment speed.

We are not suggesting that eligibility criteria should be modified to improve enrolment convenience. That would invert the proper priority order. We are saying that the feasibility impact of each criterion should be quantified before IND, so that the clinical team can make an informed decision about the criteria, rather than discovering mid-trial that a particular criterion is responsible for 40% of patient exclusions in a pool that cannot afford to lose 40%.

What simulation cannot replace

Site simulation is a pre-visit tool, not a substitute for investigator relationships, regulatory site qualification, or local patient advocacy engagement. In rare disease programs, the investigator's relationship with the patient community is often the primary driver of enrolment, and no simulation model captures the effect of a physician who is known and trusted by families in their region.

The simulation also cannot predict patient willingness to participate, travel burden, or caregiver capacity, all of which affect actual enrolment rates in rare disease trials above and beyond eligibility. A center may have a large eligible population by clinical criteria, but if the center is in a location that requires significant travel from the patient population it serves, the conversion from eligible to enrolled will be lower than the simulation predicts.

These limitations argue for using the simulation as a feasibility filter and ranking tool, not as an enrolment forecast. The output should be used to prioritize site visit resources and identify the most vulnerable eligibility bottlenecks. The enrolment forecast itself remains a judgment call that integrates site visit findings, investigator experience, and patient community intelligence in ways the simulation does not capture.

Timing in the development timeline

The practical window for this analysis is the three to six months before site feasibility questionnaires go out, typically during the study design phase when eligibility criteria are still being finalized. This timing serves two purposes: first, it allows the criteria sensitivity analysis to influence the final eligibility design. Second, it allows site selection to begin with a pre-ranked candidate list rather than an unranked one.

Once the trial enters site initiation, the utility of the simulation shifts. It becomes a monitoring tool rather than a selection tool: the predicted eligible volumes can be compared to actual screen rates as the trial progresses, identifying sites where the model substantially over-estimated feasibility and triggering an earlier escalation decision.

In rare disease programs, where the total addressable patient population is measured in hundreds or low thousands globally, every patient lost to a misactivated site or an avoidable screen failure is a resource the trial may not recover. The investment in pre-IND simulation is modest relative to the cost of site activation and far smaller than the cost of a timeline extension caused by under-powered site selection.

Explore the platform

The science described here is the basis for the platform's inference architecture. If you are working on a Phase I-II decision, request early access.

Request access Learn how it works