From The Editor | September 15, 2026

A Biostatistician's Caution Around Digital Twins In Trials

John Oncea Profile Photo

By John Oncea, Chief Editor, Clinical Tech Leader

digital twins, virtual reality, and advanced technology-GettyImages-2154151760

AI is rightfully garnering a lot of attention in our industry right now, but don’t sleep on another up-and-coming technology: digital twins, pioneered in the 1960s by NASA and coming to a clinical trial near you.

Digital twins have become one of clinical development’s more seductive ideas: build a virtual model of a patient to either reduce the number of participants needed for a placebo arm or, more radically, replace it entirely. It’s the kind of concept that shows up in vendor decks with total confidence.

So, when I sat down with Natasa Rajicic, ScD, a biostatistician with 25 years of experience spanning Pfizer, Cytel, and her own consulting practice, I expected an enthusiastic answer when I asked her what she thought about the technology. I got something more useful: a working statistician’s real-time skepticism.

Rajicic hadn’t given digital twins much thought before our conversation, telling me that had I asked her about it a week before our conversation, she would have known nothing about it. Coincidentally, that morning, while making coffee, she read a paper by Stephen Senn, an award-winning consultant statistician and former academic specializing in drug-development methodology, clinical-trial planning, regulatory strategy, statistical analysis, and training, which raised questions about the concept.

What Randomization Is Actually Doing

To understand her skepticism, you have to understand what randomization is for. When a trial randomizes patients into an active arm and a placebo or standard-of-care arm, the goal isn’t to make every patient identical, or even to ensure perfect balance between the groups with respect to prognostic factors. The goal is to make treatment assignment a matter of chance rather than judgement, which provides a basis for valid statistical inference, even though important prognostic factors may still end up unevenly distributed across groups.

“We don’t know all the things that make them similar or dissimilar,” Rajicic explained. “We kind of hope that randomization takes care of much of that stuff.” Stratification and baseline adjustment handle the variables researchers can measure. Randomization is the mechanism that hedges against the ones they can’t.

A digital twin, by definition, has to model a patient without that hedge. It has to represent, computationally, what would have happened to a specific person under the treatment they didn’t receive. That requires knowing, or at least approximating, all the variables that make one patient’s response differ from another’s. Rajicic’s objection is direct: “There’s so much you don’t know. And to me, I would have to be convinced that that’s possible.”

Why The Engine Analogy Doesn’t Hold

During our conversation, I offered a comparison: if you’re simulating the mechanics of an engine, you can trust the simulation because engineers understand engines with enormous precision. Rajicic agreed but drew the line exactly where you’d expect a statistician to draw it. Engine failure follows physical laws that are well characterized. Human response to treatment does not come with that kind of settled, mechanistic understanding, at least not for now.

It’s not that she thinks that use of digital twins has no place in trial design. In its more established form, a digital twin is used within an otherwise randomized trial, for example as part of a covariate-adjusted or efficiency-enhancing analysis that, as a result, can end up saving sample size. Rajicic’s skepticism is aimed at the more ambitious version: using a digital twin as a full stand-in for a control arm.

It’s that the confidence a digital twin implies, standing in for what a real patient would have experienced, requires a depth of biological understanding that current science doesn’t have for most conditions. Randomization exists precisely because that understanding is incomplete, and replacing it with a model built on the same incomplete understanding doesn’t resolve the uncertainty; it just moves it somewhere less visible.

 

Where She Sees AI Actually Earning Its Keep

This is the part of the conversation that mattered most to me because it wasn’t dismissive. Rajicic isn’t opposed to AI or modeling in general. She’s precise about where it currently delivers value in her own work, and where it doesn’t.

Her day-to-day use of generative AI tools is for research and writing. “It’s not designing new things,” she said. “It’s just synthesizing information that’s out there at the given moment.” She compared it to a faster, more targeted version of the literature search she used to do with journals and Google: still work, but work that used to take her far longer. For a consultant who has to get up to speed on a new indication or methodology quickly, that’s genuinely valuable, even if it isn’t glamorous.

She also pointed to a different category of tool: online, interactive statistical software that lets her simulate study designs, model dose-response scenarios, or stress-test a protocol using theoretical distributions rather than real patient data. “Maybe five, ten years ago you would have to know the very complex methodology behind it to a minutiae and then also program that,” she said. Now those tools are accessible without months of custom programming, which she called “life-changing” for her consulting work. That’s a meaningfully different use case from a digital twin: it’s simulating hypothetical scenarios to inform design decisions, not claiming to predict what a specific real patient would have experienced.

One more detail is worth noting for any clinical technology leader evaluating AI vendors: several of Rajicic’s clients now run these tools inside closed, sandboxed environments specifically so proprietary or patient-level data never touches an open system. “It’s very dangerous as a consultant, or working for a company, you don’t want to put anything, of course, that is even close to anything proprietary,” she said. The infrastructure question, not just the algorithm, is part of what makes AI usable in this industry.

A Framework For Weighing The Next Modeling Claim

Rajicic’s reasoning suggests a useful filter for evaluating any vendor pitch that leans on simulation, digital twins, or AI-driven prediction. Ask what the model claims to know about the underlying mechanism, and whether that mechanism is genuinely well understood or still substantially uncertain. Ask whether the tool is generating hypothetical scenarios to inform a design decision, which is defensible, or claiming to substitute for what an individual patient’s outcome would have been, which is a much higher bar. And ask whether the term “AI” is performing real technical work in the claim or serving as a catch-all that blurs the distinction between literature synthesis and predictive modeling.

Rajicic’s own standard is simple, and it’s the one I’d recommend to any reader evaluating these pitches: convince her the underlying uncertainty has actually been reduced, not just modeled around. Until then, she’ll keep applying the same test she applies to everything else in trial design. Not whether it’s exciting, but whether it’s been shown to work.