Building Effective, Retrospective Decision-Grade Disease Registries
By Oznur Seyhun, Cofounder, ECONiX Research; Anusha G., Senior Analyst, Beroe Inc.; and Monica Nandagopal, Senior Analyst, Beroe Inc.

The pharma industry is seeing rapid growth in real-world data (RWD), but converting it into reliable, decision-grade evidence remains challenging. Disease registries provide detailed, longitudinal, disease-specific data on progression, treatment exposure, and outcomes, but they are limited to claims or general healthcare databases. FDA and EMA guidance increasingly supports the use of registries and emphasizes clear populations, robust protocols, data quality, and appropriate analysis. A registry’s suitability depends not on its size or data quality alone but on whether it captures the data, population, and time needed to address a specific clinical, regulatory, or market-access question. This white paper outlines key considerations for building decision-grade retrospective disease registries, including defining the evidence question, selecting and linking data sources, establishing patient identification and cohort criteria, addressing data quality and bias, and ensuring regulatory and privacy compliance.1,2
Challenges In Existing Registry Designs
A registry may be suitable for its original purpose but not for a later evidence question. Retrospective use is constrained by the original design around the patient population, variables collected, outcome definitions, and data collection frequency. Decision-readiness requires assessing whether the registry fits the specific research, regulatory, safety, or reimbursement objective.
Purpose misalignment
Registries are often built around original operational objectives rather than future evidence needs, creating a gap between data relevance and decision relevance. This can limit their use for new comparative studies. For example, a registry designed to assess disease prevalence may capture demographics well but lack treatment or sequencing data, while a treatment-focused registry may lack disease-severity measures. The FDA's guidance on assessing registries for regulatory decision-making recommends assessing whether an existing registry is relevant and reliable for the specific research objective, including its required exposures, outcomes, covariates, and applicability to the target population.1
Retrospective constraints
Retrospective registry use introduces historical challenges as definitions, clinical practices, coding systems, and technologies change over time, making variables difficult to compare across years, sites, or patient groups. Patient data may also be fragmented across EHRs, claims, laboratories, and pharmacies, requiring rigorous linkage, interoperability, and harmonization. Differences across sites, care settings, and access to care can further make observed populations and measurements unrepresentative of the broader target population.
Missing variables
Registries often lack information needed for specific research or regulatory questions. Missing critical variables can prevent valid adjustment for confounding and bias, limiting reliable comparative-effectiveness assessments or health technology assessments (HTAs). For example, cancer registries may contain valuable clinical data but lack comorbidities, concomitant therapies, resource utilization, and safety outcomes needed for comprehensive regulatory evaluation.3
Gaps in overall data completion
Data completeness can mask gaps in variables critical to specific decisions. Missing disease-severity measures, biomarkers, or other key variables in important subgroups can impede unbiased adjustment and compromise comparative-effectiveness or HTA analyses.4 Organizations should adopt a decision-weighted completeness assessment that identifies which variables and patient groups are missing before committing to an analysis.
Defining the Decision and Evidence Use Case


Source: Beroe Analysis1,2
Using The Right Data For Decision Making
Once the evidence question is defined, it should be translated into a data specification that identifies the variables, measurements, covariates, and follow-up required to answer the question. Regulatory guidance recommends assessing whether the available data are relevant to the research question and whether key variables, such as exposures, outcomes, and covariates, are sufficiently reliable and complete.1,4,5 Their definition, measurement method, timing, completeness, provenance, and consistency should also be assessed. Where critical information is unavailable, linking the disease registry with complementary data sources, such as EHR, claims, laboratory and pharmacy data, imaging, genomics, or PROs, may help strengthen the evidence base.
Converting RWD into decision-grade real-world evidence (RWE) requires an intentional, question-led architecture rather than opportunistic secondary data collection. The dataset must be fit for its intended regulatory, HTA, or comparative-effectiveness purpose.
Implement a "target trial" variable specification
Rather than mining available fields post hoc, translating an evidence question into a credible analytical file requires formulating the study within a causal target trial emulation framework.6 This demands clear pre-specification of decision-critical variables — notably baseline confounding factors, specific exposure timelines, and validated outcome endpoints — to avoid structural immortal time and selection biases.
Evaluate fitness-for-use via standardized regulatory frameworks
Raw record completeness frequently masks acute deficits in variables that determine clinical decisions. Methodological frameworks published by the Duke-Margolis Center for Health Policy [7] and formal health authority standards delineate1,4 data quality across two discrete dimensions: data relevancy (availability of key exposures, outcomes, and essential confounders) and data reliability (accuracy, completeness, and provenance).
Address the decision-variable gap in specialized registries
While disease-specific registries excel at capturing granular clinical staging, histology, and molecular markers, they frequently lack non-index clinical information. Systematic evaluations of repurposed oncology and rare disease registries show recurring information gaps in performance status, functional scores, outpatient concomitant medications, and adverse event profiles, which restricts their utility for unadjusted comparative evaluation.3,20
Quantify linkage validity and selection attrition
Bridging missing variables by linking registries to administrative claims, EHRs, or national mortality registries is often necessary to complete the patient journey.8 However, investigators must evaluate linkage mechanisms using established validation frameworks, such as the RECORD-PE guidelines.9,10. Deterministic or probabilistic linkage errors non-randomly misclassify exposure and outcome windows, introducing informative attrition that must be quantified and adjusted via sensitivity analyses.
Key Data Sources And Use
The value of a retrospective disease registry depends on the quality and suitability of the data used. Since no single RWD source captures the complete patient journey, registries often combine complementary datasets based on the research question, target population, exposures, and required observation period.8
The tables below demonstrate the type of data sources used for various evidence requirements and its adoption levels, ranging from low to high
Table 1: Defining the population

Table 2: Characterizing treatment

Table 3: Measuring outcomes

Source: Beroe Analysis8,11
Linking EHRs, claims, laboratory, pharmacy, mortality, genomic, and disease-specific registry data can provide a more complete longitudinal view of the patient journey. However, linkage accuracy, duplicate records, patient matching, timing alignment, and potential linkage bias should be assessed and documented.11
Definitions And Inclusion/Exclusion Criteria
In a retrospective disease registry, patient identification directly affects the validity and interpretation of the evidence generated. The figure below demonstrates the cohort-construction rules that should be clinically clear and reproducible.8

Source: Beroe Analysis12,13
Inclusion and exclusion criteria
The figure below provides an understanding of eligibility criteria that identify patients who are clinically relevant and adequately represented in the available data.14

Source: Beroe Analysis
Data Quality, Data Privacy, And Regulations
Large-scale RWD become decision-grade evidence when supported by strong controls for data quality, traceability, privacy, and governance.5,15,16 Generating decision-grade evidence from retrospective disease registries requires moving beyond passive data warehousing to an auditable, regulatory compliant governance framework. Large volumes of RWD do not automatically yield regulatory-grade conclusions; datasets must demonstrate fitness-for-purpose through systematic validation of quality dimensions, verifiable data provenance, rigorous missing data handling, and strict adherence to data protection standards.
Implement multidimensional quality harmonization
Establishing data reliability begins with evaluating fitness-for-use across structured data quality dimensions. Following established quality ontologies, such as the Kahn harmonized typology17 and the joint HMA-EMA Data Quality Framework4, registries must undergo evaluation across conformance, completeness, and plausibility. Conformance checks ensure values align with standardized data models, while plausibility checks assess chronological event sequences, biologically implausible laboratory values, contradictory diagnoses, and patient duplication before protocol design begins.
Establish full provenance and transparent lineage
Regulatory acceptability hinges on end-to-end data traceability from primary clinical origin to the final analytic dataset. Global harmonization guidelines, such as the ICH M14 framework for non-interventional safety studies5, mandate transparent documentation of all data curation pipelines. Study protocols must include immutable audit trails covering vocabulary mapping (e.g., ICD-10, SNOMED CT, MedDRA, RxNorm), natural language processing (NLP) extraction validation metrics, variable derivation algorithms, and date-stamped data freezes.
Classify mechanisms and model missingness directly
Overreliance on crude dataset-level completeness metrics obscures informatively missing variables that induce severe confounding or selection bias. As highlighted in methodological standards for missing data,18 analysts must distinguish between administrative non-capture, absence of a clinical event, and site-level practice variations. Because missingness in routine practice is rarely missing completely at random (MCAR), simple complete-case approaches should be replaced by principled methods, such as multiple imputation, targeted maximum likelihood estimation, or quantitative bias analysis,19 complemented by rigorous sensitivity analyses.
Navigate privacy, re-identification risk, and data protection
Secondary use of linked registries must balance clinical granularity with stringent data protection mandates, such as GDPR and the HIPAA Privacy Rule. Multi-source linkage of clinical, claims, and genomic records substantially escalates re-identification risks.21 Deploying privacy-enhancing technologies (PETs) — including cryptographic tokenization, federated analytical networks, and formal k-anonymity or differential privacy thresholds — enables cross-institutional research while mitigating disclosure vulnerability.
Conclusion
Transforming retrospective disease registries into decision-grade evidence requires an intentional, end-to-end methodological framework rather than opportunistic data aggregation. While the abundance of routine healthcare data offers unprecedented opportunities to evaluate longitudinal outcomes and diverse populations outside randomized trials, large data volume alone cannot compensate for systematic biases, unmeasured confounding, and missing critical variables. Registries achieve true decision-grade utility only when study designs strictly emulate a causal target trial,6 aligning data collection directly with the regulatory, clinical, or HTA questions being asked.
A successful evidence generation strategy relies on multi-source data linkage that balances analytical completeness with methodological rigor. Bridging information gaps across EHRs, claims, lab records, and registries provides a more complete view of the patient journey, provided that linkage validity and selection attrition are transparently evaluated and reported under standardized frameworks such as RECORD-PE.10 Furthermore, organizations must move from passive data completeness to structured data reliability by embedding harmonized quality dimensions,4,17 transparent data provenance, validated natural language processing pipelines, and principled methods for missing data handling.18
Ultimately, the transition from RWD to regulatory and payer-accepted evidence demands full transparency, reproducible phenotyping, and robust governance. By adhering to global harmonization standards, such as the ICH M14 guidelines5, and deploying privacy-enhancing technologies that safeguard patient rights under modern data protection frameworks21, retrospective disease registries can deliver the reliable, verifiable, and clinically meaningful evidence required to confidently guide healthcare decisions.
References:
- US Food and Drug Administration, “Real-World Data: Assessing Registries To Support Regulatory Decision-Making for Drug and Biological Products,” December 2023. [Online]. Available: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/real-world-data-assessing-registries-support-regulatory-decision-making-drug-and-biological-products.
- European Medicines Agency, “Guideline on registry-based studies - Scientific guideline,” 2021. [Online]. Available: https://www.ema.europa.eu/en/guideline-registry-based-studies-scientific-guideline.
- M. C. Wilpshaar, D. Hilarius, L. Prada, A. Weltermann, and C. Leopold, “Repurposing Registries: Completeness of Real-World Data for Regulatory and HTA Purposes in Three Cancer-Focused Registries,” Clinical Pharmacology and Therapeutics, 2026.
- European Medicines Agency, “Data quality framework for medicines regulation,” 2022. [Online]. Available: https://www.ema.europa.eu/en/about-us/how-we-work/data-regulation-big-data-other-sources/data-quality-framework-medicines-regulation.
- European Medicines Agency, “ICH M14 Guideline on general principles on planning, designing, analysing, and reporting of non-interventional studies that utilize Real-World Data for safety assessment,” September 2025. [Online]. Available: https://www.ema.europa.eu/en/documents/scientific-guideline/ich-m14-guideline-general-principles-planning-designing-analysing-reporting-non-interventional-studies-utilise-real-world-data-safety-assessment-medicines-step-5_en.pdf.
- M. A. Hernán and J. M. Robins, “Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available,” American Journal of Epidemiology, vol. 183, no. 8, pp. 758–764, 2016.
- Duke-Margolis Center for Health Policy, “Characterizing RWD Quality and Relevancy for Regulatory Purposes,” 2018.
- C. J. Jonker, E. Bakker, X. Kurz and K. Plueschke, “Contribution of patient registries to regulatory decision making on rare diseases medicinal products in Europe,” Frontiers in Pharmacology, vol. 13, 2022.
- Agency for Healthcare Research and Quality, USA, “Registries for Evaluating Patient Outcomes: A User's Guide,” September 2020. [Online]. Available: https://www.om1.com/download/5igKaCLgTRVHMLBiqXzsHW/registries-evaluating-patient-outcomes-4th-edition.pdf.
- E. I. Benchimol, L. Smeeth, A. Guttmann, K. Harron, D. Moher, P. T. Henrik, E. von Elm and S. M. Langan, “The Reporting of Studies Conducted Using Observational Routinely Collected Health Data (RECORD) Statement,” PLOS Medicine, 2015.
- S. M. Langan, S. A. Schmidt, K. Wing, V. Ehrenstein, S. G. Nicholls, K. B. Filion, O. Klungel, I. Petersen, H. T. Sørensen, W. G. Dixon, A. Guttmann, K. Harron, L. G. Hemkens and D. Moher et al., “The reporting of studies conducted using observational routinely collected health data statement for pharmacoepidemiology (RECORD-PE),” BMJ Journals, 2018.
- National Institute of Health, “NIH's All of Us Research Program is now the largest integrated genomics and health database in the world,” June 2026. [Online]. Available: https://www.nih.gov/news-events/news-releases/nihs-all-us-research-program-now-largest-integrated-genomics-health-database-world.
- US Food and Drug Administration, “Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims,” December 2009. [Online]. Available: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/patient-reported-outcome-measures-use-medical-product-development-support-labeling-claims.
- European Medicines Agency, “Milestone 5.15 Final validated Standards Tool for Registries in HTA prepared,” 2019. [Online]. Available: https://catalogues.ema.europa.eu/system/files/2025-02/05.01.0301%20Feasibility%20Documentation%20%20-%20Registry%20Evaluation%20and%20Quality%20Standards%20Tool%20%28REQueST%29%20_
%2010-Sep-2023_Redacted.pdf. - P. Dobay and M. Sabidó, “Evaluating Data Quality by Proxy: Can We Evaluate All Dimensions of the European Medicines Agency Data Quality Framework for Registry-Based Post-Authorization Safety Studies?,” Pharmacoepidemiology and Drug Safety, vol. 35, no. 2, 2026.
- US Food and Drug Administration, “Use of Natural Language Processing to Extract Information from Clinical Text,” June 2017. [Online]. Available: https://www.fda.gov/science-research/advancing-regulatory-science/use-natural-language-processing-extract-information-clinical-text-06142017.
- E. Stubbs, J. Exley, R. Wittenberg and N. Mays, “How to establish and sustain a disease registry: insights from a qualitative study of six disease registries in the UK,” BMC Medical Informatics and Decision Making, 2024.
- M. G. Kahn, T. J. Callahan, J. Barnard, A. E. Bauck, J. Brown, B. N. Davidson, H. Estiri, C. Goerg, E. Holve, S. G. Johnson, S.-T. Liaw, M. Hamilton-Lopez, D. Meeker and T. C. Ong et al., “A Harmonized Data Quality Assessment Terminology and Framework for the Secondary Use of Electronic Health Record Data,” The Journal for Electronic Health Data and Methods, vol. 4, no. 1, p. 18, 2016.
- R. J. A. Little, R. D'Agostino, M. L. Cohen, K. Dickersin, S. S. Emerson, J. T. Farrar, C. Frangakis, J. W. Hogan, G. Molenberghs et al., “The Prevention and Treatment of Missing Data in Clinical Trials,” The New England Journal of Medicine, vol. 367, 2012.
- T. L. Lash, M. P. Fox, L. C. McCandless and S. Greenland, “Good practices for quantitative bias analysis Get access Arrow,” International Journal of Epidemiology, vol. 43, no. 6, pp. 1969–1985, December 2014.
- K. El Emam, E. Jonker, L. Arbuckle and B. Malin, “A Systematic Review of Re-Identification Attacks on Health Data,” PLOS, 2015.
About The Authors:
Oznur Seyhun is the cofounder/managing partner of ECONiX Research. She is a healthcare executive and Ph.D. candidate with over 20 years of multinational experience across pharmaceuticals, biotech, gene therapy, medtech, and digital health. She specializes in RWE, oncology research, market access, healthcare economics, commercial strategy, and evidence generation. At Novartis, she served as the affiliate RWE champion, developing data strategies to support patient outcomes, payer value, and healthcare sustainability. She is also co-director of Market Access Today, advancing evidence-based healthcare and market-access strategies globally.
Anusha G. is a senior research analyst with five years of experience in market intelligence, procurement strategy, and consulting. Her work focuses on primary research, supplier landscaping, category strategy, sourcing optimization, and risk mitigation. She has supported leading Fortune 500 pharma clients with expert-validated insights that strengthen procurement decisions and supply chain resilience. Anusha holds bachelor’s and master’s degrees in biotechnology, complemented by three years of laboratory research experience.
Monica Nandagopal is a category research analyst with more than eight years of experience in market research and consulting. Her insights have enabled top pharma companies in their strategic decisions on supplier outsourcing, category management, and planning. In the past year, she has been engaged in more than 20 market-sourcing studies, more than 10 supplier-data visualizations, and multiple rapid-response analyses for clients for global and regional requirements.