![UNC Awarded Up to $35 Million[1.1] to Lead Landmark Initiative to Build World's Largest Data Resource for Rare Disease AI](https://news.unchealthcare.org/wp-content/uploads/sites/1159/2026/08/image-300x200.png)
More than 10,000 rare diseases affect an estimated 350 million people worldwide, including as many as one in 10 Americans.
For many patients and their families, obtaining an accurate diagnosis can take years. Delayed, incomplete, or incorrect diagnoses can have serious consequences, leading to inappropriate treatment, irreversible disease progression, and excessive medical costs.
To address the challenges of rare disease diagnosis, the UNC School of Medicine and Emory University are leading a first-of-its-kind research initiative to build a comprehensive data resource. Designed to power AI-driven insights, new and existing tools will enable analysis of these large-scale clinical and biological data to help rare disease experts make earlier diagnoses, better inform medical care, and accelerate treatment discovery.
"Our aim is to create a large-scale dataset spanning approximately 2,700 to train models that can then be deployed in diagnostic settings where we don't have such rich data," said Melissa Haendel, PhD, FACMI, who is the Sarah Graham Kenan Distinguished Professor in the Department of Genetics at the UNC School of Medicine and lead scientist on the project. "If we can build models based on the rich data, we can build clinical decision support tools that either diagnose patients in less sophisticated settings or help move those patients along."

Melissa Haendel, PhD, FACMI
The four-and-a-half-year project is supported by an up to $35 million award from the Advanced Research Projects Agency for Health (ARPA-H), an agency within the U.S. Department of Health and Human Services (HHS), through their Rare Disease AI/ML for Precision Integrated Diagnostics (RAPID) program. The UNC and Emory teams will lead efforts to acquire data from patient registries and real-world sources. Other program components will focus on direct patient data acquisition and on building a research platform for the broader rare disease community.
Even with advances in genomic testing, many rare disease cases remain difficult to diagnose because relevant clinical and research data is often scattered, inconsistent, or incomplete. A more connected data ecosystem is needed to help clinicians interpret findings more effectively and bring answers to patients sooner.
"Much of the information about rare disease patients isn't centrally available because each condition affects relatively few patients," said Richard Moffitt, PhD, co-lead on the project and associate professor in the Department of Hematology and Medical Oncology at the Emory University School of Medicine. "Challenges such as limited access to specialty care, insurance barriers, and geographic distance from rare disease experts hinder effort to collect high-quality data needed to understand disease patterns and outcomes."
Led by Haendel and Moffitt, the project draws on UNC and Emory's expertise in genetics, biomedical informatics, and rare disease research to bring together diverse sources of clinical and genetic information, including health records, insurance claims, medical imaging, video technologies, and patient surveys.
To protect patient privacy, names and any other identifying information will be removed before the data enters the RAPID resource. Secure access to the data will be tiered based on the sensitivity of the data, with some datasets available publicly and others requiring data use agreements and additional safeguards.
Qualified physicians and researchers around the world will be able to apply for access to support their own research initiatives. With more data, researchers will be better positioned to reveal how rare diseases develop and progress, find patterns that can support earlier diagnosis by non-experts, design stronger and more successful clinical trials, and accelerate drug development.
"At UNC, industry partners approach us regularly with trials on a rare disease," said Haendel, who is faculty at the UNC School of Data Science and Society. "The challenge is that finding eligible patients can be extremely difficult. We currently have no means to find those patients. By securely bringing together more data, we hope to develop algorithms to identify trial participants more efficiently, accelerate research, and expand access to clinical care."
The project team is supported by a robust public-private partnership that includes leading academic institutions, rare disease advocacy organizations and industry partners. Collaborators include Across Healthcare, Combined Brain, DartNet, Datavant, EB Research, Global Genes, Johns Hopkins University, Mendelian.co, the National Organization for Rare Diseases (NORD), Queen Mary University of London, Truveta, the University of California at San Francisco, and the University of Iowa.
Additional program support will be provided by OpenAI, Anthropic, Amazon Web Services, and Google. The National Institutes of Health's All of Us Center for Linkage and Acquisition of Data and the Monarch Initiative, which are co-led by Haendel, will also be involved on the project.
This research was, in part, funded by the Advanced Research Projects Agency for Health (ARPA-H). The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the United States Government.