Authors: Roni Haas (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Michael P. Margolis (1Department of Human Genetics, University of California, Los Angeles, USA; 5Department of Psychiatry, David Geffen School of Medicine, Los Angeles, USA; 6Department of Neurology, University of California, Los Angeles, USA), Angela Wei (1Department of Human Genetics, University of California, Los Angeles, USA; 7Interdepartmental Bioinformatics Program, University of California, Los Angeles, USA; 8Department of Pathology and Laboratory Medicine, University of California, Los Angeles, USA; 9Department of Computational Medicine, University of California, Los Angeles, USA), Takafumi N Yamaguchi (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Jeffrey Feng (10Department of Radiological Sciences, University of California, Los Angeles, USA), Thai Tran (6Department of Neurology, University of California, Los Angeles, USA), Veronica Tozzo (9Department of Computational Medicine, University of California, Los Angeles, USA), Katelyn J. Queen (11Department of Medicine, University of California, Los Angeles, USA), Mohammed Faizal Eeman Mootor (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Vishakha Patil (6Department of Neurology, University of California, Los Angeles, USA), Michael E. Broudy (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Paul Tung (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Shafiul Alam (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Danielle B. Martinez (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Yash Patel (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Nicole Zeltser (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Rupert Hugh-White (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Jaron Arbet (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Christa Caggiano (13Institute for Genomic Health, Icahn School of Medicine at Mount Sinai, New York, New York), Ruhollah Shemirani (13Institute for Genomic Health, Icahn School of Medicine at Mount Sinai, New York, New York), Mao Tian (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Prapti Thapaliya (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Lora Eloyan (12UCLA Health Information Technology, Office of Health Informatics and Analytics), Lawrence O. Chen (1Department of Human Genetics, University of California, Los Angeles, USA; 5Department of Psychiatry, David Geffen School of Medicine, Los Angeles, USA; 6Department of Neurology, University of California, Los Angeles, USA), Maryam Ariannejad (4Institute for Precision Health, University of California, Los Angeles, USA), Clara Lajonchere (4Institute for Precision Health, University of California, Los Angeles, USA; 6Department of Neurology, University of California, Los Angeles, USA), Bogdan Pasaniuc (14Department of Genetics Perelman School of Medicine University of Pennsylvania), Alex Bui (3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA; 10Department of Radiological Sciences, University of California, Los Angeles, USA), Valerie A. Arboleda (1Department of Human Genetics, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 8Department of Pathology and Laboratory Medicine, University of California, Los Angeles, USA; 9Department of Computational Medicine, University of California, Los Angeles, USA), Timothy S. Chang (6Department of Neurology, University of California, Los Angeles, USA), Noah Zaitlen (1Department of Human Genetics, University of California, Los Angeles, USA; 6Department of Neurology, University of California, Los Angeles, USA; 9Department of Computational Medicine, University of California, Los Angeles, USA), Paul T. Spellman (11Department of Medicine, University of California, Los Angeles, USA), Paul C. Boutros (1Department of Human Genetics, University of California, Los Angeles, USA; 2Department of Urology, University of California, Los Angeles, USA; 3Jonsson Comprehensive Cancer Center, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA), Daniel H. Geschwind (1Department of Human Genetics, University of California, Los Angeles, USA; 4Institute for Precision Health, University of California, Los Angeles, USA; 5Department of Psychiatry, David Geffen School of Medicine, Los Angeles, USA; 6Department of Neurology, University of California, Los Angeles, USA)
Categories: Article
Source: medRxiv
Authors: Roni Haas, Michael P. Margolis, Angela Wei, Takafumi N Yamaguchi, Jeffrey Feng, Thai Tran, Veronica Tozzo, Katelyn J. Queen, Mohammed Faizal Eeman Mootor, Vishakha Patil, Michael E. Broudy, Paul Tung, Shafiul Alam, Danielle B. Martinez, Yash Patel, Nicole Zeltser, Rupert Hugh-White, Jaron Arbet, Christa Caggiano, Ruhollah Shemirani, Mao Tian, Prapti Thapaliya, Lora Eloyan, Lawrence O. Chen, Maryam Ariannejad, Clara Lajonchere, Bogdan Pasaniuc, Alex Bui, Valerie A. Arboleda, Timothy S. Chang, Noah Zaitlen, Paul T. Spellman, Paul C. Boutros, Daniel H. Geschwind
Coupling genetic profiling with electronic health records from hospital biobanks is a foundational resource for precision medicine. However, lack of ancestral heterogeneity limits discovery and generalizability. We leveraged the UCLA ATLAS Community Health Initiative, a diverse biobank with >35% non-European participants in a single health system, to inform disease prevalence and genetic risk across five continental and 36 fine-scale ancestry groups. Analyzing clinical and genetic data for 93,937 individuals, 61,797 with whole-exome sequencing (WES), we identified novel associations between genetic variants and phenotypes, including STARD7 with asthma risk in Mexican Americans and FN3K with intestinal disaccharidase deficiency across Europeans and Admixed Americans. Top decile polygenic scores (PGS) predicted patient status for many common diseases (40% of patients with Type 1 diabetes); an effect markedly diminished in non-Europeans. Exploring the distribution of ACMG ClinGen rare variants across populations demonstrated European bias in curated clinical variants. Mitigating this bias using computationally predicted deleterious variants, we identified new gene-disease associations, including EXOC1L and blood glucose level in East Asians. We identified PTPRU as a modulator of semaglutide’s effects on weight loss, and additionally found variability across ancestries and a relationship with type-2-diabetes PGS. We provide an interactive web portal for accessing cross-ancestry associations at atlas-phewas.mednet.ucla.edu. Collectively, our findings support the value of ancestral diversity in advancing precision health across a broad spectrum of populations.
The integration of electronic health records (EHRs) with genetic data is transforming biomedical research, offering unprecedented opportunities for preventing and treating common medical conditions^1^. Broad, longitudinal sampling from deeply phenotyped individuals, linked with genetic, environmental, and lifestyle data, provides many advantages over standard cohort-driven research^2^. This ability has fueled the creation of nationwide biobanks, including the UK Biobank, All of Us^3^, FinnGen^4^ and the Taiwan Biobank^5^, as well as several large-scale academic biobanks, such as Mt. Sinai’s BioMe^6^; Vanderbilt’s BioVU^7^; Geissinger’s MyCode^8^; and the Michigan Genomics Initiative^9^. Integration of these biobanks has permitted innovative collaborative efforts, such as the eMERGE consortium^10^, COVID-19 host genomics initiative^11^, and the Global Biobank Initiative^12^.
Although these existing efforts have substantially advanced genetic and biomedical discovery, their concentration on European (EUR) ancestry participants underscores the importance of increasing ancestral diversity to enhance the applicability of findings across populations findings^13–17^. As PGS are ported into clinical use, the importance of measuring diverse populations is further emphasized by recent analyses showing the continuous, linear relationship between ancestral distance from the reference population and the utility of polygenic scores^16^. Similarly, rare genetic variation is known to have substantial ancestry-specific frequency and effect size (e.g., the reduced effect of APOE4 alleles on Alzheimer’s disease risk and increased frequency of rare PCSK9 variants in those with African ancestry). Moreover, the interpretation of clinically-relevant rare variation is hampered by the EUR bias in most current studies and genetic databases^18,19^ and including non-EUR populations reveals substantial disparities in clinically-relevant rare variant frequencies^20^ and increases power for discovery^21,22^. Thus, greater ancestral diversity in large biobanks with detailed medical records strengthens efforts to advance precision health^13,23,24^.
Previously, we described early phases of our work developing the UCLA ATLAS Community Health Initiative, a biobank embedded within the UCLA Health System consisting of deidentified genetic data coupled with EHRs. The goal of ATLAS is to advance precision medicine efforts in diverse populations, both locally in Los Angeles County and more broadly across California^25–27^. Here, we report an important milestone - discoveries from the current data freeze consisting of Single Nucleotide Polymorphism (SNP) genotyping for 92,165 and whole-exome sequencing (WES) for 61,797 individuals. We perform phenome-wide association in 84,110 individuals for 776 biomedical conditions, identifying 19,431 genome-wide associations. We replicate known common and rare variant associations for a range of biomedical conditions, including differences in effects across five major continental ancestries within this single biobank. We further investigate disease diagnoses and genetic risk in fine-scale ancestral clusters, identifying many new ancestry-specific risk associations. These analyses further demonstrate the utility of coupling EHR and genetic data in populations of diverse ancestries, from large continental (broad-scale) populations to, perhaps more importantly, regional and sub-continental (fine-scale) European, American, African and Asian populations in a single, unified health care system.
To date, the UCLA ATLAS initiative has enrolled >220,000 patients, with biomaterial availability for >130,000 patients, with data collection ongoing. ATLAS enrollment largely reflects the broader composition of patients in the UCLA Health system, primarily concentrated over a broad area on the west side of Los Angeles (Figure 1a) - one of the world’s most ancestrally diverse metropolitan areas with a population of 9.6 million. The EHR initiated in 2013 enables continuous longitudinal stratification of individuals into disease states, with a mean and median of 8.6 and 7.9 years of participation per individual. As of November 2024, the UCLA ATLAS Biobank of genomic data included genotype by sequencing (GxS) and Whole Exome Sequencing (WES) data from 61,797 individuals, and custom array genotyping using the Illumina Global Screening Array from 92,165 individuals, which underwent extensive quality control (QC) prior to analysis (Supplementary Figure 1; Methods). We leveraged these data, representing ATLAS data release two to interrogate social and genetic factors that affect disease risk and health outcomes, taking advantage of the unique composition of ATLAS enriched for understudied populations.
Genotyped cohort demographics are summarized in Figure 1 and Table 1. Overall, biobank participants had a higher comorbidity index compared to non-biobank adults (3.3 vs. 1.7 mean Elixhauser index; 1-year post-collection), consistent with higher number of clinical visits, which favors enrollment (Supplementary Table 1). Common conditions were endocrine/metabolic, cardiovascular, gastrointestinal, and neoplasms; many also had immunologic, neurologic, and neuropsychiatric conditions, reflecting global trends (Figure 1c; Supplementary Figure 2)^28–30^. The ATLAS EHR data consisted of a total of 71,739,582 lab tests (including complete blood count [CBC], lipid, metabolic, HbA1c, and 25-hydroxyvitamin D panels), with a mean of 540 lab test results per patient (Supplementary Table 2). Detailed prescription data is also available with a total of 5,952,958 prescriptions; the 50 most frequent medications are listed in Supplementary Table 3.
The ATLAS population is diverse relative to most other large health system biobanks and the research-oriented UK biobank^31^. Although we primarily consider genetically determined ancestry in our analyses, rather than the social constructs of race and ethnicity^32,33^, we note substantial diversity based on patient self-reports. For instance, 63.4% of patients self-identify as White, 11.2% self-identified as Asian, 4.87% as Black or African American, 2.4% as Middle Eastern or North African, 0.9% as American Indian or Alaska Native and 0.3% as Pacific Islander. A substantial proportion (14.6%) report Hispanic or Latino ethnicity (Table 1).
Self-identified race/ancestry, a cultural-societal construct, is often used as a proxy for genetic ancestry. But, the two are not synonymous^26,32,34^ and genetic ancestry must be considered to prevent confounding in genetic association studies^33,35^. We classified biobank participants into six broad-scale ancestry groups primarily corresponding to continents via PCA (Methods): European (EUR), African (AFR), South Asian (SAS), East Asian (EAS), Admixed American (AMR) and an Unclassifiable group, aligning with the 1000 Genomes Project super-populations (Figure 1d–f; Methods). Consistent with self-reported race/ethnicity, almost a third of the biobank participants (32%) were assigned to non-EUR genetic AFR, SAS, EAS, or AMR. There was significant agreement between self-reported race in the medical record and continental ancestry – 99% of those who self-identified as White were assigned EUR or AMR populations (Supplementary Figure 3a). However, PCA across common genetic variants demonstrates more granular relationships and provides a quantitative basis for assessing the relationship between ancestry and disease^26,34^. Of participants who self-reported their race as “unknown”, 46% were assigned to the AMR ancestry, and 45% to the EUR group.
We next asked how health system usage or disease burden varied by broad-scale ancestry, observing that the mean number of encounters varied significantly among ancestry groups (p value < 1×10^−16^, ANCOVA, adjusted for genetic sex and age), with the highest numbers of total hospital encounters in individuals from AFR ancestry, followed by AMR, with the lowest values in EUR, SAS, and EAS (Figure 1g). Another way to assess health burden is by a comorbidity index score (Elixhauser Comorbidity Index^36,37^), which was similarly distributed across most ancestries, with SAS individuals having the lowest scores (SAS 1.6; all other population groups 3.0–3.1; Figure 1h). Increased encounter numbers were only partially explained by elevated comorbidity index scores, indicated by modest correlation between the two (mean total encounters vs. Elixhauser Comorbidity Index: R = 0.32, P <2.2×10^−16^; mean hospital encounters vs. Elixhauser Comorbidity Index: R = 0.26, P <2.2×10^−16^; Supplementary Figures 3b-c).
The population diversity of ATLAS enables the interrogation of the combined genetic and social/environmental effects on the risk of disease diagnoses and exploration of medical comorbidity in depth. To illustrate this, we assessed variation in medical conditions and disease prevalence across population groups for several major conditions (Supplementary Figure 4a-f; Supplementary Table 4).
We first confirmed known associations between cancer and genetic ancestry, observing the highest risk of prostate cancer in those with AFR ancestry, stomach cancer in EAS, bladder cancer in AMR, and breast cancer in EUR^38^ (Supplementary Figure 4a). We also replicated several known cardiovascular disease associations. For instance, AFR individuals were more affected by hypertension and myocardial infarction, while those with SAS ancestry had the lowest risk of atrial fibrillation despite the highest incidence of coronary atherosclerosis (Supplementary Figure 4b) – an established but seemingly contradictory risk profile^39^.
Major metabolic disorders were more prevalent in AFR participants compared to other ancestries. This finding included diagnostic codes (9- and 10-International Classification of Diseases [ICD]) related to Type 2 diabetes, hypercholesterolemia, and hyperlipidemia, supported by laboratory measurements or vitals (Supplementary Figure 5a). Type 2 diabetes was more common in all non-EUR groups relative to EUR (Supplementary Figure 4c). Previous work has suggested that some of these differences reflect social and environmental, rather than genetic, factors^40^. With regard to neurological conditions, we find a higher risk of dementia in AFR participants, but a lower risk of migraines and Parkinson’s Disease in AFR compared to EUR patients, consistent with prior analyses^41–43^ (Supplementary Figure 4d). Among neuropsychiatric disorders, anxiety and major depressive disorders were the most frequent in EUR individuals, and least frequent in the continental Asian cluster, as previously reported (Supplementary Figure 4e)^44,45^.
We also identified previously unreported associations, such as a lower risk of epilepsy in EAS compared to EUR individuals (Figure 1i; odds ratio [OR] = 0.46 with 95% confidence interval [0.32, 0.65], Pbonferroni = 2.4×10^−3^), which prior studies did not detect^46,47^. Similarly, we found significantly reduced risk for bipolar disorder in those with AMR ancestry, clarifying conflicting findings in previous smaller studies^48,49^ (Figure 1i; OR = 0.47 [0.40, 0.56], Pbonferroni = 1.5×10^−16^).
We extend previous findings that sleep apnea is highest in AFR individuals but show a significantly lower risk of sleep apnea in EAS individuals relative to EUR (Figure 1h; OR = 0.67 [0.59, 0.75], Pbonferroni = 4.05×10^−9^), which has not been definitively demonstrated^50–53^. These associations illustrate the utility of ATLAS for cross-ancestry comparisons, including under-represented populations, within a single health system biobank to reduce potential confounding due to geographic effects.
Broad-scale ancestry encompasses multiple geographically and genetically distinct fine-scale populations, which in turn reflect recent geographic or socio-demographic stratification. When a fine-scale population arises from a small group of founders, deleterious rare alleles may become disproportionately prevalent over time and drive elevated rates of disorders like epileptic encephalopathy^54^ or liver disease^55^. Similarly, fine-scale populations may exhibit significantly different frequencies of common risk variants compared to large populations^35,56,57^, many of which are understudied^13–15^.
We have previously shown that fine-scale ancestral analysis reveals critical features of genetic and environmental contributions to health, including marked disparities in disease risk^13^. We therefore sought to apply our previous identity-by-descent (IBD) total shared length approach to construct fine-scale ancestries in a population nearly three times the size of our previous work^56^. Briefly, this approach identifies genomic regions shared between individuals due to a common ancestor and defines fine-scale ancestries, which we refer to as “clusters” (Methods^57,58^), based on the total amount of IBD segments shared between patients. We labeled each cluster with both a numeric identifier (e.g., IBD-01, IBD-02) and a corresponding cluster-specific name to facilitate interpretation, using reference populations and patients' self-reported demographic information, as we did previously^56^. With this approach, we identified 36 fine-scale ancestry clusters represented by at least 30 individuals (Figure 2a–b; Supplementary Figures 5b-c; Methods), replicating previously identified fine-scale clusters and adding new ones, such as a Native Hawaiian cluster (IBD-29; n = 64) and a Bantu cluster (IBD-35; n = 32). The largest clusters in our biobank consist of Northern Europeans (IBD-01; n = 33,675) and Southern Europeans (IBD-02; n = 14,841), making up 53.2% of our biobank. The remaining clusters represent the heterogeneity of ancestral origins in the Los Angeles community, such as Armenians (IBD-16, IBD-23, IBD-31; n = 560), Ashkenazi Jews (IBD-03; n = 14,261), Iranian Jews (IBD-11; n = 707), and Filipinos (IBD-09; n = 1,438).
For a nuanced understanding of diagnostic variation, we tested the prevalence of 1,253 phecodes^59–61^ across fine-scale clusters with at least 100 patients (Figure 2c; 23 tested clusters). As an initial measure of integrity, we tested for known increases in prevalence in these fine-scale ancestries, for example replicating an increase in breast cancer, irritable bowel syndrome, and Crohn’s disease in the Ashkenazi Jewish (IBD-03) cluster (Supplementary Table 5). Similarly, we observed the known increase in gout in the Filipino cluster^62^ and Alzheimer’s disease and dementias in the Puerto Rican cluster^63^. Consistent with the global ancestry findings and previous reports, fine-scale South Asian clusters had the highest coronary atherosclerosis risk, and the Japanese cluster had a higher prevalence of hypertension^64^. Cholelithiasis was widespread in clusters from South America and Mexico, while less frequent in the Northern European cluster (Supplementary Table 5).
We additionally uncovered numerous novel associations between our fine-scale clusters and disease (Supplementary Table 5; Figure 2c). For instance, the Asian, Filipino, and combined Chinese + Korean clusters showed a decrease in risk for vitamin B-complex deficiency (Filipino: OR = 0.50 [0.36, 0.67], FDR = 2.1×10^−4^; Chinese/Korean: OR = 0.6 [0.5, 0.7], FDR = 9.4×10^−11^). Cholesterolosis of the gallbladder was substantially more widespread in several Asian clusters, including, Filipino, combined Chinese/Korean, Southeast Asian, and one of the West Asian clusters (Filipino: OR = 4.3 [2.9, 6.2], FDR = 5.6×10^−12^; Chinese + Korean: OR = 2.3 [1.8, 3.1], 3.2×10^−8^; Southeast Asian: OR = 3.8 [1.4, 8.0], 1.9×10^−2^; West Asian OR = 3.2 [1.3, 6.3], 3.1×10^−2^), and less common in Northern EUR (OR = 0.6 [0.5 −0.7], 1.0×10^−5^). Mexican and South American clusters suffered considerably more from hormones’ adverse effects in therapeutic use (Mexican American OR = 2.0–2.8 [1.3, 2.3], FDR = 1.6×10^−2^ - 2.1×10^−24^; combined South American OR = 1.9 [1.4, 2.5], 1.3×10^−4^). Glaucoma was common in the African American (IBD-06) and East Asian clusters (Chinese + Korean (IBD-05), Japanese (IBD-10)), as is already known^65^, but also in the Iranian Jewish cluster (OR = 1.9 [1.5–2.4], FDR = 2.9×10^−6^), which has not been reported before. Iranian and Ashkenazi Jewish clusters showed the highest risk for hyperplasia of prostate (Iranian Jewish: OR = 2.3 [1.6, 1.8], FDR = 1.1×10^−13^; Ashkenazi Jewish: OR = 1.7 [1.6, 1.8], FDR = 4.6×10^−66^) and bladder cancer (Iranian Jewish: OR = 1.5 [1.1, 1.9], FDR = 0.008; Ashkenazi Jewish: OR = 3.1 [1.4, 5.6], FDR = 0.02), and the lowest risk for cirrhosis of liver (Iranian Jewish: OR = 0.2 [0.03, 0.5], FDR = 3.0×10^−2^; Ashkenazi Jewish: OR = 0.4 [0.3, 0.5] , FDR = 6.8×10^−24^).
We next focused on cardio-metabolic diseases due to their high prevalence in ATLAS and global impact on public health, comparing disease risk for each fine-scale cluster within the same broad-scale continental ancestry (Methods; Figure 2d). Strikingly, among Asian clusters, the Filipino cluster had elevated risk for all tested medical conditions, which remained after adjusting for body mass index (BMI) (essential OR =1.6 [1.3, 1.7], FDR = 9.0×10^−11^; Type 2 OR = 1.4 [1.6, 2.1], FDR = 7.1×10^−6^; Coronary OR = 1.4 [1.1, 1.7], FDR = 5.3×10^3^; abdominal aortic OR = 2.8. [1.4, 5.3], FDR = 7.1×10^−3^ Hyperlipidemia: R = 1.2 [1.0–1.4], FDR = 2.4×10^−2^. Firth's bias-reduced logistic regression adjusted for BMI, sex and age).
Across fine-scale EUR ancestry clusters within ATLAS (Ashkenazi Jewish, n = 14,262; Iranian, n = 363; Iranian Jewish, n = 707; two Armenian clusters, n = 382 and 134; Lebanese, n = 183; Egyptian Christians, n = 107), the larger Armenian cluster (IBD-16) showed an elevated risk for the range of cardiometabolic conditions including Type 2 diabetes (OR = 2.0 [1.5–2.7], FDR = 2.5×10^−6^) and essential hyperlipidemia (OR = 1.6 [1.2–2.0], FDR = 8.5×10^−4^). Iranian clusters, both Jewish and non-Jewish, had a relatively higher risk of coronary atherosclerosis (Iranian Jewish: OR = 1.8 [1.5, 2.2], FDR = 4.4×10^−8^; Iranian cluster OR =1.9 [1.4, 2.5], FDR = 2.8×10^−5^), and Type 2 diabetes (Iranian Jewish: OR = 2.4 [1.9, 2.9], FDR = 6.6×10^−16^; Iranian OR = 2.2 [1.6, 2.8], FDR = 1.7×10^−6^), but not hypertension. As is known, among AMR fine-scale clusters, those in the Mexican American cluster had a high risk of Type 2 diabetes, and combined South Americans showed higher risk for hyperlipidemia (Supplementary Figure 5d).
We next leveraged genotyping in our cohort to calculate individual PGS for a range of cancer, cardiovascular, metabolic, neuro-psychiatric and autoimmune diseases (Methods) and tested their relationship to disease risk. We focused on the top end of the PGS distribution (10%), compared with the 5^th^ decile, in EUR individuals (Figure 3; Supplementary Table 6). For Type 1 diabetes, the top PGS decile was the most enriched (OR = 11.7 [7.6–19.0], FDR = 1.3×10^−25^), with 41% of the diagnosed patients in ATLAS assigned to the top PGS decile. The second most enriched trait was Crohn's disease, with 33% of diagnosed patients within the top PGS decile (OR = 5.5 [4.0–7.7], FDR = 2.8×10^−24^), followed by gout with 27% (OR = 5.11 [4.1–6.4], FDR = 2.3×10^−46^), testicular cancer with 25% (OR = 3.7 [1.2–7.3], FDR = 8.1×10^−5^) and prostate cancer with 21% (OR = 3.4 [2.8–4.0], FDR = 3.64×10^−46^) (Supplementary Figure 6a-e). On average, the top decile of risk accounted for 18.5% of diagnosed patients across 28 tested disorders. Twenty-five of 28 (89%) showed significant enrichment of patients within the top PGS decile (mean OR for significant tests = 2.9). Performance declined when we applied PGS to non-EUR populations, as expected^55–58^. This was due to both reduced sample sizes and the model fit (Supplementary Figure 6f-i), identifying only 16.7, 27.3, 36.4, and 38.5% significant associations of diseases with the top PGS decile, for SAS, AFR, AMR, and EAS, respectively. Concordantly, the top PGS decile accounted for fewer cases (on average, 12.6%, 15.0%, 15.5% and 15.9%, for SAS, AFR, AMR, and EAS). This further supports the need for larger, more ancestrally heterogeneous cohorts for clinical development of PGS^55–58^.
We next sought to understand how genetic risk for common, complex traits was distributed across broad- and fine-scale ancestry. Through phenome-wide association analysis (PheWAS) of 84,110 genotyped individuals using the Regenie framework^70^ (Methods), with analyses run separately for cases/controls within each cohort, we refined numerous known associations and identified novel associations that represent cohort-specific disease risk. After LD pruning, we identified a total of 19,431 unique variant-phenotype associations passing a genome-wide significance threshold of 5×10^−8^ (Supplementary Table 7), with 5,772 passing the most conservative Bonferroni threshold of 6.4×10^−11^ (5×10^−8^ conditioned on 776 tested phenotypes). We provide a web portal at atlas-phewas.mednet.ucla.edu to facilitate browsing of these associations.
First, we examined the distribution of APOE alleles across fine-scale clusters, highlighting the increased frequency of ε4 risk alleles in the African American (IBD-06) and Bantu (IBD-35) clusters (Figure 4a). We replicate the finding that African American individuals with ε4/ε4 haplotype have a lower risk of Alzheimer’s disease compared to other ancestral cohorts (Supplementary Figure 8). We additionally replicated two well-known genetic associations in the African American (IBD-06) cluster, between HBB rs334-A (Figure 4b; MAFIBD-06 = 5.21%) and a diagnosis of sickle cell anemia (PIBD-06 = 1.96×10^−78^; ORIBD-06 = 15.08) and between the Duffy null ACKR1 rs2814778-C polymorphism (Figure 4b; MAFIBD-06 = 77.06%) and a decrease in neutrophil count (PIBD-06 4.00×10^−31^; ORIBD-06 0.67). Using quantitative blood work findings from ATLAS, we also profiled the impact of rs334-A on mean corpuscular hemoglobin concentration (PIBD-06 = 5.21×10^−15^; βIBD-06 = 0.41), nucleated red blood cell count (PIBD-06 = 1.59×10^−10^; βIBD-06 = 0.16), and mean corpuscular volume (MCV) (PIBD-06 = 3.27×10^−8^; βIBD-06 = −0.29). Both risk variants have been concluded to be under strong selection for resistance to malaria^71^.
Across multiple Asian clusters, we confirmed the association between a variant in high LD with the --^SEA^ deletion, which causes inherited alpha-thalassemia^72^, and microcytic anemia. We observed a significant decrease in MCV in response to LUC7L rs372755452-A in the broad-scale EAS ancestry (PEAS = 6.10×10^−88^; βEAS = −1.64; MAFEAS = 0.91%), and the fine scales Chinese + Korean (PIBD-05 = 2.71×10^−52^; βIBD-05 = −1.58; MAFIBD-05 = 0.79%) and Filipino (PIBD-09 = 2.15×10^−24^; βIBD-09 = −1.77; MAFIBD-09 = 1.21%) clusters. We also replicated the finding that AMR individuals are twice as likely to carry a PNPLA3 rs738409-G missense variant (Figure 4b) that greatly increases the risk for non-alcoholic fatty liver disease, a major cause of cirrhosis that often necessitates liver transplant^73^.
We extended these findings by identifying significant associations between rs738409-G and nonalcoholic cirrhosis of liver in two Mexican American clusters (Figure 4c), IBD-04 (PIBD-04 = 2.66×10^−14^; ORIBD-04 = 1.90) and IBD-07 (PIBD-07 = 1.46×10^−8^; βIBD-07 = 1.95), with differing allele frequencies across these cohorts (MAFIBD-04 = 45.00%; MAFIBD-07 = 52.93%). In the Northern European cluster, this variant has a smaller effect size (PIBD-01 = 3.62×10^−09^; ORIBD-01 = 1.53) and is half as frequent (MAFIBD-01 = 22.88%). We identified a previously unreported association between STARD7 rs17419569-C, an intronic variant, and asthma in a Mexican American cluster (Figure 4d; PIBD-04 = 1.94×10^−9^; ORIBD-04 = 2.62; MAFIBD-04 = 2.49%). Prior work suggests that decreased STARD7 expression is associated with enhanced allergic responses in the human lung and significant increases in airway hyperresponsiveness in haploinsufficient Stard7 mice, supporting this discovery^74^. In the same cluster, we identified a novel association between gastrointestinal reflux disease (GERD) and rs74744741-C (Figure 4e; PIBD-04 = 5.79×10^−10^; ORIBD-04 = 1.63; MAFIBD-04 = 8.92%), nominating GPX7 and SHISAL2A as risk genes. GPX7 has previously been associated with carcinogenesis in the context of GERD-associated Barrett’s esophagus^75^, while SHISAL2A has unknown function, but is highly expressed in the small intestine and in lymphoid tissues^76^.
We further identified two low frequency variants associated with chronic renal failure in the AMR cohort (Figure 4f), rs112680741-C (PAMR = 2.94×10^−11^; ORAMR = 3.78; MAFAMR = 1.24%) and rs2744548-C (PAMR = 9.59×10^−11^; ORAMR = 3.17; MAFAMR = 1.72%), nominating GPLD1, ALDH5A1, and KIAA0319 as potential risk genes. None have previously been associated with kidney disease, although ALDH5A1 and KIAA0319 are highly expressed in the kidney^76^.
Lastly, we identify a new common variant association for intestinal disaccharidase deficiency (Figure 4g), the inability to completely digest sugars such as lactose and sucrose, which is a common cause of IBS-like symptoms such as persistent abdominal pain and diarrhea^77^. Across four EUR cohorts (PEUR = 3.80×10^−41^; OREUR = 1.26; MAFEUR = 33.26%) and the AMR cohort (PAMR = 4.90×10^−11^; ORAMR = 1.31; MAFAMR = 42.33%), we associate a high-frequency FN3K rs7208565-T variant with increased risk for intestinal disaccharidase deficiency. FN3K phosphorylates glycated proteins, preventing the formation of advanced glycation end-products (AGEs), and has been implicated in a wide range of conditions^78^.
The ATLAS WES catalog comprises more than 12 million autosomal variants, with a median of 9,821 missense and 149 loss of function (LOF) variants per individual (including both rare and common variants), closely matching expectations from other studies^79^. Most WES variants were rare (MAF < 1% in any broad-scale ancestry; n = 11,808,484), including 2,724,021 rare missense and 155,006 rare LOF variants (Table 2), with a median of 15 rare LOF and 403 rare missense variants per participant (Supplementary Table 8). These values varied significantly across ancestry groups (Supplementary Table 9), with AFR individuals showing the most rare or common LOF and synonymous variants (180 and 11,550), matching prior results^80^.
We selected several clinically relevant rare variants with known elevated frequencies in specific ancestral clusters and confirmed that these patterns were replicated in ATLAS. For instance, Ashkenazi Jewish (IBD-03) had the primary risk for carrying either pathogenic or likely pathogenic (P/LP) variants in BRCA1 or BRCA2 (BRCA1: OR = 47.143 [20.6, 133.0]; BRCA2: OR = 48.2 [23.6, 114.2]) (Figure 5a). Similar to past reports^81^, 4.6% of breast cancer and 11.8% of ovarian cancer patients carry either P/LP variants in BRCA1 or BRCA2. The three BRCA Ashkenazi Jewish founder alleles were the most prevalent in EUR, and particularly the Ashkenazi Jewish cluster (Supplementary Figure 9a-b). Additionally, we identified 753 carriers of variants associated with risk for Familial Mediterranean Fever (FMF), with the highest carrier frequencies in the Armenian clusters^82,83^, and the next highest carrier frequencies in West Asian, Lebanese, and Egyptian Christian clusters (Supplementary Figure 9c). Carriers of these variants had a higher risk of the amyloidosis phecode (OR = 1.3 [0.8–1.8]; Supplementary Figure 9d). The HBB:p.E7V variant, which is responsible for most sickle cell anemia cases, was carried by 273 participants, with elevated frequency in the African American (IBD-06) cluster (OR = 51.4 [39.7, 67.0]; Supplementary Figure 9e), in line with previous findings^84^. Finally, we identified 49 carriers of loss-of-function variants in the PCSK9 gene causing lowered LDL levels, based on curated protective variants^85^ within the African American (IBD-06) cluster (OR = 9.1 [4.5–17.4]) (Supplementary Figure 9f).
Given the ability to replicate known findings, we next tested ATLAS' unique ancestral heterogeneity to detect new enrichments of known rare clinically relevant variants across fine-scale clusters. We queried the full set of the curated rare P/LP ClinGen^86^ variants, which have strong clinical and genetic evidence (n = 643), identifying 5,223 unrelated carriers for ClinGen P/LP variants. Within the Filipino (IBD-09) cluster, we detected an elevated carrier frequency for variants within the Tier 1 gene, LDLR, that causes familial hypercholesterolemia (OR = 3.9 [1.4–8.8], FDR = 0.03; Figure 5a). This finding has interesting implications given the high prevalence of dyslipidemia in Filipinos^87^ (Figure 2d). Additionally, we identify an elevated carrier frequency of variants that cause non-syndromic genetic deafness within two different genes in two separate (1) the MYO15A gene in our Central American 1 (Mexican) (IBD-04) cluster (OR = 6.9 [2.6–17.4], FDR = 8.44 × 10^−4^; Figure 5a) and (2) the CDH23 gene in our East Asian 1 (Han Chinese/Korean) (IBD-05) cluster (OR = 6.0 [2.0–15.6], FDR = 6.04 ×10^−3^; Figure 5a) and East Asian 2 (Japanese) (IBD-10) cluster (OR = 28.1 [8.9–77.1], FDR = 7.38×10^−6^) that have not been previously reported. We also identified previously unreported elevated carrier frequencies in the Ashkenazi Jewish (IBD-03) cluster for Pendred syndrome variants in the SLC264A gene (OR = 3.8 [2.9–5.0], FDR = 8.61×10^−19^) and recombinase activating gene 2 deficiency in the RAG2 gene (OR = 2.9 [1.4–5.8], FDR = 0.02). This illustrates the utility of fine-scale ancestry clustering in identifying populations at higher risk of rare, monogenic disease.
Finally, we focused on the clinically actionable American College of Medical Genetics and Genomics (ACMG) genes secondary findings (SF) v3.2^88^. We asked if the total allele frequencies of ACMG ClinGen P/LP variants vary across broad- and fine-scale ancestries. Overall, 17 ACMG genes had at least one P/LP variant in ClinGen. For a more nuanced biological understanding, we divided the ACMG variants into two groups of rare LOF (n = 53) and rare P/LP missense (n = 131) variants (we defined ‘LOF’ for variants ranked as high-confidence LOF by loss-of-function transcript effect estimator [LOFTEE]^89^, which accounts for biological context beyond truncating variant identification, and ‘missense’ based on the VEP^90^ "missense_variant" annotation; see Methods). We observed the highest frequency of rare P/LP LOF variants in EUR individuals among broad-scale groups, and in the Ashkenazi Jewish cluster among fine-scale clusters (Supplementary Figure 9g). The numbers of total rare P/LP LOF alleles (as defined based on broad-scale groups) in these populations were by far the highest compared to all others (OREUR = 3.7 CIEUR = 2.6–5.4, PBonferroni = 6.3×10^−17^; ORAshkenazi Jewish = 6.5, CIAshkenazi Jewish = 5.3 – 8.1, PBonferroni = 3.5×10^−62^; Figure 5b). No differences in rare P/LP missense total counts were identified at the level of broad-scale ancestries, but Northern EUR individuals had significantly more at the level of fine-scale clusters (ORNorthern EUR = 1.4, CINorthern EUR = 1.15 – 1.8, PBonferroni = 0.02). This replicates a bias toward EUR ancestry for rare curated clinical variants, as was recently shown in All of Us^20^, but we demonstrate this in a single health system.
As clinical variant databases are primarily ascertained from EUR cohorts, we took a conservative approach to call putative P/LP missense and LOF rare variants, using a consensus of a majority (5 of 9; Methods) of state-of-the-art computational methods based on the Critical Assessment of Genome Interpretation (CAGI) project^91^ for pathogenicity prediction (Methods). Keeping our focus on ACMG reportable genes^88^, this yielded 1,602 rare LOF variants and 6,487 rare predicted damaging missense variants within this ACMG gene set. We emphasize that these are distinct from the approach taken with ClinVar variants, as variants identified in this manner do not necessarily have evidence supporting their association with disease but have the advantage that they are not biased by oversampling EUR cohorts.
We then tested differences in the overall abundance of these variants across populations. Among broad-scale populations, EAS individuals carried significantly more predicted damaging rare missense variants compared to all other groups. AFR individuals had substantially more rare LOF variants compared to other ancestries (Figure 5c–d); and EUR clusters had the lowest rare predicted damaging missense, which is consistent with published data^92^. However, among fine-scale ancestries (n>100), both the fine-scale EUR Iranian Jewish and AFR clusters had significantly higher numbers of rare LOF (as defined based on broad-scale groups). Rare predicted damaging missense alleles were also significantly more common in Iranian Jewish, along with the EAS fine-scale Filipino and Chinese + Korean clusters (Figure 5c). In contrast, the Japanese cluster carried the smallest number of rare predicted damaging missense variants in this ACMG gene list (Figure 5c–d). These observations contrast in comparison with rare variants identified in ClinGen (Figure 5b), consistent with the interpretation that ClinGen variant frequencies are biased because they are based primarily on a majority of EUR patients, as is the case for most curated clinical data sets.
To examine the influence of predicted deleterious rare LOF (dLOF) and missense coding variants (dMIS) on ATLAS phenotypes (Methods), we performed exome-wide association studies (ExWAS) across 17,537 protein-coding genes. Using Regenie’s unified gene burden association strategy (“GENE_P” test^70^), we identified 1,099 unique gene-trait associations at a significance threshold of 2.85×10^−6^ (0.05/17,537 protein-coding genes). Within the EUR cohort, we replicated numerous known associations – 45 of the top 50 associations have been previously reported^79^ (Supplementary Table 10). These include PKD1 with cystic kidney disease (Figure 5e; PEUR,LOF = 9.41×10^−42^, OREUR,LOF = 107.14 [58.14, 197.43]), TTN with primary intrinsic cardiomyopathy (PEUR,LOF = 2.95×10^−14^, OREUR,LOF = 4.52 [3.18, 6.41]), and JAK2 with polycythemia vera (PEUR,MIS = 1.14×10^−45^, OREUR,MIS = 27.42 [18.25, 41.18]).
In the Northern European (IBD-01) cluster, we detected a novel gene-level association between HNRNPA1L2 and acquired absence of the breast/breast cancer (GENE_PIBD-01 = 1.00×10^−25^; GENE_PIBD-01 = 2.21×10^−8^; Supp Table 10). Though HNRNPA1L2 has not been previously associated with breast cancer, it is highly expressed in breast invasive carcinoma^76^ and a paralog HNRNPA1 was previously implicated in breast cancer progression^93^. In the Ashkenazi Jewish (IBD-03) cluster, we also find novel associations between EPG5 and HDL cholesterol level (PLOF_MIS = 3.55×10^−10^; βLOF_MIS = 1.77 [1.22, 2.33]) as well as triglyceride level (PLOF_MIS = 1.33×10^−8^; βLOF_MIS = −1.81 [−2.44, −1.18]). EPG5 is an autophagy tethering factor classically associated with Vici syndrome^94^, a severe developmental disorder, though our findings suggest an additional role in lipophagy and lipid metabolism^95^.
In the AMR population, we identified a novel association between dLOF and dMIS variation in CLN3 and cystic kidney disease (GENE_PAMR = 8.67 ×10^−9^), supported by experimental evidence that CLN3 is highly expressed in medullary collecting duct principal cells and plays a role in osmoregulation^96^. We additionally identified novel associations between PPARG and viral pneumonia (PAMR,MIS = 1.06×10^−6^, ORAMR,MIS = 52.96 [13.45, 208.48]) as well as NADSYN1 and abnormal lung examination findings (PAMR,MIS = 1.24×10^−5^, ORAMR,MIS = 3.73 [2.13, 6.51]). PPARG is known to regulate macrophage response to pulmonary inflammation^97^, while NADSYN1 knockout in mouse models was shown to cause abnormal lung development^98^, though neither has been previously associated with respiratory traits in patients.
Within the AFR cohort and African American (IBD-06) cluster, we uncovered novel associations between DDHD2 and dysphagia (PIBD-06,LOF = 1.24×10^−6^, ORIBD-06,LOF = 8.29 [3.68, 18.71]) as well as EFCAB13 and essential hypertension (PAMR,MIS = 9.31×10^−7^, ORAMR,MIS = 13.07 [4.53, 37.67]). These findings are consistent with prior studies – mutations in DDHD2 have been shown to cause hereditary spastic paraplegia and symptoms of dysphagia^99^, though not in African Americans, while methylation studies have prioritized EFCAB13 as a risk factor for heart failure^100^. We detected an additional association between ANKZF1 and peripheral vascular disease (PAMR,LOF_MIS = 1.56×10^−6^, ORAMR,LOF_MIS = 58.83 [12.58, 275.12]), supported by a recent finding that ANKZF1 acts downstream of HIF-1α to promote angiogenesis in the human mesenchymal stem cells under hypoxic conditions^103^.
Next, in the EAS cohort, we observed a novel association between EXOC1L and glucose mass volume in plasma (PEAS,LOF = 3.76×10^−7^, OREAS,LOF = 6.45 [3.14, 13.22]). The exocyst complex, including EXOC1 through EXOC8, has previously been identified as a regulator of insulin-stimulated glucose uptake in skeletal muscle cells^101^, although EXOC1L has not been previously associated with insulin resistance or blood glucose. We further identified an association between KIF2B and nephritis/nephropathy in the Chinese/Korean (IBD-05) cluster (PIBD-05,MIS = 7.18×10^−7^, ORIBD-05,MIS = 26.18 [8.97, 76.41]); though KIF2B has not been previously associated with renal traits, other genes in the kinesin (KIF) superfamily have been implicated in renal pathology, such as KIF2C with clear cell renal cell carcinoma^102^ or KIF3A with polycystic kidney disease^103^.
Finally, we examined known risk associations to identify ancestry-specific patterns of deleterious rare variation. We examined two known GBA1 rare LOF mutations that increase risk for Parkinson’s disease^104^, p.Glu365Lys and p.Thr408Met, in the AMR and EUR cohorts. Despite similar allele frequencies within each cohort, p.Glu365Lys was only significantly associated with increased risk in EUR (PEUR,E365K = 0.014; PAMR,E365K = 0.75), while p.Thr408Met was only associated with increased risk in AMR (PEUR,T408M = 0.212; PAMR,T408M = 0.0029), displaying evidence of ancestry-specific risk stratification (Figure 5f). We also found that the AMR cohort exhibited greater decreases in LDL cholesterol level, in response to additive LOF and deleterious missense risk variation in PCSK9 (OREUR,LOF = 0.51 [0.35, 0.73]; ORAMR,LOF = 0.33 [0.19, 0.57]), compared to the EUR cohort (Figure 5g). Further, LDL variants increase LDL in an ancestry indifferent manner, as is known^105^ (Supplementary Figure 9h). These results underscore the importance of interrogating rare variants across different ancestries to capture cohort-specific genetic risk, with potential for guiding future precision medicine-based clinical interventions.
One of the main advantages of ATLAS is the consideration of dynamic changes over time. To illustrate this, we integrated all four study components (demographics, genetic ancestry, common and rare genetic variants) to study the efficacy of GLP-1 receptor agonists (GLP-1 RAs) for weight loss. We focused on semaglutide, which had the greatest number of prescriptions in the database (Supplementary Figure 10a-b), with 7,214 individuals having recorded a prescription (Figure 6a). As a general tendency, we observed a steady decrease in weight up to 60 weeks of semaglutide treatment, consistent with previous findings^106–108^ (Supplementary Figure 10c). We tested if the medication dose, route, sex, age and initial weight affect semaglutide efficacy, corroborating that dose and subcutaneous delivery versus oral delivery were positively correlated with weight loss as has been well established (Figure 6b; linear mixed PBonferroni~ = 3.2×10^−62^, rsq (explained variance) = 0.71, effect size = −0.6 ; route oral vs. P = 2.5×10 ^−20^, rsq = 0.65, effect size = 1.9). We did not detect a significant effect of patient age and sex on efficacy.
Next, we asked if semaglutide’s effects varied by broad-scale ancestry over the first 60 weeks of treatment. Ancestry and its interaction with time significantly influenced weight loss (Figure 6c; ANOVA on a linear mixed-effects model; PBonferroni = 0.002, Sum Sq = 120.5; ancestry×time: PBonferroni = 0.0005, Sum Sq = 185.20). Further comparisons between EUR and the other populations revealed less weight loss in AMR, and a slower rate of weight loss in AMR and EAS (Figure 6c; linear mixed model; ancestryAMR: PBonferroni = 2.6×10^−2^, estimate = 0.8; ancestryAMR× = 4.6×10^−3^ ,estimate = 45.6; ancestrytime: PBonferroniEAS × = 3.2×10^−3^, estimate = 76.4).time: PBonferroni
Given limited evidence on the role of inherited factors in GLP1-RAs effectiveness^109^, we used ATLAS to explore the contribution of common genetic variation to semaglutide-related weight loss. We first tested whether genetic scores related to BMI or Type 2 diabetes mellitus (DM2) could influence semaglutide efficacy and found no correlation between BMI PGS and semaglutide efficacy. However, we observed that weight loss was negatively correlated with DM2 PGS (Figure 6e; linear mixed model, PGSHigh vs. Low: PBonferroni = 2.8×10^−4^, effect size = 0.96; PGSMed vs. Low: PBonferroni = 2.8×10^−3^, effect size = 0.73). The same relationship was observed when simplifying the model using methods similar to previous approaches^109^, using the maximum weight loss record for each participant (linear regression, PBonferroni = 1.3×10^−2^, beta = 0.31; Supplementary Figure 10d). As a second step, we conducted a genome-wide association (GWAS) meta-analysis to combine data across ancestries, but did not identify any significant loci (Methods; Supplementary Figure 10e).
Finally, we used WES to test if we could identify rare variant associations of semaglutide response within our cohort (Regenie^70^; Methods). Given the modest sample size, we focused on the EUR cluster and limited our test to only proteins whose plasma abundance was recently shown to be altered by semaglutide^110^. We identified a Bonferonni-corrected significant association of weight loss on semaglutide with one gene, PTPRU (PBonferroni = 0.0047, beta = −0.834) (Figure 6f; Supplementary Figure 10f). The negative effect size indicates that PTPRU activity is negatively associated with weight loss on semaglutide (variants affecting the function of PTPRU contribute to weight loss on semaglutide), concordant with the finding of decreased PTPRU levels as a response to semaglutide treatment^110^. This association involves 37 variants, of which one is common (rs2235937; nominally associated with weight loss (P = 9.7×10^−3^, beta = −0.06, SAIGE^111^ in EUR), and the rest are rare (Supplementary Table 11). The overall frequency of rare variants in PTPRU varied across ancestries (Supplementary Figure 10g). This gene has no previous known functional relationship to weight loss or related metabolic functions. However, combined with the strong data from serum proteomics^110^, our genetic analysis nominates this protein kinase and its pathways as novel candidates worth further investigation.
Understanding of the causes of a wide variety of biomedical conditions to improve healthcare outcomes and reduce the burden of disease has been made possible by large scale population studies, starting with the Framingham Heart study^112,113^. Health system biobanks continue to advance this line of research by providing cost effective, longitudinal data in large population cohorts^6,7,9,31,57,79,114^.
Despite the known sources of errors and incomplete phenotyping in health system EHRs, multiple studies have shown that many of these factors can be mitigated, and robust analyses can be conducted using these data^7,8,31,115^ Indeed, we were able to validate previously identified genetic associations tested in the UCLA ATLAS Community Health Initiative biobank based on EHR-derived phenotypes. This includes the application of polygenic scores, which we show can have high utility in identifying patients at high risk for common disorders, such as Type 1 Diabetes Mellitus, where > 40% of our cases lie in the top PGS decile. However, this predictive power, which is derived from EUR PGS, decreases dramatically in those with non-EUR ancestry across phenotypes, further emphasizing the well-articulated need for ancestral diversity in genetic studies of human disease^66–69^. We also leverage the heterogeneity of genetic ancestries in our population, > 30% of which are non-European, to validate and extend our knowledge of disease burden and genetic risk factors in ancestrally distinct populations.
Prior work has mostly relied on utilizing broad-scale genetic ancestry to understand how ancestry contributes to health risks^3,27,114^. Genetic ancestry can be measured on a continuum. While broad-scale ancestry discerns more ancient genetic variation via PCA, fine-scale ancestry discerns more recent genetic variation via shared IBD^116^. Both resolutions of genetic ancestry are useful for identifying groups of individuals with differing disease risk, but different aspects of disease risk are identified depending on the resolution of genetic ancestry leveraged^56,57^. Identifying fine-scale ancestry clusters within our biobank has leveraged its diversity to enable findings that will benefit populations that are underrepresented in genetics and medical research, with the potential to improve ancestry-targeted screening for diseases^117^. For example, to our knowledge UCLA ATLAS consists of the largest known genetic data set of ancestrally Filipino individuals (n = 1,438), with the next largest genetic data set consisting of 1,028 individuals^118^. We identified novel findings within our Filipino cluster, which can be used to improve preventative and precision care for this population.
Through PheWAS across broad- and fine-scale ancestries, we identify numerous novel genetic associations, including rs17419569-C with asthma risk; rs74744741-C with GERD risk in the Mexican American (IBD-04) cohort; and rs7208565-T with intestinal disaccharidase deficiency across multiple EUR and AMR cohorts. By conducting ExWAS using burden masks of rare, predicted deleterious LOF and missense variants, we prioritize novel risk genes, including CLN3 for cystic kidney disease in AMR, EXOC1L for blood glucose level in EAS, and HNRNPA1L2 for breast cancer in the Northern European (IBD-01) cohort. Using our fine-scale ancestral clusters, we reveal differences in risk frequency – for instance, the combined South American (IBD-07) cohort carries the non-alcoholic cirrhosis risk variant PNPLA3 rs738409-G at a higher frequency than the Mexican American cohort where this association was first discovered. Discerning common and rare variant risk frequency, as well as effect size, across fine-scale ancestral cohorts poses a powerful tool for precision medicine – offering opportunities to optimize risk stratification, alter screening guidelines, and tailor therapeutic interventions to individuals based on their ancestry-specific risk factors.
It has been well demonstrated that diagnostic misclassification due to differences in variant frequencies across ancestries is a serious risk in homogeneous populations^19^. We emphasize that having such diversity within a single regional biobank can help mitigate against confounders, not only caused by ancestral genetic differences, but non-genetic variables that differ across countries or geographically distinct biobanks. We demonstrated the application of longitudinal EHR to gain deeper insights into health outcomes over time. We showed that weight loss on semaglutide varies by broad-scale ancestry and genetically aligns with DM2 PGS. Leveraging omics data, we display the influence of genetic variants in a semaglutide-affected candidate protein on weight loss. These were enabled by integrating dynamic EHR changes with detailed prescription information, including the medication dose and route. Lastly, our analysis of ACMG clinical variants further supports that preponderance of EUR individuals in clinical genetic databases reduces power to detect true associations due to allele frequency differences, which can lead to misclassification or loss of power^19,20^. Our findings highlight new fine-scale populations with generally low or high ACMG putative damaging variants for further investigation.
While the findings are valuable, certain limitations should be considered. First, due to the nature of EHR, our study captures a partial view of patient care, potentially affecting the completeness of our findings. Second, our analysis focuses on the risk of receiving a diagnosis for a disease, which is related but not equivalent to the underlying risk of disease development. This may influence how our findings should be interpreted in the context of true disease incidence. It is our expectation that ATLAS will continue to expand, improving its power and utility, but will also become an engine for clinical intervention, allowing researchers to rapidly identify individuals for genetic (and non-genetic) health studies and the implementation of precision medicine as the field continues to move forward.
ATLAS enrollment reflects a modest over-representation of females by self-reported sex (female 55.4%, male 44.6%), especially among patients aged 20–60. The participants are older than the population as a whole (median ATLAS male, 61; female, 56, Figure 1b; Table 1). Medical morbidity was also notably more widespread in the biobank population compared to all adult UCLA patients. This point was evidenced by a significantly larger number of diagnoses in the biobank population compared to all UCLA patients (16 vs. 10.5 mean ICD codes per patient, one year after collection; Supplementary Table 1). The biobank population experienced a substantially higher number of diagnoses across the most prevalent medical condition categories, including endocrine, cardiovascular, musculoskeletal, gastrointestinal, neoplasms, neurological, respiratory, mental, genitourinary and infections (Supplementary Figure 2). For instance, endocrine/metabolic and cardiovascular diseases were diagnosed in 52.0% and 44% of the biobank participants versus 38% and 36% of non-biobank adults, respectively. Biobank participants also experienced substantially more neoplasms (30% vs. 19%) (Supplementary Figure 2c-d). This likely led to significantly more medical encounters in the biobank population (119 vs. 64 mean total medical encounters per patient, and 16.2 vs. 12.3 mean inpatient days), which is consistent with expectations, as greater numbers of clinical visits facilitate the passive blood collection that powers our biobank. The most common diagnoses included hypertension (30% of participants), hyperlipidemia (25%), GERD (18%), anxiety disorder (17%), and depression (15%) (Supplementary Figure 2a).
As time progresses, more diagnoses have an opportunity to be added, and the rate of diagnoses is consistent across a wide spectrum of organ systems (1–2 years after collection) (Supplementary Figure 1b). We retrieved the latest results of the 39 most common laboratory blood count, lipid panel, metabolic panel, HbA1c, and vitamin D,25-Hydroxy to explore test frequencies and variation in results. The largest variability in adults was observed for bilirubin (median = 0.4, IQR = 0.3), followed by triglycerides (median = 90, IQR = 66), alanine aminotransferase (median = 22, IQR = 14), neutrophils (median = 3.73, IQR = 2.07), and LDL cholesterol (median = 97, IQR = 50). Similarly, we retrieved information on the 50 most abundant filled prescriptions to learn about treatment and prescribing tendencies. (Supplementary Table 2). As hospital visits and surgeries were the most common types of encounters in the biobank population, it is not surprising that among the most prescribed generic medications were Acetaminophen (n = 738,061 prescriptions), Ondansetron HCl (n = 636,165), Propofol (n = 633,830), and Lidocaine HCl (n = 492,838) (Supplementary Table 3).
Array genotypes were obtained using the Global Screening Array. All data was mapped to GRCh38 and dbSNP, build 147^119^. Common haplotypes and variants were imputed using the TOPMed Freeze 5 panel^26,27^ using 668,127 observed SNPs, resulting in a total of 50,757,223 high-quality called genotypes variants following imputation, an average of 2,048,050 per individual. Details regarding QC and imputation procedure were described before^26^. Minimal QC was applied to the imputed genotypes. We retained only non-duplicated, bi-allelic variants with high imputation quality (R2 > 0.7), a minor allele frequency (MAF) > 0.1%, and a missing rate < 5%. Genotypes with a missing rate > 5% were excluded from further analysis. Concordance between genotypes determined by observed array variants and targeted Illumina sequencing was approximately 99.6%. The vcf-compare command in VCFtools (v0.1.16) was used to compare genotypes between the two platforms^120^.
Demographic information, vitals, lab tests, ICD codes, and medication prescriptions were retrieved from the UCLA Data Discovery Repository (DDR), established on March 2, 2013^112^, containing deidentified patient EHR data from our health system.
Demographic, vital, and lab data was up-to-date as of Aug 24, 2024. For vitals and encounter counts, only records from in-person encounters (hospital visits, appointments, surgeries, office visits, and walk-ins) were considered. Both the latest records and median values (medians across all encounters, the latest three years, and the latest year) were calculated to reduce possible recorded typos.
Yearly encounter numbers were retrieved in 2024. We included only complete yearly data up to 2023 to ensure consistency and avoid partial data from 2024. Hospital encounters during the COVID pandemic years (2019–2022) were excluded.
For associations involving clinical phenotypes, both ICD-9 and 10 were extracted. The ICD-9 codes were mapped to phecodes using phecode Map v1.2^59^, while ICD-10 codes were mapped using phecode Map v1.2b1^60^, both of which were obtained through the PheWAS catalog^61^. In total, there were 1,253 phecodes with over 100 cases in ATLAS. For each phecode, individuals were classified as cases if they had at least two occurrences of the same phecode, with occurrences spaced at least 30 days apart. Controls were defined as individuals who had no record of the phecode in their EHR data and at least two recorded encounters in the system. Undecidable individuals were excluded.
As all samples were collected incidentally, there was interest in characterizing the population, especially in comparison to non-biobank UCLA patients. To characterize the non-biobank UCLA patients while mitigating time-dependent confounding, we included a subpopulation with an encounter within one year of ATLAS launch date. Encounters could be either inpatient or outpatient, and we summarized the relative proportion of patients with only outpatient encounters or at least one inpatient encounter (there were no patients with only inpatient encounters).
To identify clinical phenotype patterns in ATLAS and to compare these to patterns of non-biobank UCLA patients, we used all ICD codes from each patient to encapsulate past and present conditions, subject to practical challenges^121^. For the ATLAS population, we captured a snapshot using their encounter closest to their biobank sample collection date. For the non-biobank patients, we used their closest encounter to the ATLAS launch date. These encounters represented the baseline encounters for both populations. The ICD codes at these encounters were converted to phecodes and phecode groups^59–61,122^ to represent meaningful categories of disease. Phecodes were extracted using pandas v2.2.2 with Python v3.9.19. We reported the unique phecodes with a prevalence ≥5% in the UCLA ATLAS population. ICD codes at the baseline encounters were also used to calculate Charlson and Elixhauser Comorbidity Indices ^36,37,123^ both measures that predict mortality. The scores were derived using the comorbidity R function v1.0.7^124^ with R v4.1.2.
In addition to characterizing patient populations at a baseline time, we also described differences in how they changed over time - an important consideration when assessing relative disease burden across populations. Patients were enrolled in UCLA ATLAS and were encountered within the health system at different times; therefore, we controlled the interval over which we measured change. We identified patients with at least one encounter between one and two years after their baseline encounter. From this subpopulation we obtained patients’ first encounter within this window and summarized the cumulative encounters that occurred between the two times. We also used the ICD codes at the second encounter to derive new comorbidity scores, as well as extract updated phecodes. The change in comorbidity scores and increases in phecode prevalence provided insight into the evolving disposition of the UCLA ATLAS population over time.
The genetic ancestry of individuals in the ATLAS dataset was estimated by assessing their proximity to the centroids of 1000 Genomes superpopulations in principal component (PC) space. The top 20 PCs were calculated using the bed_projectPCA bigsnpr^119^ R function v1.12.2 with default parameters. For every individual, the Euclidean distance to the centroids of the five broad-scale populations (AMR, AFR, EUR, EAS, SAS) was computed. Individuals were assigned AMR or AFR ancestry if the nearest centroid corresponded to one of these populations, as these groups are well-separated in PC space. For EUR, EAS, and SAS ancestries, which exhibit more genetic overlap, a stricter distance threshold was enforced to minimize misclassification. Specifically, an individual was assigned to one of these ancestries if their distance to the nearest centroid was less than a scaled threshold, calculated Threshold=max(dist)×min(FST)/max(FST)×0.5 where max(dist) is the largest squared distance among centroids and min(FST)/max(FST) accounts for genetic differentiation. Individuals who could not be assigned to any ancestry cluster under these criteria were labeled as "admixed/unknown." Visualization was performed using Boutros Plotting General (BPG) R package v.7.1.0^125^.
To test differences in disease diagnosis across ancestries, phecodes (retrieved as described in Retrieving Phenotype Data) were associated with genetic ancestry groups using logistic regression, adjusting for age (age at diagnosis for cases and the latest age for controls) and genetic sex if applicable. Differences in encounter numbers and comorbidity index were tested using ANCOVA, adjusted for genetic sex and age. For all parts, patients under 18 years old, with ambiguous sex or with “unclassified” genetic ancestry were excluded. Visualization was performed using the BPG R package v.7.1.0^125^. The following disease-phecodes pairs were used to define disease Type 2 diabetes-Type 2 diabetes, Sleep apnea-Sleep apnea, Essential hypertension-Essential hypertension, Hyperlipidemia-Hyperlipidemia, Anxiety disorders-Anxiety disorders&Anxiety disorder&Generalized anxiety disorder, Asthma-Asthma, Parkinson's disease-Parkinson's disease, Type 1 diabetes-Type 1 diabetes, Schizophrenia-Schizophrenia, Crohn's disease-Regional enteritis, Chronic kidney disease-Chronic kidney disease, Stage I or II, Multiple sclerosis-Multiple sclerosis, Major depressive disorder-Major depressive disorder, Cerebrovascular disease-Cerebrovascular disease, Atrial fibrillation-Atrial fibrillation, Hypercholesterolemia-Hypercholesterolemia, Hyperlipidemia-Hyperlipidemia, Coronary atherosclerosis-Coronary atherosclerosis, Hypertrophic obstructive cardiomyopathy-Hypertrophic obstructive cardiomyopathy, Myocardial infarction-Myocardial infarction, Systemic lupus erythematosus- Systemic lupus erythematosus , Gout-Gout, Bipolar-Bipolar, Psoriatic arthropathy-Psoriatic arthropathy, Epilepsy-Epilepsy, Neurofibromatosis-Neurofibromatosis, Dementias-Dementias, Obesity-Obesity, Obsessive-compulsive disorders-Obsessive-compulsive disorders, Autism- Autism, Migraine- Migraine, Alzheimer's disease- Alzheimer's disease, Coronary atherosclerosis-Coronary atherosclerosis, Posttraumatic stress disorder- Posttraumatic stress disorder.
PGS were calculated using array data after imputation in EUR individuals. Related individuals based on their genetic similarity were excluded (defined using PLINK v2.0a^109^ with the relatedness coefficient --king-cutoff 0.05). PGS were calculated using pgsc_calc^126^ with the default setting but setting the –min_thres to 0.65. Logistic regression was used to associate every PGS with the corresponding trait based on phecodes (see Retrieving phenotype data). The following PGS model IDs from the PGS catalog^127^ and their phecode pairs were PGS002250-Malignant neoplasm of ovary, PGS003766-Cancer of prostate, PGS000004-Malignant neoplasm of female breast&Breast cancer, PGS001794-Thyroid cancer, PGS000079-"Melanomas of skin, PGS002264-Pancreatic cancer, PGS003395-Colorectal cancer&Colon cancer, PGS000729-Type 2 diabetes, PGS002025-Type 1 diabetes, PGS004254-Regional enteritis, PGS004699-Multiple sclerosis, PGS000134-Schizophrenia, PGS004760-Major depressive disorder, PGS001806-Malignant neoplasm of testis, PGS004687-Malignant neoplasm of bladder & Cancer of bladder, PGS000039-Cerebrovascular disease, PGS004526-Essential hypertension, PGS004706-Atrial fibrillation, PGS004784-Hypercholesterolemia, PGS002029-Hyperlipidemia, PGS003726-Coronary atherosclerosis, PGS000739-Hypertrophic obstructive cardiomyopathy, PGS004528-Myocardial infarction, PGS000803-Systemic lupus erythematosus, PGS001789-Gout, PGS002786-Bipolar, PGS000198-Psoriatic arthropathy. For prostate and testicular cancer, only males were included, and for breast and ovarian cancer, only females. An adjustment was made for age at diagnosis for cases and the latest age for controls, genetic sex when both sexes were included, and the first ten genetic principal components (PCs). For Figure 3, the top end bottom PGS deciles compared with the 5^th^ decile were considered, testing only EUR individuals. FDR was used for multiple testing correction. The same process was repeated for other ancestries as presented in the supplementary material. Visualization was generated using BPG R package v.7.1.0^125^.
ATLAS array data was merged with genotyping data from the 1000 Genomes Project^128^, the Simons Genome Diversity Project^129^, and the Human Genome Diversity Project^130^. BCFtools^131^ annotate was used to harmonize variant reference SNP ID (RSIDs), and BCFtools^131^ norm with a GRCh38 genome reference was used to standardize the genotyping data. Sites or individuals with more than 1% missing were removed using PLINK^132^ --mind and --geno. Only SNPs with MAF > 1% across all individuals were kept.
SHAPEIT5^133^ with default parameters and the distributed GRCh38 map files were used to phase genotyping data, one chromosome at a time.
A custom Python script that converts PLINK bed files to PLINK ped/map^132^ files while conserving phasing information was used to convert genotyping data. Centimorgan data for the map files were generated using the same genetic map data in SHAPEIT5.
Identity-by-descent segments were called using iLASH^134^ with the following slice_size 350, step_size 350, perm_count 20, shingle_size 15, shingle_overlap 0, bucket_count 5, max_thread 20, match_threshold 0.99, interest_threshold 0.70, min_length 2.9, auto_slice 1, slice_length 2.9, cm_overlap 1 and minhash_threshold 55. Identity-by-descent was called one chromosome at a time.
Identity-by-descent segment outliers were removed as described in Belbin et al.^57^ and Caggiano et al.^56^. Segments overlapping centromeres or the human leukocyte antigen (HLA) region were removed. Regions that may have false positive identity by descent were identified using the following process and total identity by descent per each SNP was identified by summing across all identity-by-descent segments that overlapped each SNP; SNPs with a total identity by descent greater than or less than three standard deviations from the genome-wide mean were removed.
To identify clusters, we followed the approach of Caggiano et al.^56^ and Dai et al.^58^ and applied Louvain clustering^135^. An undirected network is generated based on pairs of individuals who share identity-by-descent nodes are the individuals, edges are the total, genome-wide identity-by-descent shared as the edges. We used the Python package, NetworkIt^136^, to iteratively run Louvain clustering four times to detect fine-scale clusters.
To avoid redundancy and maximize sample size, clusters were merged in two stages to produce a final set of consensus fine-scale cohorts. First, following Caggiano et al.^56^and Dai et al.^58^, we computed pairwise Hudson’s FST using PLINK v2.0a^137^ across 378 clusters identified from the fourth layer of Louvain clustering. After removing clusters with fewer than 10 individuals, to avoid unreliable FST estimates, we merged the remaining 356 clusters into 67 clusters if the pairwise FST was less than 0.001.
In the second stage, we refined clusters using IBD sharing. For each cluster pair, we examined all inter-cluster individual pairs to (1) IBDmean, the average cM shared, and (2) IBDprop, the proportion of pairs sharing at least 3cM of IBD as detected by iLash. We defined a composite metric for cluster merging, IBDweighted = IBDmean×IBDprop, that captures both the extent and prevalence of genomic IBD sharing. Next, clusters were sorted from smallest to largest number of ATLAS participants. For each cluster, we computed the IBDweighted score with all larger clusters and calculated the mean of these pairwise values. The cluster was then merged into the most similar larger cluster – i.e., the one with the highest IBDweighted score-only if that score exceeded the mean. Otherwise, the smaller cluster was retained independently. This process merged 14 small clusters into larger parent clusters, resulting in a total of 36 fine-scale clusters with ≥ 30 individuals each for downstream analyses^56^. These fine-scale clusters were assigned unique identifiers (IBD-01 through IBD-36) and manually annotated with labels to ease interpretation.
We primarily relied on reference individuals, described in “Fine-scale Ancestry Pre-processing and Quality Control”, to add labels to clusters. When clusters did not contain reference individuals or were heterogeneous, we utilized patient-reported race, ethnicity, preferred language, and religious affiliation (in order of priority) to inform our cluster labeling. These aspects are not caused by identity-by-descent segment sharing but can be indicative of a shared culture for individuals within a cluster; these shared practices can influence a group’s demography, environment, and disease risk^138^. Our labels are not definitive and are our best attempt to generate informative assignments for each cluster.
We used logistic regression to model the association between fine-scale ancestry assignment and phecode prevalence, estimating OR with 95% confidence intervals using the logistf R package v1.26.0 for Firth's bias-Reduced penalized-likelihood logistic regression^139^. Differences were tested between every fine-scale ancestry cluster with at least 100 participants and all other ATLAS individuals. Adjustment was made for age at diagnosis for cases, and the latest age for controls, and for sex when applicable. Phecodes with diagnoses in at least 100 patients were tested, resulting in 1,253 phecodes.
The same model was used to test differences in cardio-metabolic diseases using appropriate phecodes, for each fine-scale cluster within the same broad-scale continental ancestry. For this goal, clusters were assigned to broad-scale ancestries based on the predominant ancestry match among individuals within each cluster. For this, the largest population cluster within each broad-scale ancestry was used as the reference level for associations. FDR was used for Multiple testing correction. The data were visualized using BPG R package v.7.1.0^125^.
PheWAS were conducted using the Regenie v4.0 framework^70^ for all five broad-scale cohorts (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale cohorts with at least 400 individuals (IBD-01 through IBD-15). Association testing was performed separately within each cohort for 1,437 binary traits and 41 quantitative traits, including ICD-derived diagnoses and clinical laboratory measurements (see “Phenotype Definitions”). Samples were restricted to those in predefined inclusion lists (–keep), and trait-specific covariates were provided via –covarFile, including age, sex (modeled categorically with –catCovarList sex), BMI, and the top 10 genetic principal components. Quantitative traits were divided into two analysis groups based on missingness patterns, following Regenie’s recommendation that traits with similar levels of missing data are modeled together in Step 1 (Supplementary Figure 11).
For each cohort, unimputed array genotypes in bed format were filtered using PLINK v2.00a^137^ to produce a minimal set of variants suitable for estimating genomewide polygenic effects through ridge regression in step 1 of Regenie. We included autosomal variants with call rate ≥99% (--geno 0.01), minor allele frequency ≥1% (--maf 0.01), and Hardy-Weinberg equilibrium p > 1×10^−^15 (--hwe 1e-15). Additional linkage disequilibrium (LD) pruning was applied using a sliding window of 1000 SNPs, advanced by 100 SNPs, with an r2 threshold of 0.9 (--indep-pairwise 1000 100 0.9), producing an average of 389,209 SNPs per cohort. For binary traits, step 1 was run with a minimum case count of 50 (--minCaseCount 50). For quantitative traits, step 1 was run with Rank Inverse Normal Transformation (--apply-rint) enabled to stabilize variance across lab values with differing distributions. For all traits, a block size of 1000 (--bsize 1000) was used, leave-one out cross validation (--loocv) was enabled, and sex was marked as a categorical covariate (--catCovarList sex).
Prediction files from step 1 for each trait and imputed genotypes in BGEN format were used to perform association testing in step 2 of Regenie. For binary traits, a minimum allele count of 20 (--minMAC 20) and minimum case count of 50 (--minCaseCount 50) were required, and Firth approximation was performed (--firth --approx) using a p-value threshold of 0.01 (--pThresh 0.01). For quantitative traits, a minimum allele count of 20 (--minMAC 20) was required and Rank Inverse Normal Transformation (--apply-rint) was applied. For all traits, a block size of 500 (--bsize 500) was used and sex was marked as a categorical covariate (--catCovarList sex).
Genomic DNA libraries were created by enzymatically shearing high molecular weight genomic DNA to a mean fragment size of 200 base pairs. Multiplexity of exome capture and sequencing was achieved by adding unique asymmetric 10-bp barcodes to the DNA fragments of single samples during library amplifications. Equal molar amounts of DNA samples were pooled for exome capture using a slightly modified version probe library of xGen exome research panel from Integrated DNA Technology (IDT). After PCR amplification and quantification of the captured DNA, samples were multiplexed and loaded to Illumina sequencing machines for sequencing to generate 75 base pair paired end reads. The samples in this study were sequenced using the Illumina sequencing machines including NovaSeq 6000 with S2 or S4 flow cells and the NovaSeqX with 25B flow cells.
Sequencing reads in FASTQ format were generated from Illumina image data using bcl2fastq program (v2.20, Illumina). Whole exome reads alignment and germline small variant detection were conducted using the Original Quality Functionally Equivalent (OQFE) protocol described in Krasheninina et al., 2020^140^. Briefly, raw read files (FASTQ) were mapped to the GRCh38 reference obtained from https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/GRCh38_reference_genome/ using BWA-MEM v0.7.17-r1188 in an alt-aware manner^141^. Duplicate reads were then marked with Picard v2.21.2^142^. The final CRAM files were compressed with SAMtools v1.2^131^. Germline variant detection was performed on each CRAM using a Parabricks accelerated version of DeepVariant v0.10.0 with a custom WES model ^143^, resulting in a sample-level gVCF (genomic VCF). Per-sample gVCFs were merged with GLnexus v1.4.3^144^ into joint-genotyped multi-sample project-level VCF (pVCF). Variant prediction was restricted to the exome capture region and the 100 base-pairs buffer on each side of the target regions. The WES was of high quality and reached an average coverage of 41.6-fold with a minimum of 24.2-fold, 90% having ≥ 34.7-fold in targeted regions, consistent with a previous publication^145^. The pVCF was converted to a plink file format using PLINK 1.9^132^ for downstream analyses.
These data underwent extensive quality control to ensure the absence of contamination, duplication, and other technical errors, as well as sufficient read depth to guarantee reliability and accuracy. Genetic duplicates were defined based on the aggregated genotype data of all sequenced samples. The gender was predicted based on the ratio of read coverage on chromosome Y over the whole exome read coverage. VerifyBamID v1.1.3^146^ was used to estimate sample contamination. Samples were excluded if they showed gender discordance, duplication, cross-Indvidual contamination > 5%, or coverage < 20X in more than 20% of target regions.
Alignment quality metrics were generated using Picard CollectHsMetrics v2.27.4^142^ and SAMtools stats v1.15.1 with default settings^114^. MultiQC v1.27.1^147^ was used to systematically aggregate sample-level quality metrics. RTGtools v 3.12.1^148^ was used to assess variant QC metrics, including the number of variants for SNPs, small insertions and deletions; genotype counts; heterozygous-to-homozygous ratios for each variant type and transition/transversion (Ti/Tv) ratio. Genotypes were further confirmed using an alternative approach^149^ using 500 random samples yielding a high-concordance on-target sites (median concordance of 99.5% for SNPs and indels).
For WES-based analyses, we used samples that also had array genotyping data, which was necessary to consistently define genetic ancestry. We annotated variants using Ensembl Variant Effect Predictor (VEP) v.112^90^ for 58,387 samples assigned to the EUR, AFR, SAS, EAS, or AMR ancestry class (i.e. excluding unclassified genetic ancestry samples). Annotations were performed using the GRCh38 cache and corresponding reference FASTA file. We limited our analysis to autosomal chromosomes to avoid technical artifacts in variant calling caused by the differences in ploidy between males and females, as well as the high-sequence similarity between the X and Y chromosomes in certain regions^150^. Counts for single nucleotide variants, indels, multi-allelic, synonymous, missense, and LOF variants were restricted to whole exome sequencing (WES)-targeted regions. Consistent with a previous UK Biobank WES study^151^, we classified LOF variants as those with the following stop_gained, start_lost, splice_donor, splice_acceptor, stop_lost, and frameshift. To increase reliability, we further restricted LOF variants to those flagged as high confidence by LOFTEE^89^. For multi-allelic variants, the predicted function for each alternate allele was determined using the --pick-allele option based on the default ordered set of criteria defined by VEP^90^. To determine the number of variants with MAF < 1% while accounting for ancestry-specific allele frequency differences, we retained variants with MAF < 1% in at least one ancestry group.
For multi-allelic variants, we defined the minor allele as the second most common allele (including the reference) and classified the variant as rare if the cumulative allele frequency of all alternate alleles was < 1%. For rare multi-allelic variants, we incremented a sample's count only if it carried an alternate allele with MAF < 1%. In multi-allelic cases where the minor allele was the reference or the alternate alleles had AF ≥ 1%, we incremented a sample's count if it carried any alternate allele. For rare functional variants (e.g., missense), we incremented a sample’s count only if it carried an alternate allele with MAF < 1% corresponding to that functional category.
To identify clusters where individuals were enriched for carriers of monogenic variants associated with disease, we identified pathogenic variants based on different sets of criteria. First, as a positive control, we identified variants, via literature search, where variants are known to be enriched within certain populations. Next, we used pathogenic variants that were curated by experts within the field and underwent stringent review to be considered pathogenic from ClinGen^86^. ClinGen variants filtered to identify P/LP variants with autosomal dominant inheritance, autosomal recessive inheritance, or semidominant inheritance, resulting in 3,521 variants associated with 204 conditions overall according to ClinGen. We removed one HNF4A ClinGen variants due to high allele frequencies (MAF>0.01).
To test enrichment of single alleles across clusters, “carriers” are identified as individuals who carry at least one allele of a pathogenic variant within a gene associated with a monogenic condition. We then calculated carrier frequency as the number of identified carriers over the total number of individuals with WES available and were unrelated using a kinship coefficient of 0.05 based on PLINK v2.0a^137^ king-cutoff. To identify clusters with a significantly different carrier frequency of variants, we applied a Fisher’s Exact test to the carrier frequency of each cluster against the carrier frequency of the other clusters grouped together. We then used FDR corrections for p-values generated from all Fisher’s exact tests, and selected significant clusters where at least 5 carriers were identified and FDR < 0.05.
The list of all ClinGen^86^ P/LP variants identified in any ACMG^88^ genes was extracted as described above. Variants were labeled as missense or LOF, based on VEP v.112^90^ annotations. Missense variants were defined for ‘missense_variant’ variants according to the ‘consequence’ VEP output column. LOF was defined for high-confidence “HC” LOF variants based on LOFTEE^89^. All variants were rare across broad-scale ancestries. A WES plink file with all individuals was filtered to include only the listed variants. Then, this file was broken into ancestry groups using PLINK v2.0a^137^. Related individuals were removed. The frequency of each allele in every population group was calculated with the PLINK v2.0a –freq command, and total frequencies were summed-up for rare missense and LOF variants separately. To test differences in the allele counts across populations, allele dosages were calculated with PLINK v2.0a^137^ --recode A option. Dosages were summed to calculate the total alternative (ALT) allele count within each population. The number of reference (REF) alleles was defined as twice the number of individuals in the group minus the number of ALT alleles. Fisher's Exact tests were applied to test differences between the number of REF and ALT alleles in each population compared to all other individuals not assigned to that specific group. Bonferroni correction was applied for multiple testing correction. Plotting was done using BPG R package v.7.1.0^125^.
Variants from exome sequencing were annotated using VEP v.112^90^ with the dbNSFP v4.9a^153^ and LOFTEE^89^ plugins installed. The LOFTEE high-confidence “HC” flag was used to select for LOF variants predicted to have deleterious effects.
Missense variants were assigned a 9-point deleteriousness score based on a consensus of nine missense deleteriousness prediction toolkits, similar to methods described in prior biobank-scale rare variant studies^154^. We used the Critical Assessment of Genome Interpretation (CAGI) project^91^ to prioritize well-performing tools not trained on the same features. We selected five meta-predictors – ClinPred^155^, MetaRNN^156^, BayesDel_addAF^157^, VARITY_R^158^, REVEL^159^ – and four stand-alone predictors – AlphaMissense^160^, MutPred2^161^, VEST4^162^, ESM-1b^163^ – for use in our analysis. We assigned each variant a binary score per tool based on dbNSFP rank 1 if the variant’s score exceeded the threshold score for being more likely a deleterious ClinGen or ClinVar variant than a background variant (Supplementary Figure 12), and 0 otherwise. Summing these binary scores produced a deleteriousness score ranging from 0 to 9, with predicted damaging missense variants scoring ≥ 5 retained for downstream analysis.
A WES PLINK file was filtered to include computationally predicted LOF and predicted damaging missense variants (see above) in ACMG genes^88^, providing a more comprehensive evaluation than the one based solely on ClinGen P/LP variants, which included only 17 genes. A WES plink file with all individuals was filtered to include only the ACMG putative damaging variants, and the file was split into ancestry groups (broad- and fine-scale ancestries) using PLINK v2.0a^137^. Related individuals were removed. Allele frequencies were calculated with PLINK v2.0a, and only rare variants (MAF < 1% in all broad-scale ancestries) were kept. Allele dosages per individual were extracted using PLINK v2.0a, and the total rare LOF and predicted damaging missense alleles were counted separately per patient. First, the differences between the total numbers of REF and ALT alleles across ancestries were evaluated with a Fisher's Exact test as described above (ClinGen ACMG analysis). Second, a Mann-Whitney U test with a Bonferroni correction was applied to test the difference in the distribution of rare LOF and predicted damaging missense counts per individual between any broad- or fine-scale group compared to all others. Figure 5d presents the Mann-Whitney U test results, with statistically significant differences indicated by an asterisk (*). Visualization was done with BPG^125^.
ExWAS were conducted using the Regenie v4.0 framework^70^ for all five broad-scale cohorts (AMR, AFR, EAS, EUR, SAS) and fifteen fine-scale cohorts with at least 400 individuals (IBD-01 through IBD-15). Within each cohort, imputed genotype dosages in BGEN format (--bgen) and prediction scores from PheWAS step 1 (--pred) were used to conduct gene-based association testing for binary and quantitative traits across 17,676 genes. Samples were restricted to those in predefined inclusion lists (--keep), and trait-specific covariates were provided via --covarFile, including age, sex (modeled categorically with --catCovarList sex), BMI, and the top 10 genetic principal components.
Whole-exome genotypes in bed format were filtered using PLINK v2.0a, retaining variants with a call rate ≥ 90% (--geno 0.1), minor allele count ≥ 1 (--mac 1), and Hardy-Weinberg equilibrium p-value > 1×10^−15^ (--hwe 1e-15). In addition, 784,548 variants overlapping low-complexity regions (LCRs) were excluded prior to analysis. Briefly, SNP positions were extracted from the .bim file, intersected using bedtools v2.29.1^164^ with annotated LCRs from the Genome in a Bottle Consortium^165^, and filtered from the genotype files, resulting in 12,326,160 exome variants retained for burden testing.
Filtered whole-exome genotypes in BGEN format and prediction files from PheWAS step 1 were used to perform gene-based association testing. Regenie-style annotation, set, and mask files for predicted deleterious LOF and missense variants (see “Variant Annotation”) were created programmatically using Python v3.11.9 with the polars v1.2.1 package. For binary traits (--bt), a minimum case count of 50 (--minCaseCount 50), minimum minor allele count of 5 (--minMAC 5), and block size of 1000 (--bsize 1000) were enforced. Firth logistic regression with saddlepoint approximation (--firth --approx) was applied for variants with p < 0.01 (--pThresh 0.01). For quantitative traits (--qt), rank inverse normal transformation (--apply-rint), a minimum minor allele count of 5 (--minMAC 5), and block size of 500 (--bsize 500) were enforced. Regenie’s implementation of the RGC gene-based p-value test (–rgc-gene-p) was enabled for all analyses. Additionally, variants were binned by minor allele frequency using 1% bins (--aaf-bins 0.01), and SNP-level membership for each burden mask was recorded (--write-mask-snplist).
In addition to gene-based analyses, single variant association testing was performed for selected predicted deleterious LOF and missense variants using filtered whole-exome genotype dosages in BGEN format and prediction scores from PheWAS step 1. For both binary and quantitative traits, tests incorporated the same set of covariates (age, sex modeled categorically, BMI, and the top 10 genetic principal components) and enforced a minimum minor allele count of 5 (--minMAC 5). For binary traits, Firth logistic regression with saddlepoint approximation was applied for variants with p < 0.01, along with a block size of 1000, while quantitative traits were evaluated using rank inverse normal transformed phenotypes with a block size of 500.
ATLAS GLP1-RAs prescriptions, including medication names, start and end dates, discrete dose, usage instruction, strength and route (oral or subcutaneous), were retrieved and grouped based on the following simple generic dulaglutide, semaglutide, liraglutide, exenatide, albiglutide and lixisenatide. For all statistical analyses, only semaglutide users were considered. Usage start dates were defined based on the earliest prescription start date for each patient. In the case of a missing start date, the prescription ordering date was used instead (in most cases, these two fields were identical). Overlapping medication periods were handled such that when a new prescription started, the previous one was considered to have ended. Similarly, missing end dates were determined using the start date of the next prescription when available. In the case of completely overlapping prescriptions with different instructions, the combination of both routes and the weighted average dose was considered. If discrete dose information was missing, the medication dose was extracted from the instruction’s free text field. If the instructions also omitted the dose information, it was imputed for a given medication type, based on the weighted dose average from the entire cohort on the corresponding type. Initial weight and BMI were defined as the median of all available weight or BMI measurements recorded from in-person visits, within 6 months prior to the first prescription start date. Not relying solely on a single data point helps minimize the likelihood of typos and errors in the EHR. The percentage of weight change was obtained for every weight measurement recorded on in-person visits within the period of active prescriptions between 4–60 weeks in total on semaglutide. The medication dose at each time point was defined as the weighted sum of all medication doses (doses multiplied by number of prescription weeks) by the measurement date. In case both oral and subcutaneous medications were used within a period, a combined route category was defined.
To identify the overall weight loss patterns across time, the FPCA R function from the fdapace package v.0.6.0^166,167^, which is suited to plot smoothed longitudinal data with repeated measurement, was used. This identified a consistent weight loss pattern up to 60 weeks, and sparse data points beyond ~150 weeks. Thus, we restricted all analyses to this period.
For all analysis parts that involved longitudinal data with repeated measurements, a linear mixed model with the bobyqa optimizer and an increased function evaluation limit (maxfun = 10000) was used (lmer R function; the lmerTest R package v.3.1.3^168^). In each case, ANOVA was used to identify differences between models to define the best way to model a potential nonlinear relationship between weeks and weight loss, when treating weeks as a fixed effect, based on the Akaike Information Criterion (AIC). In some cases, the best model was achieved using a restricted cubic spline (RCS) with the rcs R function from the rms package (v.7.0.0), applied to weeks. In other cases, a polynomial function of weeks provided a better fit. To plot smoothed longitudinal data with 95% CI, the fitted model values (excluding covariates) and 95% lower and upper confidence bounds, were extracted using the visreg package^169^ v.2.7.0 and visualized using the BPG package^125^.
For testing the effect of non-genetic factors, fixed effects were defined for the medication dose, route, sex, age, initial BMI, 10 first genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and weeks on semaglutide. The model included both random intercepts and slopes for weeks on semaglutide. Bonferroni correction was applied to control for multiple testing.
For testing differences across ancestries, fixed effects were defined for the medication dose, route, sex, age, initial weight a polynomial function of weeks on semaglutide, and an interaction between a polynomial function of weeks and genetic ancestry categorical class (EUR, AFR, EAS, SAS and AMR). The model included both random intercepts and slopes for the polynomial function of weeks on semaglutide. An ANOVA was performed on the model with a Bonferroni correction to assess the global effects of ancestry and ancestry×weeks interaction. When a significant result was found, post hoc comparisons between groups were conducted using the summary lmerTest function (v.3.1.3^168^), applying a Bonferroni correction to all 12 class or class × time interactions.
To test the relationship between PGS and weight loss relying on longitudinal data, scaled BMI (PGS000027) and DM2 PGS (PGS000729) were divided into three equal bins low, intermediate and high scores. Related individuals based on their genetic similarity were excluded (defined using PLINK v2.0a^137^ with the --king-cutoff 0.05), Then, for each trait, fixed effects were defined for the categorical PGS bins, medication dose, route, sex, age, initial weight, 10 first genetic PCs (excluding PC6 due to a strong collinearity with PC5 that disrupted the model convergence) and a polynomial function of weeks. Random intercepts were defined to account for repeated measures within individuals. Bonferroni correction was applied to control for multiple testing.
To test the relationship between PGS and weight loss relying on a simplified model where the maximum weight loss was considered for each patient, linear regression was applied. The model was adjusted for the medication dose and route, 10 genetic PCs, age, sex, and initial weight. Visualization was made with BPG^125^, with smoothed data and 95% CI using loess.as R function (fANOVA package v.0.6.1^170^).
The analysis was performed to test a relationship between the maximum weight loss on semaglutide and common genetic variants, for each broad-scale ancestry using SAIGE^171^. For step one, array observed variants were used following PLINK v2.0a^127^ filtering, with the --maf 0.01 --mind 0.1 --geno 0.1 --hwe 1e-6 . For the second step, array observed and imputed variants were used following plink filtering with --maf 0.01 --geno 0.05 --hwe 1e-6. A quantitative analysis was conducted with the traitType flag. The medication dose and route, the first five genetic PCs, age, sex, and initial weight were used as covariates. METAL^172^ was used for meta-analysis.
To identify genes genetically associated with weight loss, we used Regenie^70^ with an additive model for gene-level tests. Only EUR semaglutide users were used, and per each the maximum weight loss record with the corresponding number of weeks was considered. The list of candidate genes was limited to bonferroni-significant proteins whose plasma abundance was altered by semaglutide treatment^110^. We considered variants within these genes with a predicted moderate or high impact on the protein function, according to VEP v.1.2^90^. For step one, the variant list was limited to observed array SNPs following a PLINK v2.0a^137^ filtering --maf 0.01 --mac 100 --indep-pairwise 1000 100 0.9 --chr 1–22 --snps-only --geno 0.1 --hwe 1e-15. In step 2, we applied PLINK v2.0a filtering to variants across all EUR individuals using the parameters --geno 0.05 and --hwe 1e-6. Subsequently, the file was filtered to include only semaglutide users and was used in Step 2. The medication dose and route, the first 10 genetic PCs, age, sex, and initial weight were used as covariates. Bonferroni correction was applied to control for multiple testing.