Authors: Nazia Hilal (1.Division of Genetics and Genomics, Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School, Boston, MA, USA; 2.Broad Institute of MIT and Harvard, Cambridge, MA, USA), Maniteja Arava (1.Division of Genetics and Genomics, Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School, Boston, MA, USA; 2.Broad Institute of MIT and Harvard, Cambridge, MA, USA), Sangita Choudhury (1.Division of Genetics and Genomics, Department of Pediatrics, Boston Children’s Hospital, Harvard Medical School, Boston, MA, USA; 2.Broad Institute of MIT and Harvard, Cambridge, MA, USA)
Categories: Article, Aging, Basic Science Research, Genetics, Omics
Source: Circulation research
Authors: Nazia Hilal, Maniteja Arava, Sangita Choudhury
Single-cell genomics has emerged as a transformative approach to unravel the complexity of somatic variation in specific cells within the human body. This field has profound implications for understanding the role of somatic mutations in aging, cardiovascular disease, and tissue-specific pathologies. By focusing on circulating cells and cardiac-resident cells, including cardiomyocytes, within the heart, this review explores how single-cell genomics provides insights into cellular heterogeneity and clonal evolution. We discuss the implications of somatic variation for cardiovascular health, highlight technological innovation, and future directions in this rapidly evolving field. This review sheds light on national initiatives, such as the NIH SMaHT Network and the EU SOMATICART project, which aim to generate reference atlases of somatic mutations across human tissues. It also explores challenges and future directions in leveraging single-cell approaches to improve diagnostics and therapeutics in cardiovascular disease.
Cardiovascular research has been fundamentally transformed by the advent of single-cell genomic technologies, including single-cell RNA sequencing (scRNA-seq), single-nucleus RNA sequencing (snRNA-seq), and single-cell DNA sequencing. These tools have enabled researchers to dissect the cellular and molecular complexity of the heart and vasculature with unprecedented resolution, revealing a level of heterogeneity and dynamism that was previously masked in bulk tissue analyses^1,2^. Among the most exciting applications of these technologies is the exploration of somatic variations in both circulating and cardiac resident cells, a subject that has opened new avenues for understanding cardiovascular aging and disease pathogenesis.
Somatic variations, which are genetic alterations acquired during an organism’s lifetime, have long been recognized as crucial players in cancer biology^3–7^, however, their role in non-malignant conditions, particularly in cardiovascular health and disease, has only recently begun to be appreciated^8–10^. Studies of Clonal Hematopoiesis of Indeterminate Potential (CHIP) have established compelling links between age-related somatic mutations in hematopoietic stem cells and increased risk of cardiovascular events^11–15^. Building on these findings, recent work, including our own, has begun to explore the broader landscape of somatic variations in cardiac tissues, with implications for understanding aging, heart failure, and vascular disease^9,10^.
This review aims to summarize the current knowledge at the intersection of single-cell genomics, somatic variation, and cardiovascular health. We begin by outlining the technological advances that have enabled high-resolution detection of somatic variations in both blood-derived and tissue-resident cells. We then trace the conceptual evolution of somatic mutations from cancer biology to cardiovascular disease, with a focus on CHIP and emerging evidence from cardiac tissues. We explore how somatic variation affects tissue homeostasis, cellular identity, and susceptibility to disease, highlighting the distinct yet complementary insights gained from studying circulating versus cardiac-resident cells.
We also address practical considerations for researchers, including the challenges of obtaining and analyzing cardiovascular tissues, the advantages of blood-based studies, and the need for robust collaborations between clinicians and scientists. Emphasis is placed on the importance of longitudinal cohorts, sensitive detection methods, and the integration of genomic data with clinical phenotypes.
In sum, the convergence of single-cell genomics and somatic variation research offers a powerful lens through which to reinterpret cardiovascular aging and pathology. By illuminating the cellular consequences of somatic mosaicism in the heart and circulation, this field holds the potential to identify novel biomarkers, therapeutic targets, and strategies for precision cardiovascular medicine.
From conception onwards, the cells of the human body consistently endure genomic changes, whether due to inherent factors or exposure to mutagens^16–18^. Although most DNA damage is corrected and the genome is duplicated with remarkable accuracy, cells continuously accumulate somatic variations throughout their lifespan.
The concept of genetic diversity within the human population is now a well-accepted phenomenon (Figure. 1a). The Human Genome Project delivered the initial near-complete mapping of the human DNA sequence, subsequently leading to extensive initiatives where the 1,000 Genomes Project and the Pangenome Project have together cataloged genetic variation among populations as well as individuals (Figure. 1b)^19^. Every organ within a human body endures somatic genomic changes; however, each specific variants exists in just a subset of cells or may be confined to particular organs (Figure. 1c). Consequently, somatic variations are frequently characterized as mosaic^4^. DNA mutations are the inevitable consequences of errors that arise during replication-repair of DNA damage as well as aging and disease progression^20^. Due to their random and infrequent occurrence, the quantification and characterization of DNA variants in the genomes of somatic cells have been challenging^21^. These variants in DNA drive genetic diversity, alter gene function, define evolutionary trajectories, and provide targets for precision medicine and diagnostics. It is crucial to detect these variants across a wide range of abundance, i.e., variant allele frequency (VAF). Detecting low-abundance variants (e.g. <0.1–1% VAF or in individual cells) is essential for understanding human embryonic development, somatic mosaicism, and clonal hematopoiesis and uncovering pathogenic variants^14^. Additionally, with recent advances, there is an increasing recognition of significant genetic heterogeneity among individuals across various organs and cells. However, identifying these somatic variants is difficult. Unlike hereditary variants, somatic variants are confined to a few and varying proportions of cells, ranging from embryonic alterations present in the majority of cells to those found in individual cells (Figure. 2).
Somatic mosaicism is now acknowledged as a common characteristic of aging. Large-scale genomic surveys show that age-related clones of hematopoietic cells (clonal hematopoiesis) are common, affecting >10–20% of adults over 65, and are associated with chronic inflammatory states. For example, epidemiological studies find that hematopoietic clones carrying DNMT3A, TET2, and related mutations causally contribute to cardiovascular disease by amplifying innate immune signaling^14,22^. Recent cohorts have extended these in a Chinese population study, about 18% of middle-aged adults harbored CHIP, which independently increased incident coronary heart disease (CHD) risk (hazard ratio ~1.4)^23^. Notably, even small clones (variant allele fraction <2%) raised CHD risk by ~1.3-fold. This emerging evidence underscores that somatic variants – long studied in cancer – are also important in non-malignant cardiovascular aging and disease. New clinical genomic efforts are therefore integrating somatic variation into risk models and translating them into early-intervention strategies.
The problem of identifying rare variants is complex due to artifacts that mimic low-frequency mutations during DNA library preparation and sequencing^24^. Contemporary short-read sequencing technologies constrain the identification of mutations in repetitive genomic areas and are presumably less effective for detecting somatic structural alterations. Next-generation sequencing (NGS) technologies have revolutionized the field of genomics research, offering unprecedented capabilities for analyzing DNA and RNA molecules in a high-throughput and cost-effective manner^25^. However, even with ultra-deep sequencing techniques, such as whole-genome sequencing (WGS) and whole-exome sequencing (WES), the detection of somatic variations with high sensitivity and specificity remained a significant challenge due to the presence of artifactual mutations, which are introduced during sample preparation and sequencing processes^26^. The concept of consensus between sequencing reads has been widely accepted as the gold standard for increasing the accuracy of somatic mutation detection, although at a significant cost^27^. Ultra-deep sequencing techniques were initially favored for somatic variant detection, with the goal of finding consensus amongst multiple reads within a sample. However, even within a single individual, ultra-deep sequencing of multiple tissue sites revealed at most 297 early clonal variants, a fraction of the genetic diversity known to be present by single-cell sequencing^28^.
To address this challenge, researchers have developed various advanced sequencing technologies and error-correction techniques both experimentally and bioinformatically (Figure. 3), such as those that focus on generating consensus by DNA barcoding^24,29^. That is, DNA molecules are appended with a unique identifier and amplified. Multiple copies of each barcoded DNA molecule are then sequenced into reads, and a consensus sequence from the multiple reads reduces errors that are introduced by library preparation, polymerase chain reaction (PCR), and instrumentation, which would be present in only a subset of reads with the same identifier. The recent advent of duplex consensus sequencing has heralded a new generation of accuracy^30–38^. By tracking the original DNA molecule across multiple amplified reads, UMIs enable duplex consensus sequencing, where both strands of DNA are sequenced and compared. This technique dramatically reduces error rates from 10−³ to as low as 10−⁷, enabling the confident detection of extremely rare variants. Multiple techniques focus on targeted areas of the genome (Twin Strand Biosciences) or are based on restriction sites (Nanoseq)^32^or fragmentation of the genome and standard A-tailing and ligation (BotseqS^33^, CODEC^34^). Pro-Seq^35^, META-CS^36^ a TN5-duplex technology^37^, has enabled comprehensive somatic variant characterization. Technological advances have dramatically improved the fidelity and scope of somatic mutation detection, enabling unprecedented insight into ultra-rare genetic variants across diverse biological contexts. TwinStrand Duplex Sequencing (Targeted Duplex Sequencing) exemplifies this progress as a highly accurate, commercialized duplex sequencing platform that leverages molecular barcoding and consensus calling from both strands of each DNA molecule. By tagging Watson and Crick strands independently and requiring mutation concordance, error rates can be reduced to as low as one in ten million nucleotides. This approach is typically applied to targeted genomic regions, notably CHIP-associated genes such as TET2, DNMT3A, and JAK2, providing deep coverage with ultra-high fidelity; within cardiovascular research, it has proven invaluable for detecting ultra-rare clonal mutations in blood and quantifying somatic mutation burden in high-risk tissues.
Complementary advances are seen with NanoSeq (Restriction Enzyme–Based Duplex Sequencing), an error-corrected sequencing methodology that replaces random DNA fragmentation with controlled digestion by restriction enzymes. This strategy, which employs nicking endonucleases for predictable breaks and subsequent strand-specific labeling and consensus calling, improves coverage uniformity and base quality while reducing error-prone fragment ends and false positives. Though originally developed for hematopoietic tissue, NanoSeq’s robust sensitivity and accuracy render it a powerful tool for somatic mutation detection in solid tissues such as the heart and vasculature.
Other innovations are broadening the landscape, offering flexibility in throughput, scalability, and sample constraints. For example, BotSeqS (Bottleneck Sequencing Strategy) adopts a cost-effective approach to rare variant detection through extreme dilution of genomic DNA prior to amplification. This bottlenecking ensures that most PCR products originate from a single DNA molecule, which allows for consensus variant calling. Although BotSeqS does not employ duplex-based error correction, it still significantly lowers sequencing error rates, and has seen application in quantifying mutation accumulation with aging. Its main limitation lies in reduced throughput and dependence on early PCR error correction rather than duplex confirmation; nonetheless, in cardiovascular studies where only an estimate of genome-wide somatic burden is required and ultra-high accuracy is less critical, it can be a valuable option.
Emerging approaches such as CODEC (Concatenating Original Duplex for Error Correction) further refine error correction by physically linking both strands of a duplex DNA molecule into a single continuous read through intramolecular ligation, followed by short-read sequencing. Maintaining both strand orientation and sequence complementarity within one molecule, CODEC enables true duplex error correction without high redundancy or complex computational pairing, effectively reducing sequencing costs while preserving specificity. This technique is particularly promising for scenarios where duplex information is critical, but sample input is limiting, such as analyses of biopsy-level heart or vascular tissue.
Meanwhile, Pro-Seq (Proximity Sequencing) introduces an efficient technique based on short proximity ligation between genomic fragments and universal primers, yielding uniquely labeled DNA molecules for highly efficient amplification. While employing error correction through molecular consensus akin to duplex sequencing, Pro-Seq operates without full duplex capture, thus striking a balance between scalability and accuracy. This makes it particularly well-suited for targeted sequencing of lower-abundance mutational hotspots; in the cardiovascular context, it offers a practical solution for screening variants in genes involved in atherosclerosis, inflammation, or fibrotic remodeling.
Finally, in META-CS (Multiple End Tagging Amplification of Complementary Strands), duplex technology employs engineered Tn5 transposase complexes along with strand-specific adapters. This enables simultaneous fragmentation and tagging of both DNA strands, generating fragment-specific barcodes, and minimizes amplification loss caused by intramolecular hairpin formation when identical sequences are present at both fragment ends. META-CS’s multiplexing capabilities support simultaneous interrogation of numerous regions of interest, potentially including those relevant to cardiovascular homeostasis and remodeling.
Further, the advent of long-read sequencing technologies, such as Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), has the potential to overcome the limitations of short-read sequencing by generating longer sequence reads, thereby improving the accuracy and resolution of genomic data analysis^38^. PacBio and ONT have revolutionized structural variant detection. These platforms generate reads spanning thousands of base pairs, allowing researchers to resolve complex genomic regions, large insertions/deletions, and rearrangements missed by short-read platforms. However, long-read sequencing comes with a higher base-calling error rates, For Oxford Nanopore, the error rate varies depending on DNA quality and flow cell chemistry. With R10.4.1 (1D) reads, it’s typically around 2–5%. But for PacBio HiFi sequencing has an error rate of less than 0.1%. These technologies are promising for capturing somatic mosaicism, particularly in highly repetitive regions of the genome. A central challenge in somatic mutation detection is the identification of technical errors from true biological variants, especially when using long-read sequencing platforms such as Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio). These platforms provide reads that span kilobases in length, facilitating the detection of structural variants, complex rearrangements, and phasing of alleles. However, they are known for relatively high base-calling error rates ranging from 2% to 5% for ONT, compared to <0.1% for Illumina short-read platforms.
The Table 1 below highlights the main strengths and limitations of these technological platforms
Even the best sequencing protocols must contend with intrinsic tissue heterogeneity, especially in cardiovascular systems composed of diverse cell types including cardiomyocytes, fibroblasts, endothelial cells, smooth muscle cells, and resident immune cells. Bulk sequencing obscures this heterogeneity, providing averaged signals that may dilute or mask mutations in specific lineages. To resolve this, researchers have turned to single-cell and single-nucleus sequencing (scDNA-seq, snSeq). These techniques isolate individual cells or nuclei from cardiac tissue and amplify their genomes before sequencing. While technically demanding, they provide cell-specific variation profiles, enabling the detection of somatic mosaicism at unprecedented resolution. Single-nucleus sequencing, in particular, is vital for studying multinucleated and polyploid cardiomyocytes, allowing genotyping of individual nuclei and assessment of intra-cellular heterogeneity. The recent development of low-input single-cell sequencing platforms has allowed for the whole-genome profiling of hundreds of single cells, once considered infeasible due to cost and amplification bias^39^. Sophisticated algorithms such as SComatic^40^ can now detect somatic single-nucleotide variants (SNVs) and small insertions/deletions (indels) from single-cell RNA-seq or ATAC-seq data, even in the absence of matched DNA samples. Such tools use population and cell-type filters to distinguish true somatic events from sequencing errors, enabling mutation burden and signature analysis in mixed tissues. Recent reviews emphasize that single-cell genomics can now “tackle previously intractable problems,” such as resolving complex clonal patterns in scarce or post-mitotic cells^39^.
The field of somatic mutation detection is evolving rapidly, driven by technological innovations that enhance the accuracy and resolution of genomic studies. While traditional methods like ultra-deep sequencing have their limitations, the advent of new error correction techniques and duplex consensus sequencing offers promising avenues for future research.
The concept of mutations emerged in the early 20th century, marked by substantial contributions from Gregor Mendel (1860s), who established the principles of inheritance through genetic features. Later in 1901, Hugo de Vries introduced the term “mutation” to denote abrupt alterations in genetic material^41^, and in 1927, Hermann Muller demonstrated that X-rays could induce mutations in Drosophila^42^, marking the initial connection between environmental factors and genetic alterations. Although these initial discoveries highlighted inherited abnormalities, the possibility that mutations in somatic cells could induce illnesses was not explored.
Advances in carcinogen research in the 1950s-60s showed that chemicals, radiation, and other factors might induce genetic alterations^43^. Investigations on smoking and ultraviolet exposure suggested that somatic mutations influenced cancer development. In 1971, Alfred Knudson proposed that two mutational “hits” in tumor suppressor genes such as RB1, resulted in cancer^44^. This notion reinforced the somatic mutation hypothesis of cancer. Later in the 1970–80s, the discovery of oncogenes (e.g., RAS)^45^ and tumor suppressor genes^46^(e.g., TP53) established the connection between particular somatic mutations and the development and progression of tumors. Landmark studies, such as the identification of the Philadelphia chromosome in chronic myeloid leukemia^47^ and the discovery of the RAS oncogene, firmly established somatic mutations as key drivers of cancer. These findings provided the first evidence that cancer is driven by the accumulation of mutations in somatic cells. Since then, large-scale initiatives like The Cancer Genome Atlas^48^(TCGA) and the International Cancer Genome Consortium^49^ (ICGC) have mapped mutations across various cancer types, identifying driver mutations, mutation signatures, and clonal evolution patterns.
While cancer remained the primary focus of somatic variation research for many years, the potential role of these genetic alterations in other diseases began to gain attention towards the end of the 20th century as the field experienced a paradigm shift with the advent of next-generation sequencing technologies in the early 2000s, enabling comprehensive genomic analyses and expanding research beyond cancer. Early clues came from studies of aging and e.g. detection of pre-leukemic DNMT3A/TET2 mutations in healthy elders (commonly termed CHIP) was linked to atherosclerosis by 2014–2015^50^. Recent studies have indicated that somatic variations can cause a wide range of diseases, including neurological, hematological, and immune-related disorders^51^. Individuals accumulate somatic variation even in the absence of clear phenotypic effects, and this variation includes the entire spectrum of mutations observed in the germline. Somatic mutations that arise during development can cause neurological diseases like brain malformations, epilepsy, and intellectual disability. Work from Poduri et.al has shown that specific somatic variations in genes like AKT3, PIK3CA, and mTOR can lead to severe brain overgrowth disorders like hemimegaloencephaly^52^. Later, multiple groups demonstrated the role of somatic variations in aging and neurodegenerative diseases like Alzheimer’s^53,54^. Recent work also showed that somatic variations accumulate in neurons during aging (a process called gynostemium) and provided insights into the cellular processes driving somatic variation and cellular dysfunction. Further, a study from K. Sikora discovered that somatic variations in hematologic precursor cells can cause adult-onset, complex inflammatory diseases like VEXAS (vacuoles, E1 enzyme, X-linked, autoinflammatory, somatic) syndrome^55^.
In cardiology, these approaches are only now maturing. Key milestones include the 2022 discovery that human cardiomyocytes accumulate somatic SNVs at a rapid rate with age^9^. This work, and others, shifted the paradigm by demonstrating that even non-dividing heart cells develop genomic mosaicism. Likewise, cross-tissue initiatives have shown that different cell types carry distinct mutation loads, for example, epithelial cells often have higher burdens than connective tissues – highlighting the need for systemic study of mosaicism. This historical progression from cancer to aging biology underscores that somatic variants are not rare curiosities but common features of human tissues. The current era leverages this foundation to what are the functional impacts of these variations on organ decline and disease?
In the field of cardiovascular research, the idea of somatic variations as contributors to heart disease began to emerge in the late 1990s and early 2000s. Initial studies focused on mitochondrial DNA mutations and their potential role in cardiomyopathies and heart failure^8,56–58^. These early studies laid a conceptual foundation for the broader investigation of acquired variations in cardiovascular health. A pivotal advance in this field has been the discovery of Clonal Hematopoiesis of Indeterminate Potential (CHIP), the most widely studied manifestation of somatic variations in circulating cells^59,60^. CHIP arises from somatic mutations in hematopoietic stem cells (HSPCs), leading to clonal proliferation and a population of mutated circulating blood cells^8,61^. This condition affects over 10% of individuals older than 70 years and is associated with a 30–40% increased risk of mortality, primarily due to higher rates of ischemic stroke and cardiovascular disease^62^. The most commonly mutated CHIP driver genes include DNMT3A, TET2, ASXL1, JAK2, PPM1D, and TP53, which are involved in DNA methylation, chromatin regulation, and cellular proliferation. These mutations have been linked to increased risks of coronary artery disease (CAD) and other cardiovascular conditions^12^.
TET2 mutations lead to increased atherosclerotic plaque size and inflammation through mechanisms involving interleukin-1β (IL-1β) and IL-6^63,64^, while JAK2 mutations activate the JAK-STAT pathway, leading to cytokine-independent proliferation of myeloid cells and increased inflammation, accelerating atherosclerosis and plaque instability^12,65^. The impact of CHIP mutations extends beyond atherosclerosis to other cardiovascular conditions. Studies have shown that CHIP mutations are associated with poor outcomes in heart failure and aortic stenosis^66^. For example, patients with CHIP mutations undergoing transcatheter aortic valve implantation (TAVI) for aortic stenosis had worse survival outcomes^67^m. In heart failure, mutations in genes like DNMT3A and TET2 are linked to increased inflammation and worse clinical outcomes.
Until recently, most CHIP studies focused on Western populations and larger clonal expansions (variant allele fraction [VAF] >2%). However, emerging data are beginning to fill important population-level gaps. For example, a recent study by Zhao et al. (JAMA Cardiology 2024) screened a general Chinese cohort and found CHIP in ~18% of participants with mean age 54, with many clones falling in the low-VAF range (0.5–2%), well below conventional thresholds^23^. Even these small clones were associated with a 1.4-fold increased risk of coronary heart disease (CHD) over a 12-year period. These findings underscore that even modest clonal expansions can impact cardiovascular health, and that risk models should account for both germline and somatic genetics.
Beyond coronary disease, CHIP has been increasingly recognized as a systemic risk factor. Recent analyses from large population-based cohorts such as TOPMed and ARIC have linked CHIP to diverse vascular phenotypes, including heart failure, stroke, and venous thrombosis. Schuermans et al. (JAMA Network Open, 2024) reported that TET2-driven CHIP doubled the risk of heart failure with preserved ejection fraction (HFpEF), even after adjusting for traditional risk factors^68^. Separately, CHIP is now recognized as a contributor to ischemic stroke and vascular disease^22^. Mechanistic mouse studies support these human hematopoietic Tet2 or Dnmt3a mutant clones promote macrophage dysfunction and accelerate atherosclerosis in animal models, and they impair resolution of inflammation in infarcted hearts^13^.
These findings position CHIP as a potent, age-related inflammatory driver of cardiovascular disease, with implications for early detection, risk stratification, and future anti-inflammatory therapies
Following the discovery of CHIP, researchers began to investigate somatic variants directly in heart tissue. A study by Alexander Hsieh and Sarah Morton estimated that ~1% of CHD probands have a mosaic variant detectable in blood, that contributes to the heart abnormalities^69^. This work emphasized the need to study somatic variants as tissue mosaicism since cardiac tissues account for approximately 5% of all somatic mosaic variant detection in this study.
Studies using deep sequencing techniques revealed the presence of SNVs in cardiomyocytes. Cardiomyocytes harbor a surprisingly high number of SNVs, ranging from 4,000 to 30,000 per cell in healthy individuals^9^. The burden of these somatic SNVs increases with age in the human heart. Further mutational signature analysis indicates that aging results in an increased generation of, or decreased repair of, oxidative DNA lesions, highlighting the heart’s vulnerability to genomic instability. The likelihood of disrupting essential gene function in human cardiomyocytes increases significantly with age, suggesting that age-related cellular dysfunction in aged cardiomyocytes may be partially due to somatic mutations. Few of these SNVs are shared by cardiomyocytes, indicating early somatic changes in the common lineage of the cardiomyocytes.
Another study using single-cell whole-genome sequencing revealed a significant increase in SNVs in cardiac endothelial cells from smokers^70^. These excess somatic mutations exhibited Catalogue Of Somatic Mutations In Cancer (COSMIC) mutational signatures, including SBS4, SBS29, SBS40, SBS92, and ID3, indicative of DNA damage and mutagenesis associated with cigarette smoking, oxidative stress, and formaldehyde exposure^45^.
Furthermore, somatic mosaicism in cardiac tissues has been implicated in various cardiovascular conditions. For example, somatic mutations in genes like FBN1 (associated with Marfan Syndrome) have been identified in cardiac tissues, suggesting a role in structural heart diseases such as aneurysms and dissections^4^. Additionally, mutations in SCN5A and GNAI2 have been implicated in Long QT Syndrome^71–73^ and Idiopathic Ventricular tachycardia (VT)^74–76^.
Using deep sequencing analysis, Ming-Hui Chen and Ryan Doan group established that somatic mosaic variants may play a much larger role even in nonsyndromic thoracic aortic aneurysm with at least 3% of nonsyndromic patients with TAA are likely to have a genetic basis from the somatic variant^77^. In this particular study, although somatic variants were nominally enriched in nonsyndromic patients with TAA, the difference did not survive correction for multiple hypothesis testing. Nonetheless, their pilot results suggest the possibility of somatic variants contributing to the pathology of TAA.
More recently, Hilal et al. applied single-cell whole-genome and transcriptome sequencing to hearts from individuals with ischemic heart disease (IHD) versus healthy controls. They found higher somatic SNV burdens in IHD cardiomyocytes, with distinct mutational spectra suggesting disrupted DNA repair under ischemic stress^78^. These early reports imply that heart disease itself may accelerate genomic aging of cardiomyocytes.
Somatic variants may underlie a wide range of cardiovascular manifestations, from conduction system changes to the development of atherosclerotic disease which develop in Cardiac Resident tissue^71–76^, and clonal hematopoiesis of indeterminate potential (CHIP) mutations in hematopoietic stem cells^79–85^, which may exert detrimental effects throughout the cardiovascular system. Ongoing technological progress is enabling increasingly comprehensive detection and characterization of somatic variants across cardiovascular tissues (Figure 4). The somatic variants burden in individual cardiomyocytes is comparable to that in dividing blood cells, but the consequences in the heart, even a few damaging variants in an essential gene may impair contractile or electrical function. Understanding these patterns is a new frontier; forthcoming single-cell atlases of cardiac tissue (including non-myocytes like fibroblasts and smooth muscle) will clarify how mosaicism contributes to heart failure, arrhythmia, and cardiomyopathy.
The cardiovascular system is a dynamic and metabolically active organ system. Somatic variants in cardiovascular tissues could arise from a convergence of endogenous processes and environmental stressors that accumulate over time. Hemodynamic stress, oxidative injury, environmental factors, and intrinsic genomic maintenance deficiencies can all contribute to the generation and accumulation of somatic mutation^86–88^ (Figure. 5).
Intrinsically, DNA replication errors during cell division are a well-established source of mutations, particularly in proliferative cell populations such as endothelial cells, vascular smooth muscle cells, and cardiac fibroblasts. Even in largely post-mitotic cell types such as cardiomyocytes, mutations can accumulate due to unrepaired DNA damage and oxidative stress over decades of function. Mitochondrial respiration in cardiomyocytes generates high levels of reactive oxygen species (ROS), which can damage nuclear and mitochondrial DNA, leading to base modifications, strand breaks, and replication fork collapse. Deamination of methylated cytosines and spontaneous hydrolysis further contribute to baseline mutational burden in the heart.
Extrinsically, cardiovascular tissues are repeatedly exposed to systemic and local stressors that elevate mutation rates. These include smoking, air pollution, hypertension-induced shear stress, metabolic syndrome, and chronic inflammation. For example, inflammatory cytokines can induce nitric oxide and ROS production in endothelial cells, impair DNA repair fidelity, and alter the epigenetic landscape. Similarly, ischemia and reperfusion injury, common in coronary artery disease, create bursts of oxidative damage that may overwhelm DNA repair pathways. Over time, such insults can give rise to clones of cells with advantageous or neutral mutations that expand through mechanisms such as tissue remodeling, fibrosis, or compensatory hypertrophy. Environmental exposures such as cigarette smoke, for instance, contain mutagenic and carcinogenic compounds such as polycyclic aromatic hydrocarbons (PAHs), acrolein, and formaldehyde, which induce DNA adducts and oxidative lesions in vascular cells^89^.
These genotoxins may penetrate endothelial and smooth muscle layers, initiating atherogenic or fibrotic pathways through mutation accumulation. Similarly, exposure to air pollutants, especially fine particulate matter (PM2.5), ozone, and nitrogen dioxide has been associated with increased cardiovascular morbidity, with chronic inhalation leading to systemic inflammation and DNA damage^90,91^. These pollutants not only generate ROS but also alter epigenetic landscapes, further exacerbating genomic instability^92^.
Recent single-cell and whole-genome analyses of human cardiovascular tissues have revealed mutational signatures indicative of aging (e.g., SBS1, SBS5), oxidative damage (e.g., SBS18), and defective DNA repair (e.g., mismatch repair deficiency), reinforcing the notion that multiple mutagenic forces operate concurrently in the heart and vasculature. These forces may also vary spatially within tissues; for instance, border zones of infarcted myocardium or regions of vascular turbulence may serve as focal points for clonal expansion.
Age is perhaps the most universal and non-discriminatory driver of somatic mutation accumulation in cardiovascular tissues. Over time, DNA damage accrues from both intrinsic (e.g., ROS) and extrinsic sources, increasing mutation burden in long-lived cardiomyocytes. Single-cell genomic studies have revealed that aged cardiomyocytes harbor substantially more somatic single-nucleotide variants (SNVs) and structural variants than those from younger individuals^9^. Compounding this accumulation is an age-related decline in DNA repair fidelity. Pathways such as base excision repair (BER), nucleotide excision repair (NER), and homologous recombination (HR) become progressively less efficient, allowing damaged DNA to persist and propagate^93^. These mutations may disrupt genes essential for contractility, ion homeostasis, or metabolic function, contributing to functional deterioration and disease susceptibility in the aging heart.
In addition to nuclear somatic mutations, mitochondrial DNA (mtDNA) mutations accumulate in the aging heart and vasculature, contributing to cardiac dysfunction through impaired energy metabolism, increased oxidative stress, and altered mitochondrial dynamics. Numerous mtDNA mutations have been associated with atherosclerosis, ischemic heart disease, and heart failure, and distinct pathways of mitochondrial and nuclear genomic aging require different analytic and therapeutic approaches^94–96^.
Thus, dissecting the full spectrum of mutational processes in cardiovascular tissues is not only critical for understanding disease pathogenesis but also for identifying potential therapeutic windows, such as targeting inflammatory mediators or enhancing DNA repair, to limit the accrual of harmful somatic variants across the lifespan.
Cardiomyocytes are among the longest-lived cells in the human body and are largely post-mitotic. This long lifespan, high metabolic rate, and limited regenerative capacity render cardiomyocytes particularly vulnerable to the gradual buildup of somatic variants over time. These variants contribute to somatic mosaicism, cell-to-cell genomic heterogeneity, that may have functional consequences for cardiac tissue architecture, electrophysiological integrity, and stress responses. Alongside burden, aging cardiomyocytes also exhibit features of cellular senescence, including DNA damage response activation, mitochondrial dysfunction, and secretion of pro-inflammatory and pro-fibrotic factors (i.e., the senescence-associated secretory phenotype, or SASP). DNA damage caused by accumulated variants may trigger senescence pathways, while senescent cardiomyocytes may fail to effectively repair genomic lesions, creating a feedback loop that promotes tissue dysfunction. Polyploidization, a known feature of cardiomyocyte aging, may buffer against some mutational damage by providing redundant genome copies, but it may also limit proliferative renewal and exacerbate transcriptional noise. Somatic variants in genes regulating calcium handling, ion channel function, or structural integrity may confer subclinical dysfunction in a subset of cells, which, when combined with senescent phenotypes, could lead to localized arrhythmogenesis or impaired contractility.
Polyploidy is a condition in which cells contain more than two complete sets of chromosomes. In humans and other mammals, cardiomyocytes often become polyploid during postnatal development^97^. This leads to cells with DNA content of 4N, 8N, or even higher. These features present both challenges and opportunities in the study of somatic mosaicism within cardiac tissue. Traditional variant-calling algorithms assume diploidy and rely on allele frequencies around 50% (heterozygous) or 100% (homozygous) for reliable variant identification. However, in polyploid cells, allelic ratios can vary widely depending on the number of mutant alleles, the total ploidy level, and nuclear localization^98^. To address these issues, several technical adaptations have been proposed. Single-nucleus sequencing (snSeq) is a powerful approach for interrogating the genomes of individual nuclei within multinucleated cardiomyocytes^99^. Another approach is ploidy-aware variant calling, which incorporates prior knowledge of chromosomal copy number into the mutation detection pipeline^100^. These tools can model expected variant allele frequencies under various ploidy assumptions and identify statistically significant deviations. For instance, algorithms such as GATK4 have been adapted to handle non-diploid contexts, and new tools specifically designed for single-cell polyploid data are under active development^101^.
Understanding the developmental timing, regulatory mechanisms, and functional consequences of polyploidy and multinucleation is essential for interpreting how somatic mutations accumulate and influence disease pathogenesis in the heart. From a functional perspective, polyploidy and multinucleation have been linked to both protective and pathogenic roles in the heart^102^. On one hand, increased genomic content may confer resilience to stress by providing gene dosage redundancy and buffering against deleterious mutations. This is particularly relevant in long-lived, non-dividing cells like cardiomyocytes, which accumulate DNA damage over time. Polyploid cells may also resist oncogenic transformation by limiting mitotic entry and enforcing cell-cycle arrest, thus maintaining a stable phenotype despite genomic insults. On the other hand, polyploidy may facilitate mosaic expression of genes and lead to cellular heterogeneity even within clonal populations^103^. For example, if somatic mutations occur in only one or a few chromosomal copies, the resulting expression profile may be altered but not abolished, creating gradients of functional change. Moreover, multinucleation can lead to nuclear compartmentalization of gene expression and signaling, allowing mutated and wild-type nuclei to coexist within the same cytoplasm but respond differently to stimuli. These dynamics have important implications for cardiac remodeling and disease progression^97^. The evolutionary rationale for cardiomyocyte polyploidy remains a topic of debate. Some theorize that increased ploidy evolved as a response to the mechanical and metabolic demands of mammalian hearts, providing enhanced transcriptional output and metabolic robustness^103^. Others suggest that it represents a trade-off between genomic stability and regenerative potential, with polyploidy serving as a protective mechanism against replication-induced DNA damage in a high-stress environment.
Future research aiming to map the landscape of somatic mutations in polyploid and multinucleated cardiomyocytes across the human lifespan and in various disease states^99^. High-throughput single-nucleus sequencing, coupled with spatial genomics and transcriptomics, holds promise for elucidating the clonal dynamics and functional consequences of intra-cellular mosaicism. Additionally, computational models that simulate gene dosage effects and nuclear interaction networks may provide insight into how mosaic mutations influence cardiomyocyte behavior and heart function.
Somatic variations have been studied across a wide variety of human tissues, with their prevalence influenced by tissue-specific factors such as replication rate, exposure to environmental agents, and cellular repair efficiency. Mutations in DNA occur in all cells of the human body from conception to death. Skin cells, particularly keratinocytes, are exposed to UV radiation, and mutant clones are prevalent in normal skin, with the frequency rising with age. Mutations in TP53, NOTCH1, and BRAF are frequently discovered in squamous cell carcinoma, basal cell carcinoma, and melanoma^104–106^. In Proteus syndrome, somatic mutations in PIK3CA and AKT1 cause skin overgrowth^107^. Somatic mutations in TP53, CTNNB1, and TERT have been related to hepatocellular carcinoma, which is frequently coupled with hepatitis virus infection and aflatoxin exposure^108^. However, non-cancerous liver tissues also accrue somatic mutations due to their involvement in detoxification and oxidative stress. Mutations in TP53 accumulate in the normal esophageal epithelium, which has been linked to an increased risk of esophageal cancer^109^. Somatic mutations in APC, KRAS, and TP53 cause carcinogenesis in colorectal cancer^110^, whereas CDH1 and TP53 mutations are linked to stomach cancer^111^. Lung epithelial cells are heavily exposed to carcinogens, particularly tobacco smoke, and somatic mutations in EGFR, KRAS, ALK, and TP53 are prevalent in non-small cell lung cancer^112^. Breast tissue, especially during hormonal fluctuations, is more susceptible to mutation accumulation in PIK3CA, TP53, BRCA1, and BRCA2, which are the driver mutations for breast cancer^113^. Normal kidney tissue can accumulate mutations due to its high metabolic activity and vulnerability to oxidative stress. Somatic mutations in VHL and MET are responsible for renal cell carcinoma (RCC)^114^. PTEN^115^, ARID1A^116^, and POLE^117^ somatic mutations are prevalent in endometrial cancer, but TP53 and BRCA1/2 mutations occur often in the ovaries^118^. Pancreatic tissues have a high mutation rate due to secretory activity and inflammation, and somatic mutations in KRAS, CDKN2A, and TP53 are responsible for pancreatic ductal adenocarcinoma (PDAC)^119^.
Variants rates epithelia (colon, liver, skin) often show higher burdens than stroma, reflecting cell turnover and environmental exposures. Notably, single-cell approaches are revealing how mutational signatures differ by tissue – for example, ultraviolet light signatures in skin, tobacco signatures in lung, and reactive oxygen signatures in the heart^39^. Cross-tissue comparisons are becoming feasible. For instance, multi-organ sequencing of the same donors can identify clones shared between tissues (indicative of early embryonic origin) versus those that arose later. Integrating such data will help answer questions Do somatic mutations in vascular cells (e.g., smooth muscle or endothelium) mirror those in blood? Does mosaicism in kidney or brain influence heart health? Emerging cross-tissue atlases will shed light on these questions.
A central goal of somatic genomics research is to connect somatic mosaicism with patient outcomes and therapeutic response, ultimately enabling precision cardiovascular medicine. Emerging studies are beginning to integrate somatic and germline genetic information to improve cardiovascular risk prediction and therapeutic targeting. For example, genotype–phenotype analyses in large cohorts now incorporate both germline and somatic factors. Zhao et al. combined a 531-SNP CHD polygenic risk score (PRS) with CHIP status to predict to evaluate CHD risk in a large Chinese cohort. Individuals with both high PRS and CHIP exhibited a 2.2-fold increased risk of incident coronary events compared to low-risk non-carriers^23^. Notably, when inflammatory pathway SNPs were excluded from the PRS, this excess risk was substantially attenuated, highlighting that inflammation is the shared biological axis between germline and somatic risk factors.
Somatic variants are also being explored as predictive biomarkers in interventional settings. Analysis from the CANTOS trial (anti–IL-1β therapy) demonstrate that patients with TET2-mediated CHIP derived the greatest cardiovascular benefit from IL-1β blockade^13,120,121^. Such findings hint at future precision screening high-risk patients for TET2 or JAK2 clones might identify those who would benefit most from targeted anti-inflammatory or anti-fibrotic therapies. Concurrently, clinical registries are now tracking CHIP status in heart transplant and chemotherapy patients to assess its predictive value for late cardiovascular toxicity.
Large biobanks with linked electronic health records (e.g., UK Biobank, All of Us) provide unprecedented opportunities to study the clinical impact of somatic mosaicism retrospectively. These platforms enable researchers to identify individuals with CHIP or tissue-specific mutations and associate them with outcomes such as myocardial infarction, stroke, heart failure, and arrhythmias. Early results from these analyses reinforce known links (CHIP–atherosclerosis, LOY–mortality) and suggest new ones. For example, mosaic loss of chromosome X in women has been tentatively linked to heart disease risk^122^.
Advances in single-cell and spatial genomics (which was not discussed in this article) now allow researchers to map mutational burden alongside transcriptomic states and inflammatory signatures in affected cardiac and vascular tissues, enhancing mechanistic insight. Moreover, coupling genomic data with electronic health records, biobanked tissues, and imaging studies facilitates population-level analyses while preserving cellular resolution. This integrative approach holds the potential to identify high-risk individuals based on mutational profiles, predict therapeutic responsiveness, and refine risk stratification models, ultimately paving the way for precision cardiovascular care.
Looking forward, integrated genomic risk models that incorporate germline PRS, somatic mosaicism features (e.g., clone number, driver mutations, VAF), and clinical variables are poised to enhance outcome prediction. Preliminary work on such composite “genomic risk scores” is underway and will be accelerated by expanding datasets and improved detection sensitivity. These efforts mark a crucial step toward the realization of precision cardiovascular care informed by both inherited and acquired genetic variation.
The last decade has witnessed a profound transformation in biomedical research, driven by large-scale population biobanks intricately linked to granular, longitudinal electronic health records (EHRs), such as the UK Biobank, which includes approximately half a million participants, and the rapidly growing “All of Us” Research Program in the United States. These research ecosystems weave together genomics, multi-omics, imaging, lifestyle, environmental, and deep clinical data, propelling population-scale studies of somatic mosaicism and revealing how post-zygotic genetic variation shapes human disease, particularly in cardiovascular conditions. Central to this progress is the confluence of scale, diversity, and rich phenotyping, overcoming previous limitations imposed by small cohorts, technical constraints, and lack of systematic clinical tracking. Mosaicism, resulting from spontaneous DNA replication errors, environmental exposures, or aging, leads to clonal expansions of genetically distinct cell populations that range from benign to profoundly pathogenic. Biobanks enable genome-wide scans for rare mosaic variants, even at low variant allele fractions, across ancestrally and contextually diverse populations, supporting genome-first studies and amplifying discovery of ancestry- and environment-specific effects. The longitudinal clinical insights from EHR linkage further multiply their value, allowing for rigorous retrospective analyses where carriers of somatic mosaic mutations can be compared to non-carriers while controlling for confounders. Moreover, the focus on underrepresented populations provides vital data for characterizing ancestry-specific patterns of mosaicism, recognizing that mutation burden and its effects vary with genetics, environmental exposures, and social factors.
In 2023 the National Institutes of Health has launched a program, the Common Fund’s Somatic Mosaicism Across Human Tissues (SMaHT) Network, which aims to utilize state-of-the-art sequencing technologies for detecting somatic variations, including rare variants and variations in repetitive DNA regions, to create a reference catalog of somatic mutations and their clonal patterns across 19 different tissue sites from 150 healthy donors^123,124^. This consortium-led investigation on the mutational landscape across the human body will elucidate the impact of somatic changes in the genome on human biology, including their role in diseases such as cancer, neurological disorders, and other conditions affecting the brain, muscle, skin, and immune system^125^.
The expected impact of the SMaHT consortium on cardiovascular research is substantial. By providing a comprehensive catalog of somatic variants across tissues, including cardiac and vascular tissues, the consortium will offer invaluable insights into the genetic variations that occur within our bodies and their potential impact on cardiovascular disease and aging. This initiative not only enhances our scientific knowledge but also holds promise for improving clinical outcomes through better understanding and management of genetic variations in cardiovascular health.
Another significant research endeavor by the European Union under the Horizon research and innovation programme was launched in 2022, “The SOMATICART project,” which intends to examine the impact of somatic mutations in arterial wall function and their connection to age-related disorders^126^. This initiative seeks to investigate how somatic mutations accumulate in the arterial wall and their possible influence on vascular function. This line of study is critical because it aims to better understand clonality in the vascular wall, which might have important consequences for age-related vascular disorders. One of the key goals of SOMATICART is to compile an atlas of somatic mutations in the human artery wall.
The future marriage of biobanks, multi-modal tissue atlases, and international consortia promises to answer pressing questions about the prevalence, timing, cell-type specificity, and clinical consequences of pathogenic somatic mutations, and to pave the way for translating such knowledge into improved risk stratification, targeted therapy, and ultimately personalized preventive strategies. Large biobanks and their linked, multidimensional data have thus not only made somatic mosaicism studies feasible on an unprecedented scale but have changed the very foundations of cardiovascular genetics and disease risk assessment, with the potential that, within a decade, somatic mosaicism will stand alongside inherited genetics, environment, and lifestyle as a core determinant and target of personalized cardiovascular medicine.
Emerging technologies such as single-cell sequencing, and artificial intelligence are transforming our understanding of somatic mutations and their roles in cardiovascular diseases. By allowing us to detect and characterize rare mutations and clonal populations at unprecedented resolution, these technologies provide new insights into the accumulation and consequences of somatic variation in the heart. Artificial intelligence, in particular, plays a pivotal role by enabling the integration and interpretation of complex genomic and clinical datasets, facilitating the identification of pathogenic mutations, and revealing patterns that may be linked to disease progression and risk. Machine learning models can help to distinguish true biological signals from technical noise, predict clinical outcomes based on mutation profiles, and support the development of novel biomarkers and targeted therapies. Machine learning algorithms can distinguish genuine somatic variants from technical artifacts, identify subtle patterns of clonal expansion, and integrate diverse types of data, including genomics and clinical information, to reveal associations with disease. By harnessing these capabilities, AI can help predict which somatic variants might contribute to disease risk or progression, ultimately supporting the development of more precise diagnostics and targeted therapies. Integration of AI-driven approaches will thus be essential for fully realizing the potential of mosaicism research in cardiovascular health and disease
As large consortia generate expansive reference datasets and improved sequencing technologies become more accessible, these AI-driven approaches will be crucial for translating genetic discoveries into clinical practice. Linking variant data at the single-cell level with patient phenotypes and outcomes will enable the validation of findings and the refinement of predictive models. This integrated framework, combining technological advances and interdisciplinary expertise from genomics, bioinformatics, physiology, and epidemiology, will be essential for harnessing somatic mosaicism as a tool for the early detection, prevention, and treatment of cardiovascular disease. Ultimately, the intersection of single-cell genomics and AI not only deepens our biological understanding but also holds great promise for the development of personalized medicine approaches in cardiovascular health.