Authors: Alexandre Matov
Categories: Original Article, Small RNA, Longitudinal analysis, Disease detection, Early disease
Source: The Journal of Liquid Biopsy
Authors: Alexandre Matov
The current healthcare system relies largely on a passive approach toward disease detection, which typically involves patients presenting a “chief complaint” linked to a particular set of symptoms for diagnosis. Since all degenerative diseases occur slowly and initiate as changes in the regulation of individual cells within our organs and tissues, it is inevitable that with the current approach to medical care we are bound to discover some illnesses at a point in time when the damage is irreversible and meaningful treatments are no longer available.
There exist organ-specific sets (or panels) of nucleic acids, such as microRNAs (miRs), which regulate and help to ensure the proper function of each of our organs and tissues. Thus, dynamic readout of their relative abundance can serve as a means to facilitate real-time health monitoring. With the advent and mass utilization of next-generation sequencing (NGS), such a proactive approach is currently feasible. Because of the computational complexity of customized analyses of “big data”, dedicated efforts to extract reliable information from longitudinal datasets is key to successful early detection of disease.
Here, we present our results for the analysis of healthy donor samples and drug-naïve lung cancer patients’ samples, for which we identify urinary biomarkers demonstrating that small RNAs can pass through the filtration by the kidneys.
We provide a proof-of-principle that it is possible to perform non-invasive health monitoring by sequencing of urinary small RNAs and that traces of neoplastic transformation originating in organs that are not adjacent to the urinary tract, like the lungs, can also be detected in urine.
The cancer incidences and mortality rates worldwide demonstrate that for some cancers, such as lung, stomach, liver, esophagus, and pancreatic cancer as well as leukemia, the diagnosis is almost inevitably linked to a loss of life. The 5-years survival rate for patients with lung cancer is 47% for stage I and can be as low as 2% for stage IV (Fig. 1, taken from [1]), which underscores the importance of early detection. To this end, it will be beneficial to develop molecular diagnostics tests based on nucleic acid biomarkers for monitoring of neoplastic diseases as well as diseases of the nervous system and cardiovascular disease based on the longitudinal analyses of body fluids.Fig. 1Global annual cancer incidences and mortality rates. Age-standardized rate worldwide per 100,000 inhabitants. The figure demonstrates that for some cancers, such as lung, stomach, liver, esophagus, pancreas cancer, and leukemia the diagnosis is almost inevitably linked to a loss of life. The 5-years survival rate for patients with lung cancer, the “biggest killer” (third from top), is 47% for stage I and as low as 2%–20% for stage IV, which underscores the importance of early detection.Fig. 1
In the context of the analysis of samples from healthy individuals, the significance of such approaches will be in providing novel ways for detection of disease. More specifically, it will address the unmet clinical need to identify biomarkers for early detection of degenerative diseases. The identification of such biomarkers is significant because it has the potential for a positive impact on the healthcare system, both in terms of improving patient care and reducing the cost of care.
In the context of the analysis of samples from individuals diagnosed with a disease and possibly undergoing drug treatments, the significance of the approach will be in providing novel ways for treatment evaluation. Correlation between changes in biomarker levels and treatment responses will allow for early detection of a lack of therapeutic response after treatment and ultimately for optimal drug selection. It can also facilitate the discovery of biomarkers for the prediction of relapse. In the long run, longitudinal analysis of nucleic acids can allow for the development of novel targeted drugs. Recent literature has described many examples of miRs which could be targeted in disease [[2], [3], [4], [5]]. Therefore, monitoring of the expression levels of miRs in healthy individuals and in patients undergoing disease treatment will likely provide valuable datasets for the pharmaceutical industry for physicians, ultimately allowing them to select the most efficacious treatment sequence and drug combination for each individual patient.
While methods for the early detection of disease based on the analysis of body fluids [[6], [7], [8]] have become popular, the identification of biomarkers in urine is one of the few completely non-invasive approaches. In fact, it has a clear advantage over other body fluids in the simplicity of collection and processing, as urinary biomarkers are stable for a long period after sample collection. Urine samples with intact protein and RNA content can be preserved at room temperature for up to a year [9]. In addition, panels of miRs in peripheral blood and primary or metastatic tumor samples have been established for several types of tumors and our datasets allowed us to identify many of them in the urine samples we processed (see Appendix).
Overall, the analysis of urine samples is a practical way to screen a large number of individuals at risk of developing a particular degenerative disease, for instance lung cancer. Biomarkers to identify premalignant lung lesions (such as atypical adenomatous hyperplasia, adenocarcinoma in situ, and minimally invasive adenocarcinoma) are scarce. Sodium-glucose transport could potentially identify metabolically active lung premalignancy and early-stage lung adenocarcinoma [10]. Tobacco use is indicated as a cause in more than 80% of lung cancers in the US and UK, and 57.5% of lung cancers in men and 13% in women in China. Lung cancer in never smokers is the fifth most common cause of cancer-related deaths worldwide [11].
The objective of the study is to create a biobank that will generate datasets for research, based on a large number of stored samples, in order to facilitate the statistical analyses of systematic changes in our bodies when we become ill. Longitudinal studies of alterations in nucleic acids, such as miRs, will allow for the identification of panels of biomarkers of degenerative diseases, such as cancers, neurological disorders, and cardiovascular diseases. The biomarker panels will give us the opportunity to perform personalized early detection of disease and alert patients regarding a possible onset of a disease before any symptoms are present.
There are approximately two thousand miRs that have been identified in the literature and catalogued in the miRBase database [12]. Most of these miRs have been associated with the function of specific human organs and tissues, and the dysregulation of many of them has been linked to cancers [13], neurological disorders [14], cardiovascular diseases [15], and other degenerative diseases. It has also been demonstrated that it is possible to discriminate between healthy individuals of different ages as well as between the two sexes based on the analysis of miRs isolated in urine samples [16]. This work determined that the intra-individual variability is considerably smaller than inter-individual differences, which is important in the context of longitudinal analysis of healthy samples. Additionally, it demonstrates that the exact timing of urine collection minimally affects the urine miR transcriptome. Therefore, our hypothesis is that longitudinal changes in miR transcriptomes can be used to detect disease onset before the manifestation of any symptoms. As a concrete preliminary analysis of disease-specific changes in urine samples, we have chosen to highlight the specific case of lung cancer.
Our expectation is that this health monitoring approach will be able to alert healthy individuals of the emergence of a particular disease, create the impetus to discuss this alert with her/his physician, and ultimately take appropriate prophylactic action, for example, to make lifestyle changes or to elect to undergo early treatment. Our hypothesis is that urinary biomarker panels could serve as personalized health monitoring readouts by virtue of periodic measurements. Furthermore, we hypothesize that gradual changes in urinary nucleic acids levels will indicate the onset of disease based on measured changes in comparison to a baseline healthy state and the changes could reveal the type of disease by exhibiting particular changes in personalized panels of biomarkers for each disease. Ultimately, this approach will allow for overall health monitoring and early detection of the onset of any disease, and might also allow for screening for other deviant processes such as psychological disorders.
Our primary methodology is to collect body fluids for analysis. To perform research and identify biomarkers for early detection of disease, every three months we will collect body fluids and store them at negative temperatures with the objective to perform genetic sequencing and statistical analyses of longitudinal data. The urine voids will be collected via a specialized urine container, 250 mL cups custom-manufactured by Norgen Biotek, which will be shipped via a ready-to-mail enclosed envelope or deposited at a local urine collection point.
Our primary objective is the development of methodology for longitudinal analysis of large cohorts of NGS data based on miRs in body fluids for early detection of disease. This objective will consist of developing a computational analysis pipeline designed specifically to identify meaningful patterns and predictive trends in miR levels present in patient urine specimens. Specifically, we will extend the utility of the algorithms used to generate the analyses presented in this contribution, such as principal component analysis (PCA), to integrate consecutive analyses consisting of an incremental number of time points and participants in the cohorts. We will particularly aim at detection of disease prior to symptom manifestation to allow for better treatment options by the discovery, in body fluids such as urine, of biomarker panels, which can be used for the early detection of a plethora of diseases associated with aging. Our secondary objective is to discover personalized biomarker panels that allow patients to access pre-clinical research data and make informed healthcare decisions about the need for lifestyle changes and/or seeking medical advice. Using information theory, each patient's dataset will be analyzed based on longitudinal data, which will take into account the particularities of every individual and the biomarkers can be parsed in an organ/tissue or disease-specific context. Our aim is to provide every individual with the option to make data-driven decisions to slowdown the gradual loss of organ function in normal physiology and the more acute loss of function in disease, and make informed decisions to prevent the progression of age-related degenerative diseases.
The first endpoint of the study is the diagnosis of a degenerative disease (Fig. 2), which provides clinical information and allows us to benchmark our computational predictions. In this context, we present results on the detection of lung cancer in urine samples of 13 stage IV drug-naïve patients, which we compared to 15 healthy individuals without diseases that affect the function of the urinary miRs, such as diabetes and kidney disease. The healthy individuals' sample collection is limited to ages between 45 and 75, while the body fluids collection for individuals already diagnosed with a disease will be universal without any limitations. During the initial data-aggregation phase, when we will rely on a medical diagnosis to establish an endpoint, each participant will be expected to provide information from their general practitioner in regard to the type of diagnosed disease. We anticipate this initial stage to last three years, during which 7% of the participants of age 45 to 75 are expected to develop a degenerative disease. During this stage, our focus will be on identifying biomarkers for early disease detection. If we assume a thousand participants, based on statistics obtained from hospitals (Johan de Rooij, personal communication), within three years 70 participants will develop degenerative diseases of which roughly 22 participants will be diagnosed with a cancer, 4 participants will develop a neurological disorder, and 44 participants will be diagnosed with a cardiovascular disease.Fig. 2Conceptual figure of early detection of disease using longitudinal data; x-axis: Component 1, y-axis: Component 2. The numbers represent PCA axes for 12 different time points acquired over three years by taking a sample every three months. Here, the first eight samples of the same healthy individual cluster together, while for samples taken at time points 9 to 12, the vectors drift further and further away, indicating significant changes in the body in disease. Time-point 12 is when the disease is manifested with symptoms, while time point 9 represents an early detection.Fig. 2
To begin building a repository of patient samples, we will ask the participants to contribute body fluids samples, which will be stored in a biobank, and to complete a health-related questionnaire every three months. This specific sampling interval is selected because it corresponds to the longest time-interval during which a new tumor, for instance, can be expected to remain at surgically resectable levels [17]. Urine voids of 250 mL will be collected in a tube containing a Norgen Biotek preservative. Time and date of collection will be logged on a urine worksheet by the laboratory. Pacific Integrated Handling performs a custom design and delivery of long term low-temperature storage solutions based on precise biobank specifications of the refrigeration units for −20°C (urine voids) and −80°C (blood plasma) samples with seamless backup UPS auxiliary power. The continued collection of samples from the same participants will allow for accumulation of longitudinal datasets. This approach will empower those participants who develop a disease with the ability to provide additional, baseline information to their treating physicians and, once early disease biomarkers are discovered and validated, the collection of additional samples will allow us to perform treatment efficacy monitoring.
The urine voids are collected via a specialized, custom-manufactured urine container 250 mL cups and shipped at room temperature. Urine appearance will be noted in the log file as “Clear”, “Turbid”, “Yellow” or “Brown”. Before long-term storage, the 250 mL collection tube will be centrifuged for 20 min at 4°C and 500 g, and the supernatant will be aliquoted in five 50 mL tubes, while the sediment pellet containing epithelial cells will be discarded, and mixed with the preservative for long term storage at −20°C. Small RNA were extracted from de-identified 50 mL urine samples by Norgen Biotek [9].
All data analysis programs and graphical representation of the results were developed in R and Matlab. The computer code is available for download https://github.com/amatov/DiseaseDetectionUrine. To perform PCA, we used a computational topology software (Perseus [18]) and analyzed longitudinal data of 15 healthy individuals taken every two months for three different time points. PCA grouped data points of the same individuals together in triplets, which was an expected result, but we were also able to detect trends in how the values for every individual change with time (see Results). These results illustrate the transient nature of the miR levels and the importance of understanding normal biological noise and fluctuations in miR levels across the lifespan of healthy individuals.
We also aimed at discovering population-based panels of biomarkers. The small sizes of our datasets precluded the utilization of standard methods and the calculation of statistical significance, for example, by performing differential expression analysis [19]. Because of these considerations, we opted to perform information [20,21] theory-based computation to identify biomarkers. As the size of our datasets increases with time, so will our ability to make meaningful and statistically significant biomarker identifications. For the selection of disease-specific panels, several information theory approaches are available for feature selection, for instance by maximization of the mutual information [22] or by applying the maximum of the minimum criterion [23]. The methodology we utilized for the computation of relative entropy via Kullback-Leibler (KL) divergence [24] is based on an adaptive minimax rate-optimal estimator [25] of the changes in disease from healthy state(s) to cancerous lesion(s) and malignant tumor(s). Consider the KL divergence (1)D(P||Q)≜{∑i=1SpilnpiqiifP≪Q,+∞otherwise,where two patient cohorts are considered, P={p1,…,ps} and Q={q1,…,qs}, over a common set of miRs of length S (S = 947 miRs for our urinary miRs dataset). Utilizing this approach for a selection of significant features, based on two cohort of 13 stage IV lung cancer patients' urine samples and 15 healthy individuals' urine samples, by computing Eq. (1), we identified 20 biomarkers for which the tumors differ from healthy samples the most.
All programs for hierarchical clustering analysis and graphical/image representation of miRs were developed in R: https://github.com/amatov/DiseaseDetectionUrine. The KL divergence method used for biomarker selection is described and validated in [25]. The computer code is in Python and is available for download https://github.com/amatov/FragmentomicsSubclinicalDisease.
Primary and retroperitoneal lymph node metastatic prostate tumor and sternum metastatic rectal tumor tissues were dissociated to single cells using modified protocols from the Witte lab [26]. Organoids were seeded as single cells in three 30 μL Matrigel drops in 6-well plates. Organoid medium was prepared according to modified protocols from the Clevers lab [27] and the Chen lab [28]. We expressed fluorescently tagged microtubule (MT) markers in the organoids (Vid. 1) by using lentivirus-mediated transduction and had best results infecting organoids in an exponential growth phase. To optimize imaging conditions, blasticidin selection was tuned such that only 20%–30% of organoid cells are GFP-labeled [29]. This way, we achieved better contrast of imaging [30] of end-binding proteins 1 (EB1) dynamics [31].
To obtain a single cell suspension, tissues were mechanically disrupted and digested with 5 mg/mL collagenase in advanced DMEM/F12 tissue culture medium for several hours (between 2 and 12 h, depending on the biopsy or resection performed). If this step yielded too much contamination with non-epithelial cells, for instance during processing of primary prostate tumors, the protocol incorporated additional washes and red blood cell lysis [26]. Single cells were then counted using a hemocytometer to estimate the number of tumor cells in the sample, and seeded in growth factor-reduced Matrigel drops overlaid with prostate cancer medium [28]. With radical prostatectomy specimens, we had good success with seeding 3,000 cells per 30 μL Matrigel drop, but for metastatic samples organoid seeding could reliably be accomplished with significantly less cells, in the hundreds. To derive organoids from patient circulating tumor cells (CTCs), liquid biopsy samples of 40 mL peripheral blood would be collected, processed, and plated in a Matrigel-collagen-fibronectin matrix to form organoids similarly to the metastatic breast cancer organoids we cultured from mouse CTCs [32].
To transduce organoids, we modified protocols from the Clevers lab [33] to adapt to the specifics of prostate organoid culture (such as the significant differences in proliferation rates in comparison to colon and rectal organoids). We found out empirically that cells in mid-size organoids (60–100 μm in diameter) infect at much better rates than trypsinized single organoid cells, small organoids, or large organoids. These were the steps we followed to express EB1ΔC-2xEGFP in (1) Add dispase (1 mg/mL) to each well to dissolve the Matrigel at room temperature for 1 h. (2) Spin down (at 1,000 rpm for 4 min) and mix organoids with 10 μL of viral particles (enough for 1 well with three Matrigel drops of 30–40 μL with organoids containing 1–2 million cells) with 10 mg/mL Y27632 ROCK inhibitor and 10 mg/mL polybrene for 30 min. (3) Spin the organoids with viral particles for 1 h at 600 g. (4) Leave the organoids for a 6-h incubation. (5) Spin down and plate in Matrigel. (6) 1 h later, add medium. We used 0.5 mg/mL blasticidin for only one round of medium (3 days) because an increase of the density of labeled cells in the organoids reduced our ability to image MT ends with good contrast.
Organoids were imaged using transmitted light microscopy at 4x and 20x magnification, and phase contrast microscopy at 20x magnification on a Nikon Eclipse Ti system with camera Photometrics CoolSnap HQ2.
All 10x image analysis programs for prostate-specific membrane antigen (PSMA), DAPI, and CD45 segmentation and graphical representation of the results were developed in Matlab and C/C++. The wavelet transform method used, spotDetector, was described and validated in [34] and the unimodal pixel intensity thresholding, in [35]. The computer code is available for download https://github.com/amatov/SegmentationBiomarkerCTC.
We identified the CTC areas as connected-component labeling pixel lists in the epithelial tumor imaging channel (such as PSMA or occasionally cytokeratin (CK) or epithelial cell adhesion molecule (EpCAM)) for which in the nuclear imaging channel there is DAPI labeling with a statistically representative size and a circular shape. At the same time, we require that these cells are negative in the CD45 channel, that is, an absence of a leukocyte marker. Similarly, to detect neutrophils and lymphocytes, we identified clusters of bright pixels in the CD45 channel for which a nuclear area is detected in the DAPI channel and the epithelial labeling is not present. On multiple occasions, we detected double-positive (PSMA+/DAPI+/CD45+ or CK+/DAPI+/CD45+ or EpCAM+/DAPI+/CD45+) and double-negative (PSMA-/DAPI+/CD45-or CK-/DAPI+/CD45-or EpCAM-/DAPI+/CD45-) cells, which we classified in separate bins.
At least, 1917 human miRs have been described in the literature [12], and we have detected 947 of them in urine. The levels of miRs change as a function of age and systematic differences exist between the two sexes [16]. There are three main categories of degenerative diseases – neoplastic (tumors and cancers), nervous system-related (for example, Alzheimer's disease (AD) and Parkinson's disease (PD)), and cardiovascular (for example, hypertension, coronary disease, and myocardial infarction). The etiology of these diseases typically involves some combination of aging and poor lifestyle choices. It is, therefore, conceivable that the collection and analysis of urinary miR samples longitudinally would allow us to delineate changes in physiology associated with a particular disease in its early (asymptomatic) stage.
Fig. 3 shows an example of the miRs we detected in the urine of a healthy individual over four months by collecting and sequencing urine samples every two months, that is at three time points. Note that not all of the miRs we detected in urine were present in all subjects for each of the time points. We aimed at discovering personalized biomarker panels, which would allow a healthy individual to make informed decisions based on an alert about the need for lifestyle changes or seeking medical advice. In this context, we sought to identify meaningful patterns and predictive trends in miR levels present in patient urine and blood specimens. The miRs levels at different time points are intentionally shown separately in Fig. 3 for the reader to see the changes occurring in time, that is, longitudinally. This type of visualization embeds information regarding time progression, which is essential to appreciate that the regulation of our organs could be changing gradually in time – for multiple, or a panel of, biomarkers – toward a disease state.Fig. 3Expression levels (on a logarithmic scale) of 947 miRs sequenced by NGS from urine samples for one of the 15 healthy individuals over three time points. An example of visualization of urine data based on 947 miRs on the x-axis versus the miR abundance levels (in arbitrary units) on the y-axis*.* Time point 1 levels are displayed in red, time point 2 levels are displayed in blue, and time point 3 levels are displayed in green.Fig. 3
The treatment of lung cancer is often hampered by late diagnosis [[36], [37], [38]]. When discovered early, the disease is curable via surgical intervention as adenocarcinoma nodules can simply be resected [39]. For patients with locally advanced or metastatic disease, multiple drug treatment options are offered, yet there are rarely curative regimens available [40]. Immunotherapy seems to be a promising approach for a curative therapy [41], but it too has limitations as there are very few patients who respond to it and the associated toxicity is of considerable concern [42]. While current research efforts in the area of drug development have focused on identifying biomarkers for immunotherapy susceptibility [[43], [44], [45]], our approach to improving the standard of care in lung cancer treatment is to identify biomarkers for early detection of the disease, when curative surgery is still a viable option.
Approximately 190 miRs [[46], [47], [48], [49], [50], [51], [52], [53],[53], [53], [54], [55], [56], [57], [58], [59], [60], [61], [62], [63], [64], [65], [66]] have been linked to lung cancer based on transcriptomics data of specimens derived from primary tumors and/or blood. The top 25 most common biomarkers have been identified in at least three independent screens and some are listed in as many as six or more, suggesting that they are disease biomarkers of high fidelity. The fidelity of biomarkers we select will also be based on whether they have been identified in both the primary tumors and in blood samples, and we will start the health monitoring of lung cancer by accumulating statistics on these 190 biomarkers (see Appendix) with a special emphasis on the top 25 most established ones.
In addition to establishing this miR biomarker panel for lung cancer, we have also identified similar but smaller (that is, based on fewer established blood biomarkers) panels (see Appendix) for colorectal and prostate cancer (in males), breast, uterus, cervical, and colorectal cancer (in females) as wells as urinary pancreatic cancer biomarkers [67]. The necessity to consider different sets of biomarkers for males and females, depending on age, is discussed in the next section. The biomarker lists will be updated as new literature becomes available and whenever our own work allows for the identification of novel biomarkers.
Lung cancer is the leading cause of cancer-related mortality and claims more lives each year than all other major cancers combined due to the fact that lung cancers are generally diagnosed at an advanced stage because patients lack symptoms in the early stages of the disease. Based on our literature search, we identified a panel of 25 high fidelity lung cancer biomarkers, based on miRs each previously linked to lung cancer development, progression, and drug resistance in multiple papers (between three and eight, see Appendix) on the analyses of blood of lung cancer patients and primary lung tumors [[46], [47], [48], [49], [50], [51],[53], [54], [55], [56], [57], [58], [59], [60], [61], [62], [63], [64], [65], [66]]. Below, we are providing information on the effects on lung physiology of these 25 biomarkers (see the list with the 25 miRs in Fig. 4).Fig. 4Longitudinal changes of lung cancer-related urinary miRs in a healthy individual. An example of the up-regulation of 16 urinary biomarkers within a panel of 25 high fidelity lung cancer biomarkers, based on miRs each previously linked to lung cancer development, progression, and drug resistance in multiple (between three and eight) publications on the analyses of blood of lung cancer patients and primary lung tumors. x-axis: list of the biomarkers, from left to (1) miR-21-3p, (2) miR-21-5p, (3) miR-140-3p, (4) miR-140-5p, (5) miR-155, (6) miR-200b-3p, (7) miR-200b-5p, (8) miR-223-3p, (9) miR-223-5p, (10) miR-221-3p, (11) miR-221-5p, (12) miR-145-3p, (13) miR-145-5p, (14) miR-150-3p, (15) miR-150-5p, (16) miR-200a-3p, (17) miR-200a-5p, (18) miR-205-3p, (19) miR-205-5p, (20) miR-210-3p, (21) miR-210-5p, (22) miR-339-3p, (23) miR-339-5p, (24) miR-93-3p, (25) miR-93-5p. y-axis: miR abundance levels; we performed data normalization using the quantile normalization method recommended for NGS data based on single color experiments.Fig. 4
We measured an increase in the levels of 16 of these 25 miRs for one of the healthy individuals in our baseline cohort of 15 people (Fig. 4). We did not consider those miRs for which there have been reports in the literature of a downregulation in disease because their levels might appear decreased in urine for other physiological reasons than lung cancer. The most commonly published biomarkers of lung cancer, which we found to show an increase longitudinally in our urine datasets are miR-21-3p, miR-140-3p, and miR-93-3p. The nicotine-induced miR-21-3p promotes chemoresistance in lung cancer by negatively regulating FOXO3a [68]. miR-140-3p is the most stably expressed plasma exosomes miR in lung cancer patients [69]. miR-93 is up-regulated in a wide variety of malignancies, such as lung cancer, and can serve as a biomarker for drug resistance [70].
The results of this longitudinal analysis indicate that the analysis of urinary miRs in healthy individuals longitudinally has the potential for providing biomarkers for early detection of disease, even for organs, like the lung, that are not adjacent to the bladder.
After establishing that it was feasible to detect published cancer miRs also in urine samples, we sought to investigate whether healthy individuals retain similar patterns of their miR abundance profiles longitudinally via PCA (Fig. 5). To this end, we processed samples from 15 healthy individuals collected every two months for three time points, that is, 45 samples in total. This analysis demonstrated that even if there is a certain level of variability, (i) each three longitudinal samples from the same healthy individual cluster very close together — at least for two of the three time points — and (ii) the longitudinal sample triplets for each of the different healthy individuals form distinct clusters, which are separate from those formed by the sample triplet from other individuals. In Fig. 5, sample triplets for each individual are color-coded in the same color – pink for individual #1, red for individual #2, cyan for individual #3, blue for individual #4, brown for individual #5, yellow for individual #6, orange for individual #7, and so on. This result demonstrates that in normal physiology urinary miR panels can identify the same person longitudinally.Fig. 5Longitudinal analysis of 45 samples from 15 healthy individuals; x-axis: Component 1 (21.4%), y-axis: Component 2 (13.1%). PCA analysis of longitudinal data of healthy individuals taken at three different time points demonstrates the transient nature of the miRs levels. The panel consists of 15 healthy individuals who provided urine void samples in the Netherlands. Each triplet 1-2-3 in the same color in the PCA scatter plot belongs to the same healthy individual over three time points. The time step at which these samples were collected and sequenced was two months. One of the triplets is labeled 2-3-4 because the first sample was not processed correctly and that required the collection of another one at a later, fourth, time point.Fig. 5
We used a PCA library FactoMineR [71] and our analysis results indicated that urinary miRs also allow us to perform multivariate analysis (Fig. 6) for stratification of patient cohorts based on different characteristics such as age, sex, and smoking status individually or in groups. The separation between the five different clusters is expressed in the angular orientation of the vectors as well as in the length (amplitude) of the vectors. The labeling of the samples is as follows – the first letter indicates whether it a lung cancer patient (p) or control/healthy individual (c), next is the one digit or two digit sample number with the cohorts (consisting of 13 stage IV lung cancer patients and 15 healthy individuals), next is the gender (m/f), followed by the age. The last symbol indicates the smoking status that is either indicated as a nonsmoker (n) or is not provided (x), which we assume to represent the smokers. On the lower left quadrant, we see a clear cluster (Cluster I) of three assumed smokers of age 70–80 years. Up, vertical there are two male smokers in their mid-50s (Cluster II). In the upper left quadrant there are four (three of them very short) vectors representing a subset of participants of age mid-60s until 70 (Cluster III). There are three vectors pointing to the right of individuals (possibly all nonsmokers) of age 71 and 72 years (Cluster IV) to which can also be associated sample c9m71n (for which the patient age is the same). In the lower right quadrant, there is a cluster of three healthy people, all nonsmokers, of age 48–50 years (Cluster V). Our analysis indicates that urinary miRs levels can be utilized for multivariate analysis and patient stratification in lung cancer, in particular if patient characteristics, such as co-morbidities (for example, hypertension) and additional medical information from the patient electronic health records is included.Fig. 6Multivariate clustering of lung cancer patients' and healthy individuals' urine data; x-axis: Component 1, y-axis: Component 2. The separation between the clusters is expressed in the angular orientation of the vectors as well as in their length (amplitude). The labeling is as follows in patient (p) or control/healthy (c), next is the two-digit sample number, next is the gender (m/f), followed the age and smoking status – nonsmoker (n) or not provided (x), which we assume to represent the smokers in most cases.Fig. 6
Computing KL divergence has been proposed as an approach for detecting sudden deterioration of complex diseases by dynamical network biomarkers [52], which can be accomplished with the use of one patient sample. The distribution-embedding scheme transforms the data from the observed state-variables with high level of noise to their distribution-variables with low level of noise, and thus significantly reduces signal fluctuations [72]. Increasing the dimensions of the observed data by moment expansion that changes the system from state-dynamics to probability distribution-dynamics allows the derivation of new data in a high-dimensional space, but with much lower noise levels. Single-sample KL divergence has thus been proposed as an approach to detect the early-warning signal of critical transition during the progression of lung cancer. The algorithm was applied to datasets of lung squamous cell carcinoma, lung adenocarcinoma, and acute lung injury [73]. As the method identifies the critical state (or tipping point) at a single sample level and identifies signaling biomarkers, it can be of great potential in personalized pre-disease diagnosis in lung cancer. The genes it selected were involved in lung cancer signaling and cellular processes, such as cytoskeleton organization, chromosome condensation, regulation of cell division, programmed cell death, among others [73], demonstrating a wide clinical applicability of KL divergence calculations.
One of our objectives has been to discover personalized disease-specific miRs biomarker panels based on significant changes in organ or tissue regulation in disease for each individual. As a first step in this regard, we aimed at discovering population-based panels of biomarkers. The preliminary methodology we utilized for computation of relative entropy via KL divergence is based on an adaptive minimax rate-optimal estimator [25] of the changes in disease from healthy state(s) to cancerous lesion(s) and malignant tumor(s). We set our threshold at 1.2 bits divergence empirically and this allowed us to identify 20 biomarkers discriminative of lung cancer (Fig. 7). All 20 biomarkers we identified (see the list with the 20 miRs in Fig. 7) have previously been published in the literature on lung cancer based on analyses of primary lung tumors and blood from lung cancer patients [62,74,75]. The results from our analysis demonstrates the suitability of this method for the identification and selection of disease-specific biomarkers. Even if this computation is based on small sample cohorts and it is not patient-specific, it demonstrates, for the first time, the possibility to detect lung cancer in urine samples.Fig. 7KL divergence between drug-naïve lung cancer patients and healthy individuals. miRs levels in a cohort of 28 participants (13 stage IV drug-naïve lung cancer patients and 15 healthy individuals) allowed for the selection of a panel of 20 discriminative biomarkers out of 947 (743 in these two cohorts) detected by NGS in urine (1) miR-891a-5p, (2) miR-196a-5p, (3) miR-200a-5p, (4) miR-577, (5) miR-141-3p, (6) miR-29c-3p, (7) miR-95-3p, (8) miR-29b-3p, (9) miR-361-5p, (10) miR-429, (11) miR-335-5p, (12) miR-421, (13) miR-628-3p, (14) miR-660-5p, (15) miR-29a-3p, (16) miR-4454, (17) miR-330-3p, (18) miR-194-5p, (19) miR-532-3p, (20) miR-1271-5p. The numbers on the figure show the ranking in divergence (for instance, miR-532-3p is the most divergent). All 20 biomarkers have previously been published in the literature on lung cancer based on analyses of primary lung tumors and blood from lung cancer patients (see Appendix). The miRs are on the x-axis and the KL divergence (in bits) is on the y-axis.Fig. 7
To establish the clinical significance of the 20 urinary biomarkers we discovered, our next step was to identify their concrete function based on previously published studies on the analysis of peripheral blood and tumor tissue samples. Below is information on the effects of the biomarkers on tumor physiology and the miRs are listed based on their divergence levels of healthy individuals from lung cancer patients in the cohorts we analyzed, as shown in Fig. 7. In addition, we demonstrate the discriminative effectiveness of one of the biomarkers, which for miR-200a-5p has a Youden index of 0.85 (Fig. 8).Fig. 8Receiver operating characteristic curve analysis. miR-200a-5p has previously been published in four different papers on the analyses of blood of lung cancer patients and primary lung tumors (see Appendix). Red asterisk indicates the highest Youden's score index (sensitivity + specificity – 1). AUC, area under the curve.Fig. 8
miR-891a-5p downregulation results in suppressed tumor cell migration, proliferation and invasion, while the upregulation of miR-891a-5p leads to enhanced tumor cell activity and promotes tumorigenesis non-small cell lung cancer (NSCLC) [75].
miR-196a-5p is a putative diagnostic biomarker and its expression levels are increased significantly in NSCLC tissues compared with non-tumor adjacent normal tissues [74].
miR-200a-5p exhibits a tumor suppressive role by regulating proliferation and apoptosis in NSCLC, and could target RHPN2 [76].
miR-577 inhibits cell proliferation and invasion in NSCLC [77].
miR-141-3p promotes the proliferation of NSCLC [78] and reduces pulmonary hypoxia and reoxygenation injury [79].
miR-29c-3p is significantly decreased in the plasma of lung cancer patients, especially in those with early-stage lung cancer, and can be used as a biomarker discriminating between NSCLC and small cell lung cancer (SCLC) [80].
miR-95-3p inhibits the invasiveness of metastatic lung cancer through downregulation of cyclin D1 [81].
miR-29b-3p reverses cisplatin resistance by targeting COL1A1 in NSCLC cells [82].
miR-361-5p plays an oncogenic role in lung cancer through the regulation of SMAD2 [83].
miR-429 expression is upregulated in NSCLC, promotes metastasis, and is a potential therapeutic target and diagnostic biomarker [84].
miR-335-5p is significantly decreased in parenchymal lung fibroblasts of smokers [85].
miR-421 confers paclitaxel resistance by binding to the KEAP1 3′-untranslated region and predicts poor survival in NSCLC [86].
miR-628-3p promotes apoptosis and inhibits migration in lung cancer cells by negatively regulating HSP90 [87].
miR-660-5p plays a role in lung tumorigenesis and could be targeted in lung cancer therapy [88].
miR-29a-3p prevents NSCLC tumor growth and cell proliferation, migration, and invasion by inhibiting the Wnt/β-catenin signaling pathway [89].
miR-4454 has the potential to be used as a biomarker for early diagnosis of lung cancer [90].
miR-330-3p promotes metastasis and epithelial-mesenchymal transition via GRIA3 in NSCLC [91].
miR-194-5p downregulates the expression of RAC1 and can prevent metastasis in NSCLC [92].
miR-532-3p inhibits metastasis and proliferation of NSCLC by targeting FOXP3 [92].
miR-1271-5p has the potential to be used as a biomarker for early diagnosis of lung cancer [93].
Several of these 20 miRs have been shown to regulate other organs than the lungs as well. This study, to our knowledge, is the first investigation that indicates the ability to detect a panel of miRs as lung cancer biomarkers in urine. Fig. 9 shows the results of PCA of the patient and healthy cohorts based on the 20 biomarkers only. To improve the separation, we have added color coding of miR-532-3p, which improves the result for patients #4 and #13.Fig. 9PCA plot demonstrates spatial separation between lung cancer patients and healthy individuals; x-axis: Component 1 (55.46%), y-axis: Component 2 (10.02%). Analysis based on a panel of 20 miRs selected via KL divergence out of 947 miRs sequenced by NGS. An additional separation between lung cancer patients and healthy individuals is provided by color coding of miR-532-3p levels. P1-P13 denotes the 13 lung cancer patients, C1t2-C15t2 denotes the 15 healthy (control) individuals in time-point 2 (out of 3), that is, the middle time point. All patient samples, besides P4 and P13, clearly cluster to the left side of the plot. In addition, the level of miR-532-3p (see the color coding) separates P4 and P13 from most of the control samples.Fig. 9
The cancer incidences and mortality rates worldwide demonstrate that for some cancers, such as lung, stomach, liver, esophagus, leukemia and pancreas cancer, the early diagnosis is of critical importance. We will develop algorithmic solutions to facilitate the longitudinal analyses of transcriptomic miR data (about 1,917 miRs have been described in the literature and we have detected 947 of them in urine) from tens of thousands of patients and the early detection of disease. To achieve this end, we will use multivariate data and multi-dimensional clustering analyses as well as information theory approaches based, for instance, on relative entropy computation.
During treatment of lung cancer, the approach presented in this contribution can be combined with an interrogation of patient-derived lung cells, such as CTCs (Fig. 10) and organoids (Fig. 11) [94]. To isolate CTCs in peripheral blood, microfluidic techniques [95] combine size and surface specificity to exploit differences between CTCs and erythrocytes — but capture many leukocytes (Fig. 10A). Methods for CTC enumeration that do not utilize any enrichment are an alternative for collecting all CTCs within a sample (Fig. 10B). Specialized microfluidic devices for hydrogel-based capture and light-induced release of viable CTCs can lead to high purity and high yield CTC isolation [[96], [97], [98]].Fig. 10CTCs captured with or without an enrichment method. (A) Microfluidic device for capture of live CTCs. Scale bar equals 7 mm. The inset shows a zoom-in of a small area of the chip to better see the captured cells, almost all of which are leukocytes. (B) Cells imaged on a coverslip without sample enrichment. Figure PSMA (HuJ591, tumor marker, red), CD45 (leukocyte marker, green), DAPI (nuclear marker, blue). The diameter of the patient CTC in the middle of the image is about 30 μm. Scale bar equals 100 μm. The three insets show detected by our computer vision algorithm [99] CTC in each of the three imaging channels – PSMA, DAPI, and CD45. Magnification 4x. Inset scale bar equals 30 μm.Fig. 10Fig. 11Lung organoid derived from a 70 mg chest-wall resection of a patient with metastatic rectal cancer. Day 25. Magnification 20x, transmitted light microscopy. Scale bar equals 150 μm. The inset shows fluorescent EB1 comets at the tips of polymerizing MTs (see also Vid. 1). Inset scale bar equals 1 μm.Fig. 11
PSMA, a cell surface glycoprotein that is commonly overexpressed by prostate cancer cells relative to normal prostate cells, provides a validated target in tumors of the prostate [100]. The expression of PSMA is elevated in other epithelial tumors [101,102], such as tumors in the kidney [103], breast [104], thyroid [105], liver [106], pancreas [107] as well as the lung [108]. PSMA is expressed in both NSCLC and SCLC [109], which suggests a potential clinical utilization in the treatment of lung cancer. We developed software for image segmentation of PSMA areas [110] in order to extract morphology metrics regarding the changes after drug treatment [99] in CTCs to evaluate drug resistance.
Due to the high number of positive PSMA spots per field, the low contrast, and very low signal-to-noise ratio (SNR), the segmentation of PSMA images is technically challenging. Since the DAPI images have a higher contrast and lower noise in comparison to the PSMA and CD45 channels, we use this channel in the process of detection of PSMA utilizing a multiscale wavelet method [34]. Since we expect the same number of PSMA spots as in DAPI images, we use the wavelet segmentation as “seed” areas based on the pixel lists selected in the DAPI images. To improve the boundaries of the PSMA seed areas, we apply an active contour algorithm [111,112] to the original PSMA pixel values [110]. For CD45 segmentation, we cannot use the same method as the one for segmenting PSMA images because a DAPI label is not present for every CD45 label, that is, not every leukocyte expresses a nuclear marker. The CD45 images exhibit an intermediate image contract, which, even if the SNR is lower than in DAPI images, is higher than that of PSMA images. For these reasons, to segment CD45 images, we apply a damped sine wave [113] filtering as a preprocessing step, which then allows us to segment the low-pass filtered CD45 images using an unimodal pixel intensity thresholding [35]. The combination of the three channels allows the detection of CTCs as PSMA+/DAPI+/CD45– cells in noisy image datasets, which is the case in particular at low magnification (10x). In high-magnification images (63x), the PSMA channel only is sufficient to segment the images and detect CTCs, without having to use the DAPI and CD45 channels [110,114].
To anticipate drug resistance and elucidate mechanisms of resistance, we will derive primary tumor and metastatic organoids from resected tissues from lung cancer patients, and treat them with drug regimens ex vivo. Fig. 11 shows a lung organoid of about ½ mm in diameter we derived after processing tissue from resected rectal cancer metastasis protruding from the sternum. Patient-derived organoids [115] can be labeled with fluorescent markers of key cellular proteins, which deliver a detailed insight into the metabolic changes in patient tumor cells after drug treatment.
We have proposed a computational approach for the selection of drug regimens in solid tumor oncology based on the analysis of EB1 and MT dynamics in patient cells [29,32,94,116]. Novel MT destabilizing drugs have been shown to improve the treatment of advanced or metastatic NSCLC with no driver mutations when used in combination with docetaxel [117], but not when compared to sotorasib in lung adenocarcinomas with KRAS^G12C^ mutations [118]. Our quantitative live-cell imaging algorithm can facilitate patient stratification and the refinement of the drug regimens.
We hypothesize that the modulation of the MT and actin cytoskeleton could sensitize lung cancer cells to immunotherapy. Neoadjuvant anti-PD-1-based immunotherapy might become the standard of care in lung cancer [119], which further underscores the importance of the early detection of lung tumors. Further, a blood-based miR panel has been shown to have a prognostic value for overall survival in advanced NSCLC treated with immunotherapy [120] and as urinary samples are collected in a non-invasive procedure (that is more convenient than a blood draw), prognostic readouts can be obtained this way very frequently.
Our analyses can provide novel tokens to foundation models that perform longitudinal analysis [121] of patient records and improve their ability to predict a future differential diagnosis. Difficult to diagnose diseases can be detected ahead of time when the metrics we propose are added to the model to generate an accurate patient's clinical picture.
The methodology we propose may allow the development of molecular diagnostics tests based on nucleic acid biomarkers for monitoring of neoplastic diseases, diseases of the nervous system, and cardiovascular diseases based on the longitudinal analyses of body fluids. What is novel in this contribution is the description of a methodology to detect urine biomarkers for an organ – the lungs – that is not adjacent to the bladder, such as the prostate, which is most commonly investigated using urine samples. Further, our approach may provide novel ways for treatment evaluation. Urine voids can be collected at home, the procedure we outline is fully non-invasive, the samples are easy to process (as opposed to blood), the biomarkers are stable in urine at room temperature for a week, which allows flexibility in regard to shipment and processing, and urine is known to contain disease biomarkers. Correlation between changes in biomarker levels and treatment responses will allow for the early detection of a lack of response during treatment and ultimately for optimal drug selection. It will also facilitate the discovery of novel biomarkers for the prediction of disease relapse.
In the long run, longitudinal analysis of nucleic acids will allow for the development of novel targeted drugs. Recent literature has described many examples of miRs that could be targeted in disease [[2], [3], [4], [5]]. Therefore, monitoring of their levels in healthy individuals and patients undergoing disease treatment will likely provide valuable datasets for the pharmaceutical industry as well as practicing physicians, ultimately allowing them to select the most efficacious treatment sequence and drug combination for each patient.
We propose the utilization of the described here analytics to generate embeddings [122] for a generative transformer network in which the tokens are the small RNA metrics of the organoids. The embeddings will also include live-cell dynamics [114,123] before and after treatment with different drugs. We will retrain a transformer network with a new set of tokens – with shape descriptors rather than words and with growth and killing curves (that is, sigmoid curves) [110,99] rather than sentences of human speech, that is, we will retrain a LLM or a foundation model with organoid morphology data. This will allow us to train the network to classify each sample as sensitive or resistant and additionally identify subclasses with different mechanisms of resistance. Further, after training a generative transformer network with organoid texture as well as live-cell dynamics datasets, we will be able to model disease progression too. Upon validation, our computer vision approach for analysis of live-cell dynamics can train AI agents [124] and become a clinical tool for drug selection based on LLMs or foundation models.
IRB (IRCM-2019-201, IRB DS-NA-001) of the Institute of Regenerative and Cellular Medicine. Ethical approval was given.
Approval of tissue requests #14–04 and #16–05 to the UCSF Cancer Center Tissue Core and the Genitourinary Oncology Program was given.
The patient blood samples analyzed were from clinical studies with IRB protocols 0804009740 and 0707009283 at Cornell Medicine.
Informed consent was obtained from all subjects and/or their legal guardian(s). The research was performed in accordance with the Declaration of Helsinki.
A.M. conceived the study, performed cell culture, microscopy imaging, and data analysis, prepared figures, and wrote the manuscript.
The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request. Lung cancer miR datasets can be found https://github.com/amatov/DiseaseDetectionUrine/blob/main/LungCancer_miRs.RData.
No funding was received to assist with the preparation of this manuscript.
No conflict.