Authors: Akiko Urabe (Department of Pathology and Clinical Laboratories, National Cancer Center Hospital East, Kashiwa, Chiba, Japan; Department of Gastroenterology and Endoscopy, National Cancer Center Hospital East, Kashiwa, Chiba, Japan), Masahiro Adachi (Department of Pathology and Clinical Laboratories, National Cancer Center Hospital East, Kashiwa, Chiba, Japan), Naoya Sakamoto (Division of Pathology, Exploratory Oncology Research & Clinical Trial Center, National Cancer Center, Kashiwa, Chiba, Japan), Motohiro Kojima (Department of Pathology and Clinical Laboratories, National Cancer Center Hospital East, Kashiwa, Chiba, Japan), Shumpei Ishikawa (Division of Pathology, Exploratory Oncology Research & Clinical Trial Center, National Cancer Center, Kashiwa, Chiba, Japan), Genichiro Ishii (Department of Pathology and Clinical Laboratories, National Cancer Center Hospital East, Kashiwa, Chiba, Japan), Tomonori Yano (Department of Gastroenterology and Endoscopy, National Cancer Center Hospital East, Kashiwa, Chiba, Japan), Shingo Sakashita (Division of Pathology, Exploratory Oncology Research & Clinical Trial Center, National Cancer Center, Kashiwa, Chiba, Japan)
Categories: Original Article, artificial intelligence, epithelium, esophageal cancer, submucosa, surface histomorphology
Source: Cancer Science
Doi: 10.1111/cas.16426
Authors: Akiko Urabe, Masahiro Adachi, Naoya Sakamoto, Motohiro Kojima, Shumpei Ishikawa, Genichiro Ishii, Tomonori Yano, Shingo Sakashita
The depth of invasion plays a critical role in predicting the prognosis of early esophageal cancer, but the reasons behind invasion and the changes occurring in invasive areas are still not well understood. This study aimed to explore the morphological differences between invasive and non‐invasive areas in early esophageal cancer specimens that have undergone endoscopic submucosal dissection (ESD), using artificial intelligence (AI) to shed light on the underlying mechanisms. In this study, data from 75 patients with esophageal squamous cell carcinoma (ESCC) were analyzed and endoscopic assessments were conducted to determine submucosal (SM) invasion. An AI model, specifically a Clustering‐constrained Attention Multiple Instance Learning model (CLAM), was developed to predict the depth of cancer by training on surface histological images taken from both invasive and non‐invasive regions. The AI model highlighted specific image portions, or patches, which were further examined to identify morphological differences between the two types of areas. The 256‐pixel AI model demonstrated an average area under the receiver operating characteristic curve (AUC) value of 0.869 and an accuracy (ACC) of 0.788. The analysis of the AI‐identified patches revealed that regions with invasion (SM) exhibited greater vascularity compared with non‐invasive regions (epithelial). The invasive patches were characterized by a significant increase in the number and size of blood vessels, as well as a higher count of red blood cells (all with p‐values <0.001). In conclusion, this study demonstrated that AI could identify critical differences in surface histopathology between non‐invasive and invasive regions, particularly highlighting a higher number and larger size of blood vessels in invasive areas.
According to the Global Cancer Observatory 2020, ^1^ esophageal cancer accounts for 3.1% of new cancer cases globally and is the sixth leading cause of cancer‐related deaths. ESCC or adenocarcinoma is the most common esophageal cancer worldwide. ^2^
Treatment options for esophageal cancer include surgery (including endoscopic resection (ER)), chemoradiotherapy (CRT), and chemotherapy. Of these options, surgery and CRT are the most curative treatments, whereas ER is the least invasive. ER is appropriate for early‐stage T1a cancers limited to the epithelium (T1a‐EP) or superficial lamina propria mucosa (T1a‐LPM), which have a low risk of lymph node metastasis. The indication for ER is expanded to patients with more advanced T1a lesions that extend deep into the lamina propria (T1a‐MM) or T1b lesions with minor SM invasion (<200 μm). The risk of lymph node metastasis is 10%–20%. In these cases, additional CRT or surgery is performed for T1b or greater depths with pathological evaluation. The decision between endoscopic and surgical resection depends on the invasion depth, risk factors, and extent of lymph node metastasis. ^3^
The rate of lymph node metastasis varies with the depth of involvement and the treatment required. In other words, the prognosis depends on the depth of invasion. However, there is limited knowledge of why invasion occurs and why it is associated with a poor prognosis. In this study, we would like to help clarify the mechanism by looking at the differences in the tissues of the non‐invasive and invasive areas.
AI has become a widespread tool used also in endoscopy and pathology. ^4^ , ^5^ , ^6^ , ^7^ , ^8^ , ^9^ , ^10^ Digital pathology images are used for pathological diagnosis. ^11^ Particularly, WSI, which refers to the digitized images of glass‐slide specimens, enables microscopic observations on a personal computer using WSI‐specific analysis software. These images are also used for tissue analysis, image processing, and annotation. ^12^
CLAM ^13^ , ^14^ is a weakly supervised machine learning model developed for histopathological image classification. The model combines MIL with an attention branch network. MIL has been applied in pathological diagnosis because it allows classification on a case‐by‐case basis, rather than on a patch‐by‐patch basis. ^15^ , ^16^ , ^17^ , ^18^ Attention branch network refers to the mechanism through which a neural network focuses on a specific part of the input data and assigns different levels of importance to different elements. Thus, the model not only classifies but also extracts patches that significantly influence classification. In this study, we applied the CLAM model to investigate potential differences in the histopathology of superficial tissues between areas with and without invasion in esophageal tumors. In particular, using endoscopically resected ESCC specimens at SM depth, we aimed to investigate whether the model could predict depth by training AI on superficial tissue samples from regions without and with invasion (referred to as EP and SM, respectively) and, if so, which areas to focus on.
We included 75 patients (Table S1) who underwent ESD for ESCC with a pathological depth of SM1 or SM2 at the National Cancer Center Hospital East, between January 1, 2019, and March 15, 2023. Cases without superficial exposure and cases with slides borrowed from other hospitals (n = 10) were excluded. This study was approved by the Ethical Review Committee of the National Cancer Center Hospital East (2023‐029).
First, ESD specimens were fixed in 10% formalin and sectioned parallel to the short axis in 4‐mm slices, which were used to prepare H&E‐stained sections. Next, H&E slides containing the EP and SM portions in subject cases were selected. For EP assessment, we then selected the sections with the EP portion farthest from the SM portion of the same lesion. In cases where it was difficult to determine the furthest portion at both ends, both slides were used. Then, when more than one SM portion was present, all regions were used for evaluation. In total, 156 slides were used (EP, 82; SM, 74). Finally, the specimens were scanned at ×40 magnification using NanoZoomer s360 (Hamamatsu Photonics, Shizuoka, Japan) to obtain WSIs.
To examine the differences in surface histology between the EP and SM sections, we annotated the surface portion of the lesion on each slide and extracted histology using QuPath version 0.4.0, a publicly available annotation tool for digital slides ^19^ (Figure S1). The annotation sites in this study were cancerous in the superficial layer for both the EP and SM regions. The annotation size was limited to a maximum of 4 mm, matching typical biopsy specimen dimensions; hence, the AI analysis would be relevant for clinical biopsy samples. For SM invasion, the maximum depth from the superficial mucosal layer was always annotated. For SM lesions with multiple deep portions below the MMs, each distinct area was separately annotated. The most superficial cancerous mucosal sections located furthest from areas of maximum‐depth invasion were selected and annotated for non‐invasive lesions. For each lesion, the superficial mucosal layers from the sections with and without invasion were extracted and utilized as training data. Two pathologists, including a specialist, double‐checked the annotation sites to ensure accuracy.
The CLAM model was used for AI development. When preprocessing the slide images, CLAM segmented the image into small patches. The annotated slide images were cropped into 128 × 128‐, 256 × 256‐, and 512 × 512‐pixel regions to create patches (Figure 1A). The sections with SM invasion were labeled as 1 (SM), and the EP sections were labeled as 0 for model training. For model development and evaluation, a 10‐fold Monte Carlo cross‐validation strategy was used, in which the training, validation, and testing subsets were randomly derived from the cohort. Each fold was divided randomly into training (80% of the cases), validation (10%), and testing (10%) sets. Performance was assessed using the area under the receiver operating characteristic curve (AUC) and ACC. The AUC and ACC values were compared to determine the size (128, 256, or 512 pixels) used to create the patches. The model was trained using the adaptive moment estimation optimizer with a learning rate of 2 × 10^−4^. We used the default algorithm for other parameters and did not perform data augmentation. The training process ended at 200 epochs if validation loss did not decrease from its previous minimum for 20 consecutive epochs.

For each patch used for inference in CLAM, an attention score was calculated based on the proportion of contributions to the prediction. The patch with the highest attention score in the model and the best AUC value was extracted. They were then examined for features based on which the inference model classified them as SM or EP. In this study, the total area of the nuclei, the major and minor diameters and circumference of cancer‐cell nuclei, the presence of blood vessels, and the presence or absence of distinct nucleoli in each patch were examined. In patches with blood vessels, the total area and number of erythrocytes and variability in findings related to nuclei were also evaluated. The presence of blood vessels and distinct nucleoli were assessed visually. For erythrocytes and vascular cavities, manual annotation was performed and analyzed using ImageJ. ^19^ The last analysis of tumor nuclear parameters was performed after manual annotation of the tumor area using QuPath. ^20^ Regarding the work performed by people, all tasks were conducted by two pathologists.
The analysis was performed on an Ubuntu 20.04 Linux system with an A100 GPU (NVIDIA, Santa Clara, CA, USA). Statistical analysis was performed using R version 4.3.0 (The R Project, Vienna, Austria), with a chi‐squared test for categorical variables and a non‐operating characteristic (ROC) curves were equal variance t‐test (Welch test) for continuous variables for comparison between the two counts. The F‐test was used to test for variance. Logistic regression analysis was performed for multivariate analysis. p‐values <0.05 were considered statistically significant.
We examined the possibility of EP‐ and SM‐like features extracted by the AI used to assign EP or SM. For MM, 89 cases diagnosed by ESD between January 1, 2019, and March 15, 2023, were utilized. As with SM cases, the superficial layer of the deepest esophageal cancer portion was annotated, with multiple annotations if multiple deep spots were present. A total of 103 muscularis mucosa (MM) slides were annotated. Because there were many LPM cases, 84 LPM cases diagnosed by ESD between January 1, 2020, and March 15, 2023, were used. In total, 103 LPM slides were annotated.
Based on CLAM, a prediction model for tumor depth (EP or SM) was created from surface histology. In total, 128‐, 256‐, and 512‐pixel models were created (Figure 1A), and the test AUC and ACC values were compared. The mean test AUC values for 128, 256, and 512 pixels were 0.69, 0.87, and 0.72, respectively (Figure 1B). The mean test ACC values were 128, 256, and 512 pixels were 0.61, 0.79, and 0.64, respectively (Figure 1C). The 256‐pixel model was superior to the 128‐ and 512‐pixel models in terms of both test AUC and ACC values; therefore, we created 256‐pixel models. All the 10 models trained at 256 pixels achieved AUC >0.68 and ACC >0.68 on the test set (Table 1). The models with the highest AUC (1.0) and ACC (0.938) values for the test data were selected for further analysis. ROC curves were drawn using matplotlib, a Python library.
For each patch, the CLAM model calculated an attention score based on its contribution to prediction. Representative patches that received the highest scores were extracted from the CLAM model with the best AUC performance (256‐pixel model). These high‐scoring patches (Figures 2 and S2) were assumed to contain key distinguishing features, dense vasculature, and red blood cell numbers in the SM patches (Figures 2B and S2) versus higher nuclei numbers in the EP patches (Figures 2A and S2), and were used for further histological evaluation.

The attention score was utilized to extract patches with histological features that the AI model for depth prediction focused on. Significant differences were observed between the EP and SM patches in terms of the presence of blood vessels. However, no significant differences were observed in nucleolus clarity. The AI model focused on patches without blood vessels in the EP region, while the proportion of blood vessels increased in the SM region (Table 2 and Figure S3).
As the presence of blood vessels was a significant feature of the AI model, further comparisons were made between the SM and EP patches that contained vessels (Table 2). The total blood vessel area was significantly larger in SM patches than in EP patches (p < 10^−5^; Figure 3A,B). Similarly, red blood cell counts were significantly higher in SM patches than in EP patches (p < 10^−5^; Figure 3A,C). These findings indicate that the AI model may predict invasion depth by focusing on differences in blood vessel morphology between superficial and deep areas.

Morphological variations in the nuclei (Figure 4A) in the EP and SM patches were compared using the AI high‐interest images. T‐tests were used to compare mean values related to the size of nuclei, including area, maximum and minimum diameter, and circumference. F‐tests were used to assess differences in the variance. Although none of the results were significantly different with respect to the mean, variations in the area and minimum diameter of nuclei were significantly different in the EP and SM samples (F
81,71 = 0.542, p = 0.004, and F
81,70 = 0.669, p = 0.04, respectively; Figure 4B–E). These results suggest that regions with invasive SM characteristics may contain more diverse micronuclei and larger nuclei than EP regions.

In the previous study of EP and SM lesions, the probability of being determined as EP was 97.6% and 8.6%, respectively, whereas, in the present study, 57.3% of LPM and 37.9% of MM were classified as EP (Figure 5A). EP and SM patches are presented in Figure 5B. SM patches contained more blood vessels than EP patches as well as when SM and EP were compared.

We demonstrated that AI could predict the depth of invasion from surface histological findings alone. We also obtained an understanding of the structural differences between infiltrated and non‐infiltrated areas.
This study has three major implications. First, AI was able to not only classify, but also identify differences in morphology that were difficult for pathologists to notice. For instance, the AI model focused on patches of increased vascular density in the SM region; however, this is not a point of particular interest to many pathologists when diagnosing biopsy material.
Second, the ability of AI to infer SM invasion from the evaluation of surface tissues suggests that detectable changes have already occurred in these layers. Although endoscopic observation is limited to the surface epithelium, invasion depth can potentially be predicted through careful examination of surface histological features. Thus, superficial tissues may contain subtle evidence indicative of deep SM spread that AI models can leverage for diagnosis and staging.
Third, a detailed analysis of the patches with high attention scores on which the AI model focused revealed an emphasis on blood vessel features rather than cancer‐cell morphology. This study used CLAM to determine the patches that contributed the most to model predictions. CLAM learns by extracting only the most informative patches with high attention scores. Examination of these patches provided insight into the histological features that the model used to predict invasion depth. The prominence of blood vessel differences suggests that the AI model may predict depth by focusing on vascular changes between areas with and without invasion. This is a remarkable finding because pathologists traditionally do not emphasize vascular morphology in any cancer staging system. However, pronounced vascular changes, visible even in superficial tissues, may serve as key signals for predicting underlying invasion. AI revealed overlooked histological differences with clear diagnostic significance. The key difference between areas with and without invasion that enabled AI‐based depth prediction was quantitative information about the tumor stroma, particularly vascular morphology.
Although these observations regarding vascular architecture are interesting pathological findings, endoscopists routinely use vascular patterns to predict invasion depth based on NBI or other image‐enhanced endoscopy with magnification. ^21^ , ^22^ The current results suggest that endoscopic visualization of vascular changes may reflect underlying histopathology.
When metrics related to the size of nuclei were compared, no differences were found in size or diameter, but variation analysis revealed differences in heterogeneity, whereby regions with invasion showed a greater diversity of small and large nuclei.
A discussion of biological mechanisms is given here as an addition. It is said that the malignant transformation of cancer is essentially caused by changes in cancer cells, such as the accumulation of genetic mutations. In the present results, although the variation in size was observed in cancer cells between non‐invasive and invasive areas, there was no significant difference in nuclear size, long diameter, or roundness, and a difference was observed in blood vessels. This suggests that vascular changes may be involved in cancer invasion. Although it is unclear whether this is the cause or the result, the fact that similar characteristics are also observed to some extent in LPM and MM cancers, even if not SM, suggests that changes in the surrounding area, such as blood vessels, may be responsible for the invasion. ^23^ , ^24^ , ^25^ , ^26^
Ultra‐high magnification endocytoscopy ^27^ , ^28^ allows up to 580‐fold magnification compared with the previous maxima of ×80–100 to evaluate histological findings without biopsy. Images captured by the endocytoscopic system have been used to evaluate the differential diagnosis of ESCC and benign lesions using AI, and favorable preliminary results have been reported. ^29^ While evaluating the invasion depth in ESCC using this system, patterns regarding blood vessels and individual cells should be resolved, even by detecting previously invisible nuclei. The histological features highlighted by AI in the present study, such as vascular and nuclear patterns, may be observed using endoscopy in the future. Further studies will be required, as ultra‐high‐resolution endoscopy develops, to determine the optimal diagnostic criteria based on histopathological correlations, such as those demonstrated in this study.
The detection of differences that pathologists did not observe using CLAM is significant. The CLAM model learns only from patch images, which improves accuracy, and incorrect patches do not advance learning. By analyzing the patches with high attention scores in an accurately trained model, the histological differences between EP and SM samples that drive model predictions are revealed. Thus, we used AI's patch‐level focus as a tool to uncover distinguishing features that a pathologist may not be aware of. AI guides us to subtle but predictive vascular differences between areas with and without invasion that pathologists do not routinely examine but that can aid in diagnosis. This demonstrates that the interpretation of an optimized AI model can reveal new perspectives on diagnostic tissue morphology that are invisible to the human eye.
This study demonstrates the potential of biopsy histopathology to quantitatively diagnose the depth of invasion and provide a qualitative diagnosis of cancer. Although biopsies typically only provide information on whether cancer is present, the ability to determine depth from biopsy analysis can better guide clinical decision‐making and treatment planning. However, challenges remain in translating these findings into clinical biopsy interpretation, such as uncertainty regarding whether the biopsy sample adequately captures the invasive tumor front. In this study, the AI model assessed the ability to accurately diagnose the depth of tumor invasion in the histology of biopsy specimens but was not applicable to biopsy material. Specifically, it was tested using pathology images from 89 cases of biopsy material collected between 2020 and 2023, and their depth was determined through subsequent ESD and surgery. However, limited findings were obtained (Figure S4), with an AUC value of 0.49 indicating that the model's performance was comparable with random guessing and could not effectively distinguish between the two classes. This is due to differences in how ESD and biopsy specimens are viewed, potentially affecting the model's adaptability. Nevertheless, this study represents an important step toward leveraging standard biopsy procedures for more precise and quantitative staging through AI‐assisted pathological assessment of surface tissues.
This study has limitations. First, it aimed to identify histological differences and did not validate a system that can predict depth using AI. Hence, multicenter studies and the creation of independent test datasets are necessary. Second, we focused on blood vessels and found corresponding differences but we could not exclude the possibility that the pathologist may have missed other differences. Last, although we used a portion of the tissue, we did not evaluate the entire cancer superficially, and the possibility that selection bias due to the annotation site may have affected the results cannot be ruled out.
In conclusion, this study demonstrated the potential of AI‐based analysis of superficial histology to predict tumor invasion depth. The most important predictive features were the pronounced vascular differences between invasive and non‐invasive stromal areas, rather than epithelial tumor‐cell features. Therefore, standard biopsy pathology can be used for quantitative assessments of the invasion status. More broadly, this study highlights the underappreciated stromal vascular changes associated with invasion that may reflect important biological differences in tumor aggressiveness.
Akiko Urabe: Conceptualization; data curation; formal analysis; investigation; methodology; resources; software; supervision; validation; visualization; writing – original draft; writing – review and editing. Masahiro Adachi: Data curation; formal analysis; methodology; writing – review and editing. Naoya Sakamoto: Writing – review and editing. Motohiro Kojima: Writing – review and editing. Shumpei Ishikawa: Methodology; writing – review and editing. Genichiro Ishii: Writing – review and editing. Tomonori Yano: Writing – review and editing. Shingo Sakashita: Conceptualization; funding acquisition; methodology; project administration; software; validation; visualization; writing – review and editing.
This work was supported by JSPS KAKENHI (grant nos. JP20K22859 and JP21K06899). A part of this study was supported by the National Cancer Center Research and Development Fund (grant no. 2021‐A‐07).
Dr. Genichiro Ishii is an Editorial Board member of Cancer Science. Other authors do not have any COI to declare. Ishii Genichiro received research grants from Daiichi Sankyo, Inc., Ono Pharmaceutical Co., Ltd., Noile‐Immune Biotech, Takeda Pharmaceutical Company Limited, Sumitomo Dainippon Pharma Co., Ltd., Nihon Medi‐Physics Co., Ltd., and Indivumed GmbH, H.U. Group Research Institute and consulting fee from Takeda Pharmaceutical Company Limited. Tomonori Yano received lecture fees and research grants from Olympus.
Approval of the research protocol by an Institutional Reviewer Board: This study was approved by the Ethical Review Committee of the National Cancer Center Hospital East (2023‐029) and conforms to the provisions of the Declaration of Helsinki.
Informed Consent: N/A.
Registry and the Registration No. of the study/ N/A.
Animal Studies: N/A.