Authors: Parhat Yasin, Yasen Yimit, Abuduainijiang Abulimiti, Haopeng Luan, Cong Peng, Maihemuti Yakufu, Xinghua Song
Categories: Article, Knee injuries, Artificial intelligence, Deep learning, Multi-label classification, AI-Driven application, Health care, Medical research
Source: Scientific Reports
Authors: Parhat Yasin, Yasen Yimit, Abuduainijiang Abulimiti, Haopeng Luan, Cong Peng, Maihemuti Yakufu, Xinghua Song
Knee abnormalities, such as meniscus tears and ligament injuries, are common in clinical practice and pose significant diagnostic challenges. While traditional imaging techniques—X-ray, Computed Tomography (CT) scan, and Magnetic Resonance Imaging (MRI)—are vital for assessment. However, X-rays and CT scans often fail to adequately visualize soft tissue injuries, and MRIs can be costly and time-consuming. To overcome these limitations, we developed an innovative AI-driven approach that allows for the detection of soft tissue abnormalities directly from X-ray images—a capability traditionally reserved for MRI or arthroscopy. We conducted a retrospective study with 4,215 patients from two medical centers, utilizing knee X-ray images annotated by orthopedic surgeons. The YOLOv11 model automated knee localization, while five convolutional neural networks—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—were adapted for multi-label classification of eight meniscus tears (MENI), anterior cruciate ligament tears (ACL), posterior cruciate ligament injuries (PCL), medial collateral ligament injuries (MCL), lateral collateral ligament injuries (LCL), joint effusion (EFFU), bone marrow edema or contusion (CONT), and soft tissue injuries (STI). Data preprocessing involved normalization and Region of Interest (ROI) extraction, with training enhanced by spatial augmentations. Performance was assessed using mean average precision (mAP), F1-scores, and area under the curve (AUC). We also developed a Windows-based PyQt application and a Flask Web application for clinical integration, incorporating explainable AI techniques (GradCAM, ScoreCAM) for interpretability. The YOLOv11 model achieved precise knee localization with a mAP@0.5 of 0.995. In classification, ResNet152 outperformed others, recording a mAP of 90.1% in internal testing and AUCs up to 0.863 (EFFU) in external testing. End-to-end performance on the external set yielded a mAP of 86.1% and F1-scores of 84.0% with ResNet152. The Windows and web applications successfully processed imaging data, aligning with MRI and arthroscopic findings in cases like ACL and meniscus tears. Explainable AI visualizations clarified model decisions, highlighting key regions for complex injuries, such as concurrent ligament and soft tissue damage, enhancing clinical trust. This AI-driven model markedly improved the precision and efficiency of knee abnormality detection through X-ray analysis. By accurately identifying multiple coexisting conditions in a single pass, it offered a scalable tool to enhance diagnostic workflows and patient outcomes, especially in resource-constrained areas.
The online version contains supplementary material available at 10.1038/s41598-025-21895-6.
Knee trauma often results in soft tissue injuries that require prompt diagnosis and appropriate management to prevent long-term complications such as post-traumatic osteoarthritis and chronic pain^1,2^. The knee is a complex structure, made up of bones, cartilage, ligaments, and tendons that all work together to provide stability and enable activities like walking, running, and jumping. The knee’s complex structure—comprising bones, cartilage, ligaments, and tendons—makes it particularly vulnerable to injuries that can significantly impact mobility and quality of life^3^. When left undiagnosed or improperly treated, these injuries frequently lead to persistent functional limitations, with post-traumatic osteoarthritis affecting millions globally and causing substantial disability^4,5^. Early and accurate detection of knee abnormalities thus remains crucial for optimizing patient outcomes and preventing progressive joint deterioration^6^.
Currently, magnetic resonance imaging (MRI) and arthroscopy serve as the gold standards for diagnosing soft tissue knee injuries, particularly those involving the anterior cruciate ligament (ACL), posterior cruciate ligament (PCL), and menisci^7,8^. However, these diagnostic modalities present significant limitations. MRI examinations are costly, time-consuming, and not universally accessible, especially in resource-constrained settings^9^. Arthroscopy, while providing direct visualization of intra-articular structures, is invasive and carries procedural risks^10^. Conventional radiography (X-ray), despite being widely available and cost-effective, has traditionally been considered inadequate for soft tissue evaluation, limiting its utility in comprehensive knee injury assessment^11,12^. Another major drawback was the subjective interpretation of images, which relied solely on the radiologist’s experience^13^. These constraints highlight the need for more accessible diagnostic approaches that can provide reliable information about both bony and soft tissue structures.
Artificial intelligence (AI), particularly deep learning, has emerged as a promising solution to enhance the diagnostic capabilities of conventional imaging modalities^14^.
AI-based detection and classification algorithms are essential in constructing fully automated pipelines^15^. The detection algorithm plays a crucial role in automatically identifying organ anatomy, replacing the need for manual organ anatomy segmentation^16^. On the other hand, the classification algorithm classifies abnormalities of the detected organ into specific categories^17^. Together, detection and classification algorithms provide a comprehensive approach to diagnosis, helping radiologists and clinicians make more informed decisions and ultimately improve patient care. Recent advances in computer vision have enabled AI systems to recognize complex patterns in medical images that may not be readily apparent to human observers^18,19^. Multi-label classification in medical imaging allows a model to identify multiple abnormalities within a single image. This is particularly useful for knee abnormalities, where multiple conditions often coexist. In the context of knee imaging, AI offers the potential to extract additional diagnostic information from X-rays beyond their conventional application^20^. Our study hypothesized that a multi-label deep learning approach could identify signatures of soft tissue injuries on plain radiographs, potentially transforming the diagnostic utility of X-rays in knee trauma. This research aimed to develop and validate an end-to-end AI-driven model capable of simultaneously detecting multiple knee abnormalities from X-ray images, including meniscus tears, ligament injuries, and bony abnormalities, thereby enhancing diagnostic efficiency while utilizing a more accessible imaging modality^21,22^.
The Ethics Committees of The First People’s Hospital of Kashi Prefecture and The Sixth Affiliated Hospital of Xinjiang Medical University approved this retrospective study. Since the study utilized de-identified data and was retrospective in nature, the requirement for individual agreements and written informed consent was waived.
We enrolled adult patients from two medical centers between January 2019 and December 2024. The inclusion criteria were as (1) patients were 18 years of age or older; and (2) they had undergone knee X-ray imaging for suspected or confirmed knee abnormalities. Subsequently, this preliminary cohort was rigorously refined through a set of exclusion criteria to ensure clinical relevance and data quality for our AI-driven analysis. Individuals were excluded (1) no specific knee abnormality was ultimately confirmed, encompassing conditions such as meniscus tears (MENI), ACL, PCL, medial collateral ligament injuries (MCL), lateral collateral ligament injuries (LCL), joint effusion (EFFU), bone marrow edema or contusion (CONT), or other soft tissue injuries (STI); (2) the presence of any suspected condition could not be rigorously established through a combination of detailed medical history (e.g., knee pain, trauma, or instability for ≥ 2 weeks), comprehensive physical examination (e.g., tenderness, swelling, positive stress tests), specific MRI findings (detailed in supplementary materials), or surgical verification where applicable (e.g., arthroscopy for MENI, ACL, PCL); (3) MRI images failed to meet predefined quality standards, such as unclear visualization of knee structures or the presence of significant motion artifacts; (4) patients presented with comorbidities that could confound imaging interpretation or algorithmic performance, including inflammatory arthropathies (rheumatoid, psoriatic, or gout), infectious processes (pyogenic arthritis or tuberculosis osteomyelitis), primary or metastatic bone tumors, metabolic bone diseases, or severe osteoporosis; or (5) the acquired standardized knee X-ray images did not meet the required quality standards. The First People’s Hospital of Kashi Prefecture (center 1) were randomly divided into training, validation, and internal testing datasets at a ratio of 1.5:1.5. The remaining patients from The First Affiliated Hospital of Xinjiang Medical University (center 2) served as the external testing sets. (See Fig. 1)
Fig. 1Enrolled population recruitment process.
We defined diagnostic criteria for each knee abnormality based on clinical history, physical exams, imaging, and, where applicable, arthroscopy. Not all patients underwent arthroscopy, as was reserved for cases requiring surgical intervention, such as persistent symptoms or complex injuries unresponsive to non-operative management^23–25^. MRI and arthroscopy have limitations; for example, MRI can miss subtle lesions like ramp tears, and arthroscopic accuracy depends on the operator. We accounted for this by cross-referencing multiple data sources. For injuries such as MENI and ligament tears or injuries (e.g., ACL, PCL, MCL and LCL), diagnosis relied on the integration of clinical history, standardized physical examination maneuvers, and imaging confirmation. Patients typically presented with specific injury mechanisms—twisting motions for meniscal damage, pivot injuries for ACL tears, dashboard trauma for PCL injuries, and valgus or varus stress for collateral ligament injuries. Physical examination employed validated clinical tests including McMurray and drawer tests for meniscal and cruciate ligament assessment, respectively, while stress testing evaluated collateral ligament integrity. MRI provided definitive structural assessment, demonstrating signal abnormalities, morphological changes, or complete discontinuity of affected structures. In contrast, EFFU and CONT represented pathophysiological findings rather than discrete injuries. Joint effusion was identified through clinical examination revealing swelling, warmth, and ballotable patella, with MRI confirming fluid accumulation within the joint space. Bone marrow edema presented as persistent deep pain following blunt trauma, with characteristic MRI signal patterns showing low T1 and high T2 intensities without requiring surgical correlation. STI encompassed a spectrum of muscle, tendon, and fascial damage, diagnosed through localized clinical findings and MRI demonstration of tissue disruption or inflammation. (See Table 1)
Table 1Physical examination findings for knee abnormality Assessment.ConditionAbnormalityPrimary physical examination findingsMENIMeniscus TearMedial/lateral joint line tenderness, pain on hyperflexion, positive McMurray test.ACLAnterior Cruciate Ligament TearPositive anterior drawer test, Lachman test instability, positive pivot-shift test.PCLPosterior Cruciate Ligament InjuryPositive posterior drawer test, posterior tibial sag.MCLMedial Collateral Ligament InjuryTenderness along MCL, medial joint swelling, pain/instability on valgus stress test (graded).LCLLateral Collateral Ligament InjuryTenderness over LCL, lateral joint swelling, pain/instability on varus stress test.EFFUJoint EffusionVisible/palpable joint swelling, warmth, ballotable patella (fluid wave test).CONTBone Marrow Edema/ContusionDiffuse bony tenderness (e.g., femoral condyles, tibial plateau), often without specific ligamentous signs.STISoft Tissue InjuryLocalized tenderness, bruising, swelling, or reduced strength specific to the affected muscle/tendon.
A standardized knee X-ray protocol was established to ensure consistent and reliable assessment of joint abnormalities. All participants underwent weight-bearing anteroposterior (AP) and lateral radiographs following international guidelines. The AP radiographs were performed with participants standing upright, feet slightly apart, and patellae facing forward. The X-ray beam was centered over the joint space, with the tube voltage set to 65 kVp and an exposure of 10 mAs. The source-to-image distance was maintained at 115 cm. These images provided clear visualization of joint space narrowing, osteophytes, and alignment, which are critical for diagnosing conditions like osteoarthritis. The lateral views were acquired with participants standing with the affected knee closest to the detector, flexed at approximately 20–30 degrees. The X-ray beam was directed from medial to lateral, with a 10-degree caudal tube angulation to enhance visualization of the joint compartments and patellofemoral articulation. To ensure consistency across centers, the same imaging system, a United Imaging uDR 760i digital radiography unit, was used. For this view, the tube voltage was set to 60 kVp with an exposure of 8 mAs, and the source-to-image distance was maintained at 115 cm. This protocol ensured high-quality, reproducible images suitable for AI-based analysis. For the purpose of this AI-driven study, only the AP view images were used for model training and evaluation.
All knee MRI examinations (Siemens Magnetom Skyra 3T) were conducted on clinical MRI scanners using dedicated knee coils. A comprehensive MRI protocol was employed to provide detailed assessment of the knee joint. The standardized protocol commenced with a tri-planar localizer sequence to ensure accurate slice positioning. This was followed by a T1-weighted Turbo Spin Echo (TSE) sequence in the sagittal plane, acquired with a Repetition Time (TR) of 550 ms, an Echo Time (TE) of 11.0 ms, and a 150-degree flip angle. This T1 sequence used a 3.0 mm slice thickness, a 160 mm Field of View (FoV), and achieved a spatial resolution with a voxel size of 0.3 × 0.3 × 3.0 mm³. Crucial for detecting fluid-related pathologies and evaluating soft tissue structures, a series of Proton Density (PD)-weighted TSE sequences with fat saturation were subsequently acquired. These PD TSE FatSat sequences were performed in the sagittal, coronal, and transverse planes, sharing common parameters including a TR of 3000 ms, a TE of 30.0 ms, a 150-degree flip angle, a 3.0 mm slice thickness, and a 160 mm FoV, yielding a voxel size of 0.5 × 0.5 × 3.0 mm³. To complement the PD sequences and further enhance the visualization of edema and other fluid collections, a T2-weighted TSE sequence with fat saturation was also obtained in the transverse plane, characterized by a longer TR of 4000 ms and a longer TE of 90.0 ms, while maintaining the same 3.0 mm slice thickness, 160 mm FoV, and 0.5 × 0.5 × 3.0 mm³ voxel size.
Knee X-ray images were acquired from two different scanners at medical centers in DICOM format, which is a standard that encapsulates both the image data and its associated metadata. To facilitate analysis, we converted these files into PNG format, preserving the original pixel information while discarding extraneous metadata. Each image underwent normalization by scaling its pixel intensities to the range [0, 1], determined by the minimum and maximum values within the image. This step minimized variations stemming from differences in X-ray exposure and imaging equipment across the centers.
To define the focus of the subsequent analysis, two experienced orthopedic surgeons independently annotated the region of interest (ROI)—the knee region—in each image using the free open-source tool LabelMe(https://github.com/wkentaro/labelme). Both surgeons, specializing in musculoskeletal conditions, performed the annotations blindly, without access to clinical details or knowledge of each other’s markings, thereby minimizing potential bias in delineating the knee area. In cases where discrepancies arose regarding the boundaries of the knee region, the surgeons engaged in discussions to resolve differences and reach a consensus. This process ensured consistency and accuracy in the annotations, establishing a reliable foundation for the subsequent stages of the study. (See Fig. 2)
Fig. 2Overview of Wole Study.
To enable consistent input for downstream multi-label classification tasks we developed of a knee detection model using YOLOv11 (https://github.com/ultralytics/ultralytics), which aimed to automate the localization of anatomical knee radiographers^26^. The model was initialized with pretrained weights from the YOLOv11n architecture, which had been trained on the COCO dataset, and subsequently fine-tuned annotated knee X-rays. This preprocessing step played a critical role in ensuring consistent input for downstream multi-label classification tasks, thereby enhancing the reliability and generalizability of the overall pipeline. (See Supplementary Appendix 1)
For each knee radiograph, a ROI encompassing the critical joint structures was extracted based on rectangular coordinates provided by medical experts. To prepare these ROIs for model input, we implemented a standardized preprocessing pipeline using a custom Python script with the OpenCV library. For each image, the script first cropped the rectangular ROI as defined by the corresponding expert annotation. To handle the variable dimensions of these ROIs while preventing anatomical distortion, we employed a padding technique to preserve the original aspect ratio. Each cropped rectangle was first scaled such that its longest side was resized to 224 pixels. Subsequently, zero-padding (black pixels) was added to the shorter side until the final image dimensions reached the required 224 × 224 square input size. This process ensured that all input images were of a uniform size without squashing or stretching the underlying anatomical features. The final, preprocessed images were then saved as PNG files for subsequent model training and evaluation.
In this study, we ensured the integrity of model evaluation by strictly preserving the validation and inner testing datasets in their original, unaltered forms to prevent data leakage and facilitate an unbiased assessment of performance on unseen data. To enhance model robustness and generalizability, data augmentation was applied exclusively to the training datasets, a practice widely recognized for improving model performance. We implemented a custom augmentation pipeline that included rotations at 30-degree increments (ranging from 0 to 330 degrees), horizontal and vertical flips, and combinations of these transformations applied to extracted ROI images. This approach generated a diverse set of augmented images, enabling the model to generalize effectively across various knee X-ray orientations and conditions while maintaining the authenticity of the validation and testing processes. By enriching the training data without compromising the integrity of the evaluation datasets, we ensured a rigorous and objective assessment of the model’s performance. (See Supplementary Appendix 2)
We developed a multi-label deep learning framework designed to detect eight distinct knee abnormalities—MENI, ACL, PCL, MCL, LCL, EFFU, CONT, and STI—from anterior-posterior knee radiographs. To achieve this, we adapted six convolutional neural network architectures—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—each pretrained on ImageNet, by modifying their final classification layers to support multi-label outputs. Spatial augmentations, such as random rotations of ± 15°, horizontal flips, and zoom variations of ± 10%, were applied to enrich the training dataset and improve model robustness. The models were trained over 30 epochs using the Adam optimizer with an initial learning rate of 10e-4 and employed the BCEWithLogitsLoss function to handle multi-label classification. A learning rate scheduler dynamically adjusted the rate based on validation AUC plateaus to optimize performance. Model selection was guided by the highest mean AUC achieved on the validation set, ensuring the most effective architecture was chosen for the task. (See Supplementary Appendix 3)
In this study, we evaluated the performance of various models using seven widely recognized metrics for multilabel classification. These metrics included Mean Average Precision (mAP), Overall Precision (OP), Overall Recall (OR), Overall F1-measure (OF1), Per-class Precision (CP), Per-class Recall (CR), and Per-class F1-measure (CF1). To further compare the overall classification performance of each model, we utilized receiver operating characteristic (ROC) curves and calculated the area under the curve (AUC). The detailed formulations and explanations of each metric were provided to ensure a comprehensive understanding of their application and significance in the context of this research. Below are the detailed formulations and explanations of each
The Average Precision (AP) for the (i)-th category is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ :A{P}{i}=\frac{1}{\left|{G}{i}\right|}{\sum:}{k=1}^{n}{P}{k}\times:re{l}_{k}
(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:\left|{G}_{i}\right| $$\end{document}) represents the number of samples in the (i)-th category. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:{P}_{k} $$\end{document}) represents the precision at the (k)-th prediction. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:{\text{rel}}_{k} $$\end{document}) is an indicator function that equals 1 if the (k)-th prediction is a true positive, and 0 otherwise. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:n $$\end{document}) represents the total number of predictions. #### Mean average precision (mAP) The Mean Average Precision (mAP) across all categories is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:mAP=\frac{1}{\left|\mathcal{Y}\right|}{\sum\:}_{i=1}^{\left|\mathcal{Y}\right|}A{P}_{i} $$\end{document} where: (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:\left|\mathcal{Y}\right| $$\end{document}) represents the total number of categories.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:A{P}_{i} $$\end{document}) represents the average precision for the (i)-th category. #### Per-class precision (CP) The Per-class Precision (CP) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:CP=\frac{1}{K}{\sum\:}_{i=1}^{K}\frac{T{P}_{i}}{T{P}_{i}+F{P}_{i}} $$\end{document} where: (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:K $$\end{document}) represents the total number of categories.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:T{P}_{i} $$\end{document}) represents the number of true positives for the (i)-th category.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:F{P}_{i} $$\end{document}) represents the number of false positives for the (i)-th category. #### Per-class recall (CR) The Per-class Recall (CR) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:CR=\frac{1}{K}{\sum\:}_{i=1}^{K}\frac{T{P}_{i}}{T{P}_{i}+F{N}_{i}} $$\end{document} where: (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:K $$\end{document}) represents the total number of categories.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:T{P}_{i} $$\end{document}) represents the number of true positives for the (i)-th category.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:F{N}_{i} $$\end{document}) represents the number of false negatives for the (i)-th category. #### Per-class F1-measure (CF1) The Per-class F1-measure (CF1) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:CF1=\frac{1}{K}{\sum\:}_{i=1}^{K}\frac{2\times\:T{P}_{i}}{2\times\:T{P}_{i}+F{P}_{i}+F{N}_{i}} $$\end{document} where: (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:K $$\end{document}) represents the total number of categories.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:T{P}_{i} $$\end{document}) represents the number of true positives for the (i)-th category.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:F{P}_{i} $$\end{document}) represents the number of false positives for the (i)-th category.(\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:F{N}_{i} $$\end{document}) represents the number of false negatives for the (i)-th category. #### Overall precision (OP) The Overall Precision (OP) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:OP=\frac{{\sum\:}_{i=1}^{K}T{P}_{i}}{{\sum\:}_{i=1}^{K}\left(T{P}_{i}+F{P}_{i}\right)} $$\end{document} where: (TP) represents the total number of true positives across all categories.(FP) represents the total number of false positives across all categories. #### Overall recall (OR) The Overall Recall (OR) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:OR=\frac{{\sum\:}_{i=1}^{K}T{P}_{i}}{{\sum\:}_{i=1}^{K}\left(T{P}_{i}+F{N}_{i}\right)} $$\end{document} where: (TP) represents the total number of true positives across all categories.(FN) represents the total number of false negatives across all categories. #### Overall F1-measure (OF1) The Overall F1-measure (OF1) is calculated \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \:OF1=\frac{2\times\:OP\times\:OR}{OP+OR} $$\end{document} where: (TP) represents the total number of true positives across all categories.(FP) represents the total number of false positives across all categories.(FN) represents the total number of false negatives across all categories. ### Application development We developed two software tools to advance our medical imaging research. First, we created a desktop application for the Windows operating system using the PyQt framework. This application provided a robust platform for processing and analyzing imaging data, leveraging PyQt’s capabilities to build an intuitive user interface. Second, we constructed a web-based application using the Flask framework, which integrated both frontend and backend components^27^. ### Model explanation We applied multiple Class Activation Mapping techniques—including GradCAM, HiResCAM, ScoreCAM, and others—targeting the final convolutional layers to highlight key image regions for each positive class^28^. These heatmaps were generated from preprocessed images, resized to 224 × 224 pixels and normalized, then processed on GPU or CPU as available. Visualizations were saved individually and later combined into a single grid, integrating the original image with CAM outputs for a clear, cohesive summary of the model’s focus across conditions, tailored for medical researchers. ### Statistical analysis The normality of the data was assessed using the Shapiro-Wilk test, and continuous variables were expressed as mean ± standard deviation (SD) for normally distributed data or as median with interquartile range (IQR) for non-normally distributed data. Statistical analyses were performed using R Version 4.4.1, while machine learning (ML) models were developed and analyzed using Python 3.10.5 with the Scikit-learn library. Model performance was evaluated using sensitivity, specificity, F1-score, and area under the curve (AUC) metrics to ensure a comprehensive assessment of predictive accuracy and reliability. These methodological steps ensured that the analyses adhered to rigorous standards, enabling robust interpretation of the results and alignment with the study’s objectives. ## Results ### Enrolled patients characteristics The study enrolled a total of 4,215 patients, comprising 3,088 from Center 1 and 1,127 from Center 2. The mean age of participants was 42.6 years (standard deviation [SD] = 9.4) in Center 1 and 43.2 years (SD = 6.3) in Center 2. Females predominated in both cohorts, accounting for 60.1% (*n* = 1,858) in Center 1 and 61.6% (*n* = 694) in Center 2, while males represented 39.9% (*n* = 1,230) and 38.4% (*n* = 433), respectively. Condition prevalence varied significantly between the centers, as detailed in Supplementary Table 1. In Center 1, the most common conditions included EFFU (*n* = 2,114), MENI (*n* = 2,832), and LCL (*n* = 2,356), while Center 2 reported lower numbers, with EFFU (*n* = 976), MENI (*n* = 869), and LCL (*n* = 705) being the most frequent. These findings provided a detailed baseline characterization of the enrolled population, highlighting both demographic distributions and condition prevalence at the study’s initiation. (See Supplementary Table 2) ### Knee localization performance The localization model YOLO, designed for knee detection, underwent evaluation as depicted in Fig. 3A–D, offering a detailed assessment of its performance across multiple metrics. Figure 3A displayed the Recall-Confidence Curve, where the recall rate began at 1.00 at a confidence threshold of 0.000, reflecting a perfect recall at minimal confidence levels. As the confidence threshold increased, the recall rate declined sharply, illustrating the inherent trade-off between confidence and recall. Figure 3B presented the F1-Confidence Curve, with the F1 score reaching its peak of 1.00 at a confidence threshold of 0.652. This peak demonstrated an optimal balance between precision and recall at this specific threshold. Figure 3C provided an additional perspective on the Recall-Confidence relationship, consistent with the findings in Fig. 3A, further emphasizing the steep decline in recall as confidence thresholds rose. Figure 3D illustrated the Precision-Recall Curve, where precision remained consistently high at 0.995 mAP@0.5, underscoring robust performance across varying recall levels. Collectively, these results highlighted the model’s capacity to achieve high precision and recall within specific confidence thresholds, offering valuable insights into its reliability and stability. The shapes of the curves indicated that the model performed optimally within a defined confidence range, beyond which performance metrics deteriorated rapidly. This analysis emphasized the critical importance of selecting appropriate confidence thresholds to ensure a balance between diagnostic accuracy and reliability in clinical applications. By addressing these considerations, the study aimed to refine the practical utility of the model, ensuring its suitability for real-world medical imaging tasks. Fig. 3Localization YOLO model performance. (**A**) The Recall-Confidence Curve shows the recall rate starting at 1.00 at a confidence threshold of 0.000 and decreasing sharply as confidence increases. (**B**) The F1-Confidence Curve reaches 1.00 at a confidence threshold of 0.652, indicating an optimal balance between precision and recall. (**C**) Another perspective on the Recall-Confidence relationship, consistent with Fig. 3A, emphasizing the steep decline in recall as confidence increases. (**D**) The Precision-Recall Curve shows precision remaining high at 0.995 mAP@0.5, indicating robust performance across various recall levels. ### Multi-label classification performance In this study, we assessed five deep learning models—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—for their ability to classify eight common knee MENI, ACL), MENI, ACL, PC, MC, LC, EFFU, CONT and STI. We measured performance using the AUC metric across three validation, internal testing, and external testing. Figure 4A–C, provide a comprehensive evaluation of various models’ performance across different sets. Figure 4A displays the ROC curves for the validation set, comparing models such as ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19. The AUC values for these models are 0.758, 0.722, 0.646, 0.722, and 0.632, respectively. Figure 4B shows the ROC curves for the internal test set, maintaining a similar trend in model performance. Figure 4C presents the ROC curves for the external test set, further confirming the consistent performance of ResNet152 across different datasets. ResNet152 consistently outperformed the others, achieving the highest AUCs for most conditions in all phases. During validation, it recorded AUCs of 0.803 for MENI and 0.801 for STI, while DenseNet121 reached 0.748 for EFFU, and ShuffleNetV2 scored 0.781 for STI. MobileNetV3 and VGG19 trailed, with VGG19 posting notably low AUCs of 0.560 for ACL and 0.579 for PCL. In the internal testing phase, ResNet152’s AUCs ranged from 0.750 for ACL to 0.861 for MENI, though DenseNet121 surpassed it for EFFU with an AUC of 0.886 against ResNet152’s 0.813. ShuffleNetV2 also performed strongly, achieving 0.834 for MENI and 0.833 for EFFU, while MobileNetV3 and VGG19 lagged, particularly for PCL and CONT. External testing reinforced ResNet152’s lead, with AUCs from 0.764 for CONT to 0.863 for EFFU, followed closely by ShuffleNetV2 at 0.844 for EFFU. DenseNet121 held steady but was outpaced by ResNet152 in most categories, while MobileNetV3 and VGG19 struggled, especially with MENI and ACL. (**See** Table 2) Fig. 4Model performance across different sets. (**A**) ROC curves for the validation set, comparing ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19. (**B**) ROC curves for the internal test set, showing consistent performance trends. (**C**) OC curves for the internal test set, showing consistent performance trends. Table 2AUC value of each model.ModelMENIACLPCLMCLLCLEFFUCONTSTIValidationResNet1520.803 (0.785, 0.822)0.748 (0.732, 0.765)0.752 (0.74, 0.763)0.778 (0.765, 0.789)0.74 (0.726, 0.755)0.643 (0.571, 0.711)0.799 (0.786, 0.811)0.801 (0.789, 0.813)DenseNet1210.746 (0.727, 0.764)0.68 (0.664, 0.698)0.746 (0.733, 0.758)0.722 (0.709, 0.735)0.677 (0.663, 0.691)0.748 (0.7, 0.794)0.701 (0.686, 0.717)0.758 (0.745, 0.772)MobileNetV30.747 (0.727, 0.766)0.587 (0.568, 0.606)0.643 (0.63, 0.656)0.63 (0.616, 0.644)0.637 (0.622, 0.652)0.598 (0.536, 0.66)0.661 (0.645, 0.677)0.666 (0.652, 0.682)ShuffleNetV20.748 (0.729, 0.77)0.734 (0.717, 0.75)0.74 (0.727, 0.753)0.73 (0.717, 0.744)0.712 (0.697, 0.728)0.638 (0.573, 0.701)0.693 (0.678, 0.709)0.781 (0.768, 0.795)VGG190.691 (0.671, 0.712)0.56 (0.543, 0.579)0.579 (0.565, 0.594)0.633 (0.618, 0.647)0.554 (0.538, 0.57)0.661 (0.614, 0.704)0.677 (0.662, 0.693)0.701 (0.686, 0.717)Internal testingResNet1520.861 (0.842, 0.879)0.75 (0.732, 0.766)0.791 (0.78, 0.802)0.818 (0.806, 0.829)0.822 (0.81, 0.834)0.813 (0.768, 0.863)0.811 (0.8, 0.823)0.803 (0.79, 0.815)DenseNet1210.839 (0.818, 0.857)0.71 (0.692, 0.727)0.727 (0.714, 0.738)0.755 (0.743, 0.768)0.701 (0.686, 0.716)0.886 (0.859, 0.909)0.726 (0.713, 0.74)0.713 (0.699, 0.726)MobileNetV30.743 (0.721, 0.765)0.611 (0.594, 0.628)0.631 (0.616, 0.644)0.683 (0.669, 0.695)0.705 (0.691, 0.718)0.737 (0.679, 0.792)0.644 (0.629, 0.659)0.669 (0.655, 0.683)ShuffleNetV20.834 (0.816, 0.852)0.715 (0.699, 0.731)0.711 (0.699, 0.724)0.776 (0.764, 0.787)0.773 (0.76, 0.786)0.833 (0.781, 0.877)0.75 (0.737, 0.764)0.739 (0.724, 0.754)VGG190.708 (0.687, 0.729)0.546 (0.529, 0.563)0.543 (0.529, 0.557)0.558 (0.544, 0.571)0.633 (0.619, 0.647)0.66 (0.621, 0.699)0.571 (0.555, 0.587)0.584 (0.569, 0.6)External TestingResNet1520.831 (0.792, 0.875)0.783 (0.745, 0.819)0.78 (0.753, 0.81)0.77 (0.738, 0.802)0.799 (0.768, 0.831)0.863 (0.774, 0.936)0.764 (0.73, 0.8)0.786 (0.754, 0.816)DenseNet1210.745 (0.697, 0.793)0.707 (0.662, 0.749)0.729 (0.699, 0.76)0.731 (0.697, 0.765)0.693 (0.658, 0.727)0.836 (0.73, 0.928)0.669 (0.629, 0.706)0.743 (0.709, 0.777)MobileNetV30.683 (0.63, 0.732)0.605 (0.559, 0.651)0.636 (0.605, 0.672)0.628 (0.593, 0.662)0.62 (0.583, 0.657)0.801 (0.727, 0.866)0.631 (0.592, 0.668)0.657 (0.618, 0.693)ShuffleNetV20.764 (0.717, 0.814)0.734 (0.69, 0.776)0.727 (0.695, 0.759)0.747 (0.717, 0.779)0.715 (0.68, 0.751)0.844 (0.741, 0.925)0.735 (0.698, 0.77)0.763 (0.729, 0.794)VGG190.671 (0.622, 0.717)0.559 (0.517, 0.604)0.628 (0.594, 0.661)0.603 (0.57, 0.637)0.576 (0.539, 0.612)0.706 (0.583, 0.817)0.631 (0.59, 0.67)0.656 (0.62, 0.693)MENI: Meniscus Tears; ACL: Anterior Cruciate Ligament Tears; PCL: Posterior Cruciate Ligament Injuries; MCL: Medial Collateral Ligament Injuries; LCL: Lateral Collateral Ligament Injuries; EFFU: Joint Effusion; CONT: Bone Marrow Edema or Contusion; STI: Soft Tissue Injuries. We investigated the performance of five deep learning models—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—in classifying multiple knee conditions simultaneously from medical images. (See Table 3) ResNet152 demonstrated consistent superiority, achieving the highest mAP in all 87.4% during validation, 90.1% in internal testing, and 88.5% in external testing. It also achieved the highest F1-scores, with CF1 and OF1 values of 85.8% in validation, 86.4% in internal testing, and 85.8% in external testing. DenseNet121 and ShuffleNetV2 exhibited notable strengths, particularly in recall metrics during validation. DenseNet121 achieved CR and OR values of 91.7%, while ShuffleNetV2 recorded 91.6% for both, surpassing ResNet152’s 88.8%. However, ResNet152 maintained higher precision, with CP and OP values of 83.200% in validation, compared to 80.1% for DenseNet121 and 80.700% for ShuffleNetV2. During internal testing, ResNet152 continued to lead across all metrics, though DenseNet121 and ShuffleNetV2 remained competitive, with CR values of 91.1% and 90.2%, respectively, compared to ResNet152’s 88.1%. External testing further confirmed ResNet152’s dominance, as it achieved a mAP of 88.5% and CF1/OF1 values of 85.8%. Figure 5A–I, provide a comprehensive evaluation of various models’ performance across different sets. Figure 5A–C display the ROC curves, calibration curves, and precision-recall curves for the validation set. Figure 5A shows the ROC curves for models such as MENI, ACL, PCL, MCL, LCL, EFFU, CONT, and STI, with AUC values ranging from 0.724 to 0.801. Figure 5B presents the calibration curves, showing the alignment between predicted probabilities and actual outcomes. Figure 5C illustrates the precision-recall curves, highlighting the trade-off between precision and recall. Figure 5D–F show the same metrics for the internal test set, with similar trends in model performance. Figure 5G–I further validate the models’ performance on the external test set, confirming the robustness of the best-performing models. Table 3Multi-label classifier performance.NetworkmAP (%)CP (%)CR (%)CF1 (%)OP (%)OR (%)OF1 (%)Validation SetResNet15287.48283.20088.80085.80083.20088.80085.800DenseNet12185.28780.10091.70085.40080.10091.70085.400MobileNetV381.75677.90089.30083.00077.90089.30083.000ShuffleNetV285.52180.70091.60085.70080.70091.60085.700VGG1980.14775.30087.90078.70075.30087.90078.700Internal Testing SetResNet15290.11184.90088.00086.40084.90088.00086.400DenseNet12186.53979.30091.10084.70079.30091.10084.700MobileNetV382.22877.30088.80082.30077.30088.80082.300ShuffleNetV287.29280.20090.20084.80080.20090.20084.800VGG1978.43174.80087.90077.80074.80087.90077.800External Testing SetResNet15288.52383.10088.70085.80083.10088.70085.800DenseNet12185.43578.30091.00084.00078.30091.00084.000MobileNetV380.66076.30088.80081.70076.30088.80081.700ShuffleNetV286.08680.20088.90084.20080.20088.90084.200VGG1979.66476.70087.00077.10076.70087.00077.100mAP: Mean Average Precision, CP: Class Precision, CR: Class Recall, CF1: Class F1-Score, OP: Overall Precision, OR: Overall Recall, OF1: Overall F1-Score. Fig. 5Model performance across different sets. (**A**: ROC curves for the validation set, comparing models such as MENI, ACL, PCL, MCL, LCL, EFFU, CONT, and STI. (**B**) Calibration curves for the validation set, showing the alignment between predicted probabilities and actual outcomes. (**C**) Precision-recall curves for the validation set, highlighting the trade-off between precision and recall. (**D**) ROC curves for the internal test set, showing similar trends in model performance. (**E**) Calibration curves for the internal test set, confirming the robustness of the best-performing models. (**F**) Precision-recall curves for the internal test set, further validating the models’ performance. (**G**) ROC curves for the external test set, confirming the robustness of the best-performing models. (**H**) Calibration curves for the external test set, showing consistent performance across different datasets. (**I**) Precision-recall curves for the external test set, highlighting the models’ reliability. We evaluated the performance of five deep learning models—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—in an End-to-End workflow designed to classify knee conditions directly from ROIs extracted by a YOLO model on an external testing set. Figure 6A–B provides a comprehensive assessment of the models’ performance across different datasets. Figure 6A illustrated the ROC curves for ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19, with AUC values of 0.765, 0.706, 0.646, 0.734, and 0.629, respectively. These curves highlight the trade-off between sensitivity and specificity, with ResNet152 demonstrating the best performance. Figure 6B presents ROC curves for specific knee conditions, including MENI, ACL, PCL, MCL, LCL, EFFU, CONT, and STI, with AUC values ranging from 0.727 to 0.873. MENI consistently outperformed other models in terms of AUC, indicating superior diagnostic accuracy and robustness. Among the models, ResNet152 achieved the highest mAP of 86.1%, along with strong CF1 and OF1 scores of 84.0%. DenseNet121 followed with a mAP of 83.5% but excelled in recall, achieving CR and OR values of 91.4%, surpassing ResNet152’s 87.9%. ShuffleNetV2 also performed robustly, with a mAP of 85.0% and CF1/OF1 scores of 83.8%, while MobileNetV3 and VGG19 lagged, achieving mAPs of 80.2% and 79.7%, respectively. (**See** Table 4) Fig. 6End-to-end performance. (**A**) ROC curves for models such as ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19, showing their respective AUC values. (**B**) ROC curves for models such as MENI, ACL, PCL, MCL, LCL, EFFU, CONT, and STI, showing their respective AUC values. The figures collectively highlight the superior performance of ResNet152 and MENI across different sets and datasets. Table 4End-to-end workflow performance.NetworkmAP (%)CP (%)CR (%)CF1 (%)OP (%)OR (%)OF1 (%)ResNet15286.180.587.984.080.587.984.0DenseNet12183.577.591.483.777.591.483.7MobileNetV380.275.889.081.675.889.081.6ShuffleNetV285.079.289.283.879.289.283.8VGG1979.776.387.177.276.387.177.2 ### End-to-end application development In this study, we extracted and fine-tuned the ResNet152 model, trained a YOLO model, and developed an end-to-end application, as illustrated in Fig. 7. The application featured a web-based interface capable of directly processing DICOM or image format files, as shown in Fig. 7A ([http://43.143.217.126:9070/knee/](http://43.143.217.126:9070/knee/)), and a Windows-based interface with similar functionality, depicted in Fig. 7B. The integration of these models into a user-friendly application enabled seamless handling of medical imaging data, streamlining the workflow for clinicians and researchers. Fig. 7Application interface. (**A**) Web Application interface that could directly handle DICOM or Image format file. ([http://43.143.217.126:9070/knee/](http://43.143.217.126:9070/knee/)) (**B**) Windows Application interface that could directly handle DICOM or Image format file. Figure 8–F, provide a comprehensive visual documentation of a 43-year-old male sports teacher’s right knee injury sustained during a football game. Figure 8A and B show arthroscopic images of the knee, revealing a tear of the anterior cruciate ligament (ACL) and meniscus. Figure 8C and D, and Fig. 8E present MRI scans in different a T2-weighted fat-suppressed coronal view, a sagittal view, and an axial T2-weighted image. These MRI images demonstrate mild bone marrow edema in the distal femur and proximal tibia, a tear of the ACL, significant joint effusion, and surrounding soft tissue edema. Figure 8F shows the results from the window application (DeepKneeXR) multi-lesion X-ray screening system, which accurately identified abnormalities in the meniscus (MENI), effusion (EFFU), continuity (CONT), subchondral bone (STI), and ACL. The DeepKneeXR system’s findings are consistent with both the arthroscopic and MRI assessments, highlighting its effectiveness in detecting knee injuries. Fig. 8Demonstration of windows application. (**A**) Arthroscopic image showing a tear of the ACL. (**B**) Arthroscopic image showing a tear of the meniscus. (**C**) T2-weighted fat-suppressed coronal MRI scan showing bone marrow edema and ACL tear. (**D**) Sagittal MRI scan showing joint effusion and soft tissue edema. (**E**) Axial T2-weighted MRI scan showing detailed structures of the injured knee. (**F**) DeepKneeXR X-ray screening system interface displaying localized abnormalities with high accuracy, matching the findings from arthroscopic and MRI assessments. Figure 9A–F, provide a comprehensive visual documentation of a 23-year-old male college student’s right knee injury sustained during a football game. Figure 9A and B show arthroscopic images of the knee, revealing a tear of the anterior cruciate ligament (ACL) and lateral meniscus. Figure 9C and D, and Fig. 9E present MRI scans in different a coronal fat-suppressed view, a sagittal view, and an axial view. These MRI images demonstrate relatively small bone marrow edema in the distal femur, abnormal high signal intensity in the ACL, significant joint effusion, and edema in the surrounding lateral soft tissue of the knee joint. Figure 9F shows the results from a developed web application, which accurately identified abnormalities in the meniscus (MENI), effusion (EFFU), subchondral bone (STI), and ACL. The web application’s findings are consistent with both the arthroscopic and MRI assessments, highlighting its effectiveness in detecting knee injuries. Fig. 9Demonstration of web application. (**A**) Arthroscopic image showing a tear of the ACL. (**B**) Arthroscopic image showing a tear of the lateral meniscus. (**C**) Coronal fat-suppressed MRI scan showing bone marrow edema and ACL tear. (**D**) Sagittal MRI scan showing joint effusion and soft tissue edema. (**E**) Axial MRI scan showing detailed structures of the injured knee. (**F**) Web application interface displaying localized abnormalities with high accuracy, matching the findings from arthroscopic and MRI assessments. ### Model explanation Figure 10 demonstrated the performance of various explainable artificial intelligence (XAI) techniques, emphasizing their capacity to deliver interpretable and clinically actionable insights into complex injuries. The figure provided a comprehensive visual analysis of a 26-year-old male’s left knee injury, which exhibited multiple concurrent pathologies, including meniscal injury (MENI), anterior cruciate ligament (ACL) tear, posterior cruciate ligament (PCL) tear, medial collateral ligament (MCL) injury, lateral collateral ligament (LCL) injury, effusion (EFFU), contusion (CONT), and soft tissue injury (STI). Fig. 10Comparison of explainable AI methods. Visual analysis of a 26-year-old male’s left knee injury, which exhibited multiple concurrent pathologies, including meniscal injury (MENI), anterior cruciate ligament (ACL) tear, posterior cruciate ligament (PCL) tear, medial collateral ligament (MCL) injury, lateral collateral ligament (LCL) injury, effusion (EFFU), contusion (CONT), and soft tissue injury (STI). ## Discussion Knee injuries are among the most prevalent musculoskeletal conditions encountered in clinical practice, significantly impacting patient mobility and quality of life^29^. Accurate and timely diagnosis is essential, as it can prevent severe outcomes, including post traumatic osteoarthritis (PTOA), which is the most serious consequences of knee injury^30^. For ACL injuries, the pooled prevalence of PTOA is approximately 37.9% for operatively treated cases and 40.5% for non-operatively treated cases over a median follow-up of 14.6 to 15 years and mainly impacts young adults in the United States^31^. Knee abnormalities are common and can result from trauma, degenerative diseases, or congenital conditions. Meniscus and ligament injuries, particularly to the anterior cruciate ligament (ACL), posterior cruciate ligament (PCL), and medial collateral ligament (MCL), are frequent concerns. ACL injuries are the most widespread, with an estimated 100,000 to 200,000 cases annually in the U.S^32^. Meniscus injuries often occur with ACL tears. In a study of ACL patients, 77% also had meniscus injuries, with medial meniscus tears being more common. These injuries are associated with early onset post-traumatic osteoarthritis, particularly when combined with ACL damage^33^. MCL injuries are also common in conjunction with ACL injuries. In one study, 22% of patients with ACL ruptures had MCL injuries. MCL injuries often occur with bone contusions and are associated with more severe knee trauma^34^. While not as frequently mentioned in the provided data, PCL injuries are less common than ACL injuries but can occur in high-impact trauma situations^35,36^. In current clinical settings, MRI and knee arthroscopy are commonly employed for the diagnosis of knee injuries. However, MRI is often prohibitively expensive, leading to extended wait times for patients^37^, while arthroscopy, though informative, is an invasive procedure that carries inherent risks and is not universally accessible, particularly in resource-limited areas^38^. Furthermore, these diagnosing modalities are not always necessary for every case, which underscores the need for a more efficient and less traumatic initial screening tool. In daily clinical practice, X-ray examination is a widely available, cost-effective and initial screening imaging modality that can be utilized in various healthcare settings^39^. The research aims to provide initial clinical decision support by generating automated reports that highlight detected abnormalities and suggest potential diagnoses, thereby improving clinical workflow efficiency and aiding radiologists in their assessments. The study will evaluate the model’s performance in terms of accuracy, sensitivity, specificity, and overall clinical utility, comparing its diagnostic capabilities to gold standards to establish its effectiveness as a supplementary tool in clinical practice. Finally, the research will explore the implications of these findings for future research directions in AI-driven medical imaging and the integration of such technologies into routine clinical practice, ultimately seeking to enhance patient outcomes and healthcare efficiency. Deep learning models have been used to detect musculoskeletal disorders like fractures, cartilage lesions, and osteoarthritis, showcasing the power of multi-label classification to improve diagnostic accuracy and predict prognosis in complex cases^36,40,41^. Single-label classification models struggle to capture the complexity of medical conditions, as they identify only one condition per image. This limitation is particularly evident in knee imaging, where multiple abnormalities often coexist. For instance, a study on knee MRI showed that deep learning models improved inter-reader agreement by identifying and grading multiple lesions in cartilage, bone marrow, and ligaments, issues that single-label models might miss^42^. The inability of single-label models to recognize multiple conditions simultaneously can lead to incomplete diagnoses and suboptimal treatment plans^36^. Recognizing co-existing abnormalities in knee imaging is vital for accurate diagnosis and treatment planning. Multi-label classification enables a comprehensive assessment, identifying conditions like ligament, meniscus, and bone injuries in a single imaging session^35,43^. The integration of AI-driven multi-label classification in knee imaging can support radiologists by reducing the likelihood of oversight and enhancing diagnostic precision^13^. In this study we presented a novel approach by developing a fully automated deep learning pipeline based on X-ray images, designed for the early screening of knee abnormalities. By harnessing advanced artificial intelligence techniques, this tool aims to enhance diagnostic accuracy, facilitate prompt treatment decisions, and ultimately improve patient outcomes in knee injury management. We evaluated five models—ResNet152, DenseNet121, MobileNetV3, ShuffleNetV2, and VGG19—against eight common knee conditions, assessing their performance across validation, internal testing, and external testing phases using metrics like AUC, mAP, and F1-scores. ResNet152 consistently led the pack, achieving AUCs as high as 0.863 for effusion and mAP scores up to 90.1% in internal testing, showcasing its reliability in pinpointing knee abnormalities. DenseNet121 and ShuffleNetV2 also performed well, particularly in recall, with scores exceeding 91% during validation, hinting at their value for detecting specific injuries. In contrast, MobileNetV3 and VGG19 struggled, especially with conditions like ACL and meniscus tears, where AUCs dipped below 0.65. Beyond model comparison, we integrated ResNet152 into a practical end-to-end application, available as both web-based and Windows-based tools, which processed imaging data seamlessly and aligned with arthroscopic and MRI findings in real-world cases, such as a sports teacher’s ACL tear and a student’s meniscus injury. The addition of explainable AI techniques further enriched our approach, shedding light on complex injuries like concurrent ligaments and soft tissue damage in a young patient’s knee. These findings underscored the power of advanced imaging tools to support clinical decision-making. The integration of these models into an end-to-end application enabled seamless DICOM/image processing via web and desktop interfaces, bridging AI research with clinical workflows. Case examples corroborated the system’s diagnostic accuracy, with DeepKneeXR findings aligning closely with MRI and arthroscopic assessments in complex injuries. Explainable AI techniques further enhanced clinical interpretability, localizing pathologies such as meniscal tears and ligament injuries. Traditional imaging techniques such as X-rays, CT scans, and MRIs each have distinct limitations in the context of knee injury diagnosis. Conventional X-rays primarily provide insights into significant bone fractures, but they fall short in visualizing abnormalities related to ligaments, menisci, and soft tissue injuries, which are crucial for comprehensive knee assessments^44^. While CT scans offer a more detailed view of bone structures and can identify subtle fractures, they still do not provide information about soft tissue pathologies, limiting their utility in diagnosing comprehensive knee injuries^45^. On the other hand, MRI is the gold standard for soft tissue evaluation, capable of revealing ligament tears, meniscal injuries, and bone contusions^46,47^. However, its high cost and the associated long wait times for patients can delay necessary interventions, thereby exacerbating patient suffering. In contrast, AI-driven solutions can significantly mitigate these challenges. By extracting additional insights from X-ray images, AI can identify abnormalities that cannot be found with naked human eyes^48^. By leveraging AI to analyze X-ray data, the model can provide valuable insights that traditional X-ray examination alone cannot capture, thus enhancing the diagnostic process^49^. AI has increasingly been applied to the detection and classification of knee abnormalities using various imaging modalities such as X-rays, CT scans, and MRIs. These AI-driven approaches aim to enhance diagnostic accuracy, reduce interpretation time, and improve inter-reader agreement among clinicians. AI applications in X-ray imaging have also shown promising results. A study utilizing a Vision Transformer model for detecting bone tumors in children’s knee X-rays achieved high accuracy, sensitivity, and specificity, suggesting that AI can facilitate early diagnosis and improve treatment outcomes^43^. Similarly, deep learning models have been applied to automate the diagnosis of knee osteoarthritis using X-ray images, achieving high classification accuracy, particularly in distinguishing between normal and severe cases^50^. Previous studies on automated knee injury lesion assessment have primarily focused on MRI based severity staging, with most literature emphasizing lesion detection through binary classifiers^51^. Liu et al.^52^. presented a fully automated deep learning system designed to detect ACL tears in knee MRI images, using arthroscopy as the reference standard. The system employs two deep convolutional neural networks to isolate the ACL, followed by a classification network to identify structural abnormalities. In a retrospective analysis of 350 subjects, the detection system achieved a sensitivity and specificity of 0.96. The performance was comparable to that of clinical radiologists. However, the study employed three individually trained CNNs in a cascaded manner, which increased the training burden, in addition the study involved relatively small samples and transfer learning wasn’t used. Bien et al.^53^. utilized MRNet, a two-dimensional deep learning architecture for binary classification, to detect ACL and meniscus lesions. The system achieved AUCs of 0.965 (95% CI: 0.938, 0.993) for ACL tear detection and 0.847 (95% CI: 0.780, 0.914) for meniscal tear detection. In our study, we achieved an accuracy of 85% in identifying ACL injury and a meniscus injury accuracy of 90%. The above-mentioned articles focused on MRI and mainly binary classifications. The current diagnostic standard for evaluating soft tissue injuries in the knee remains MRI or CT, given their superior ability to visualize ligaments, cartilage, and other non-osseous structures, while X-rays are primarily limited to assessing bony anatomy and joint alignment^54,55^. referrals and optimizing the use of advanced modalities^56^. The proposed algorithm not only aids in initial assessments but also streamlines the diagnostic pathway, suggesting MRI only for patients with severe conditions and potentially reducing unnecessary examinations for those with less severe issues. This approach ultimately enhances patient outcomes and promotes healthcare efficiency by ensuring timely and appropriate interventions tailored to individual patient needs. In daily clinical practice, the prevalence of multiple concurrent knee abnormalities necessitates a nuanced diagnostic approach. Traditional single-label models often restrict the analysis to one condition at a time, which can overlook the complexity of knee injuries that frequently present with multiple pathologies, such as ligament tears alongside meniscal damage^57,58^. However, our AI-driven multi-label deep learning approach leverages subtle radiographic features—such as joint space narrowing, effusions, and bony alignment changes—that may serve as indirect indicators of underlying soft tissue pathology, especially when trained on multi-label datasets that capture complex interrelationships between findings^56,59^. Importantly, this method is positioned as a screening or triage tool rather than a direct diagnostic replacement for advanced imaging, offering significant technological innovation by potentially flagging cases that warrant further MRI/CT evaluation^60^. The accessibility, cost-effectiveness, and portability of X-ray imaging, combined with AI’s ability to process large datasets and recognize nuanced patterns, could profoundly impact care in resource-limited settings—such as rural, remote, or disaster environments—by reducing unnecessary referrals and optimizing the use of advanced modalities^56^. Building upon the foundational work presented, several promising avenues for future research were identified to enhance the clinical utility and robustness of our AI-driven knee abnormality screening model. First, to bolster the generalizability and mitigate potential biases stemming from our dual-center dataset, future studies should focus on enrolling larger, multi-center cohorts from diverse geographical and clinical settings. This expansion must include a specific focus on recruiting non-Asian patient populations to prospectively validate the model’s performance across different ethnic backgrounds, a critical step for ensuring equitable and unbiased clinical deployment. Second, while this study established the effectiveness of using only AP radiographs, subsequent research will aim to incorporate additional X-ray views, such as lateral, oblique, or even dynamic stress views. Integrating these complementary projections could provide more comprehensive anatomical information, potentially enhancing the model’s ability to detect subtle pathologies, particularly those affecting soft tissues, and further improving diagnostic accuracy for conditions currently challenging to identify on AP images alone. Third, the rapid advancements in deep learning necessitate continuous model refinement. Future work will explore updating or integrating alternative, more advanced model architectures beyond the YOLOv11 framework utilized for localization, potentially moving towards end-to-end solutions that streamline the diagnostic pipeline and further boost precision. While our study established a robust baseline using CNN architectures, a key future direction will be to explore the potential of Vision Transformers (ViTs). Finally, because X-ray imaging lacks detail in visualizing soft tissues, future research must integrate complementary clinical modalities. Combining X-ray AI predictions with patient clinical histories, physical examination findings, laboratory test results, or even limited, targeted MRI data could lead to a more comprehensive and accurate diagnostic system, ultimately facilitating a more personalized approach to knee injury management and aligning with the principles of precision medicine. ## Limitations Our study possessed several limitations that warrant consideration. First, regarding the study cohort and ground truth, the dataset was derived from only two medical centers, which may limit the broader generalizability of our findings. Furthermore, while MRI served as our primary reference standard for soft tissue injuries, arthroscopic confirmation was not performed on all patients, which may have introduced variability in the ground-truth labeling. Second, concerning the input data, a significant constraint of our approach was the exclusive use of anteroposterior (AP) radiographic images. The omission of lateral and other specialized views inherently limited the anatomical information available to the model and likely constrained its diagnostic performance for pathologies better visualized from other angles. In addition, the specific X-ray and MRI acquisition protocols were not fully standardized between the participating institutions, introducing a potential source of data heterogeneity. Third, from a methodological standpoint, our investigation centered on the performance of a specific architecture based on YOLOv11 without a comparative analysis against other deep learning models. The model was also designed for binary classification of soft tissue injuries and did not perform more granular grading. Finally, the analysis did not incorporate patient clinical histories or physical examination findings, which are integral to a comprehensive diagnostic assessment in practice. Future large-scale, multi-center studies that incorporate multiple radiographic views, standardized protocols, and clinical data are necessary to validate and enhance these findings. ## Conclusions Our study introduced a pioneering AI-driven, X-ray-based multi-label deep learning model for screening knee abnormalities, demonstrating a novel capability to overcome traditional diagnostic limits of plain radiographs. By first employing a robust YOLOv11-based localization model to precisely focus on the knee joint region, and then leveraging a multi-label classification framework, our system effectively identified and classified multiple co-existing conditions simultaneously. This targeted approach allowed the model to extract subtle radiographic cues that correlated with soft tissue injuries, contributing to its notable accuracy. This innovation boosts diagnostic efficiency, facilitates severity-based treatment prioritization, and critically, by reducing reliance on costly, invasive tests, enhances access to timely care, particularly in resource-limited settings. This model holds significant promise to transform initial knee injury diagnostics, ultimately improving patient outcomes and healthcare effectiveness. ## Supplementary Information Below is the link to the electronic supplementary material. Supplementary Material 1