Authors: Nida Baig, Kabir Syed Gyasudeen, Tanmoy Bhattacharjee, Jahanzeb Chaudhry, Sabarinath Prasad
Categories: Research, Accuracy, Artificial intelligence, Diagnosis, Lateral cephalometric, Orthodontics
Source: BMC Oral Health
Authors: Nida Baig, Kabir Syed Gyasudeen, Tanmoy Bhattacharjee, Jahanzeb Chaudhry, Sabarinath Prasad
Compare the accuracy and diagnostic concordance of three commercially available AI-based lateral cephalometric tracing software.
Sixty-three lateral cephalometric radiographs were analyzed using semi-automatic (Dolphin Imaging Systems LLC) and AI-based software programs (WebCeph™, Cephio, and Ceppro DDH Inc.). Intra- and inter-observer reliability were assessed for human expert measurements, and repeated-measures one-way ANOVA was used to compare the AI and human expert measurements. The diagnostic performance was evaluated using sensitivity and specificity tests.
Human expert reliability was excellent (ICC > 0.9) for most cephalometric parameters. Compared to human experts, significant differences were observed for all three AI-based cephalometric programs (WebCeph™ – 10 of 11, Cephio – 7 of 11, and Ceppro DDH Inc. – 7 of 11 cephalometric measurements). Variations exceeding two units were noted for most parameters, and differences in defining the sagittal and vertical skeletal patterns, dental, and soft tissue characteristics were observed.
All three AI-based tracing programs showed inaccuracies compared to human expert measurements and lacked reliability in measuring key cephalometric parameters. Clinicians should exercise caution when relying solely on AI-based analyses for orthodontic treatment planning and assessment.
The online version contains supplementary material available at 10.1186/s12903-024-05032-9.
More than seventy years have passed since the first cephalometric analysis was published [1]. During this period, the shortcomings of lateral cephalometry have been better understood. Notably, in a list of the most cited papers in orthodontics [2], the top spot was held by an article on the sources of error in cephalometric measurements and analysis [3]. The integration of artificial intelligence (AI) into lateral cephalometry offers a promising opportunity to decrease errors, automate analyses, optimize diagnostic accuracy, and reduce measurement bias, particularly when assessing treatment outcome [4]. Recently, fully automated AI-based software / applications have been introduced in an attempt to improve the accuracy and reliability of cephalometric measurements. However, their performance can be affected by the developmental stage of dentition [5], congenital anomalies [6], and image characteristics [7]. Sources of inaccuracy affecting the performance of AI-based lateral cephalometric applications also include the training dataset (digital vs. analog), the presence of artifacts in the images, operator errors during image calibration, and possibly even differences in their underlying algorithms.
Of the commercially available AI-based cephalometric tracing software / applications [8–10], the effectiveness of CephX, AudaxCeph, Ceppro, WebCeph^TM^, CephNinja, CephNet, and CS Imaging V8 have been studied. However, information regarding the algorithms that form the basis of only a few AI-based lateral cephalometric programs is available. For instance, WebCeph^TM^ is known to be based on a deep learning algorithm that uses a convolutional filter and pooling layer to extract features and analyse patterns [11] and the CEPPRO DDH Inc. platform is based on the deep learning algorithm You-Only-Look-Once version 3 (YOLOv3) [12]. On the other hand, information on the underlying algorithm of AI-assisted lateral cephalometric tracing programs such as the CEPHIO program is unavailable. Although AI-based approaches using deep learning algorithms have been shown to be relatively accurate for detecting landmarks on cephalometric radiographs, it remains unclear whether they are similar to, more accurate, or less accurate than clinicians [13]. To ensure clinical usefulness and dependability in orthodontic practice, it is essential to assess the accuracy of AI-assisted lateral cephalometric tracing solutions. Therefore, the accuracy, efficiency, and diagnostic concordance of the three commercially available AI-assisted lateral cephalometric tracing software were compared in this study. We hypothesized that AI-assisted lateral cephalometric tracing software based on different algorithms can achieve accuracy and efficiency comparable to semi-automatic lateral cephalometric tracing software for key cephalometric measurements.
The study was approved by the institutional review board of Mohammed Bin Rashid University (MBRU IRB-2023-381) under the exempt category. General consent to use the data for scientific purposes was obtained from the patients at the time of registration and the need for consent was waived by the institutional review board of Mohammed Bin Rashid University.
Lateral cephalometric radiographs for this retrospective study were obtained from the Dubai Dental Hospital. Pretreatment lateral cephalograms of subjects who reported for orthodontic treatment at the Dubai Dental Hospital between January 2021 and October 2023 were screened using patient management software (Dental 4 Windows, Centaur Software^Ⓒ^).
G*Power software (version 3.1, Heinrich-Heine-University Dusseldorf, Germany) was used for statistical power calculation. The sample size calculation found that for a power of 80% and a significance level of 0.05, a minimum sample size of 35 images was required, similar to the results of a previous study [14]. Lateral cephalograms (n = 63) of subjects in the late mixed / permanent dentition without cleft lip and palate, asymmetry, or previous orthodontic treatment were included in the study. Lateral cephalometric radiographs with technical or patient-related artifacts were excluded. All cephalometric radiographs included in this study were obtained using the same cephalometric machine (Veraviewepocs 3D R100 Pan/Cep, J. Morita Corporation, Japan) and the following magnification factor 1, pixel-X 96.0 μm, pixel-Y Size: 96.0 μm, kVp: 90, mA:8.8, and exposure 5.7 s). To ensure complete anonymity, patient identifiers on all cephalometric radiographs were masked using the markup toolbar (MacBook Air, MacOS Catalina Version 10.15.7; Apple Inc., Cupertino, CA, USA) before analysis.
The cephalometric landmarks chosen in this study were defined to ensure consistency in human landmark identification (Table 1). Nine hard-and two soft-tissue cephalometric landmarks were visually located and labelled using digital tracing software (Dolphin Imaging Systems LLC, Chatsworth, USA) by a third-year orthodontic resident (NB). All landmarks were verified and appropriate adjustments were made by an experienced (> 15 years) orthodontist (SP). Once the accuracy of the manually located landmarks was verified, measurements were automatically derived. The second set of tracings was performed by a recent orthodontic graduate (KG). No more than five tracings were performed in a single session. Cephalograms (n = 30) were retraced by both investigators after a minimum interval of four-weeks.
Table 1Hard and soft tissue cephalometric landmarks and cephalometric planesLandmarkAbbreviationDefinitionSellaSCenter of the pituitary fossa of the sphenoid bone.NasionNIntersection of the internasal suture with the nasofrontal suture in the midsagittal planeA pointADeepest point of the curve of the maxilla, between anterior nasal spine (ANS) and dental alveolus.PorionPoHighest point of the ear canal; most superior point at the external auditory meatus.OrbitaleOrLowest point of the roof of the orbit; most inferior point of the external border of the orbital cavity.PogonionPgMost anterior point on the mid-sagittal symphysis.GnathionGnMidpoint between the most anterior and inferior point on the bony chinMentonMeMost inferior point of the symphysis.GonionGoMost convex point along the inferior border of the ramusSubnasaleSnPoint where the nose connects to the center of the upper lipSoft tissue PogonionPg’Point on the anterior curve of the soft tissue chin. Cephalometric planes Sella- Nasion planeSNSella - NasionFrankfort–Horizontal planeFHPorion - OrbitaleMandibular PlaneGo-MeLine joining Gonion and Menton
The sequence of steps followed for the generation of the tracing reports for the fully automated AI-based cephalometric analysis briefly includes uploading to the online platform of the AI-based (WebCeph^TM^, Cephio, and Ceppro DDH Inc.) software, selecting the custom analysis, and automatic tracing followed by generation of the tracing results without additional human input. A randomly chosen number (n = 10) of radiographs were re-uploaded into all three AI-based programs after a week. The time taken by the three different AI-based tracing software and human investigators was recorded using a digital stopwatch. The cephalometric values of the variables were recorded on an online spreadsheet for statistical analysis.
SPSS for Windows (version 23.0; SPSS Inc., Chicago, IL, USA) was used for statistical analyses. The level of significance was set at p < 0.5. Interclass correlation coefficients (ICC) were calculated to determine intra- and inter-investigator reliabilities of the two human experts and AI measurements. The median values of the measurements obtained by the two experts were used for further analysis. The performance of the three AI programs versus the human expert median was evaluated using repeated-measures one-way Analysis of Variation (ANOVA). A post-hoc pairwise comparison with Bonferroni correction was performed to identify individual differences among each AI program and the human output for each parameter.
To further understand the diagnostic performance of AI-based cephalometric programs compared with that of human experts, confusion matrices were created. The true positive, true negative, false positive, false negative, sensitivity, and specificity values were calculated using confusion matrices. The human expert values were chosen as the reference standard.
For AI-based measurements generated by the three software after a one-week interval, reliability tests showed excellent agreement (ICC > 0.99). Intra-class agreement values for the human investigators were also high (ICC > 0.9). Between two human investigators, inter-class agreement values were > 0.9 for all parameters, except SNA (ICC = 0.89). Hence, the median of human measurements was used as the reference standard for subsequent analyses. Repeated-measures one-way ANOVA showed statistically significant differences among the four entities (human and three AI programs) for all the parameters. The graphs of post-hoc pairwise comparisons with Bonferroni corrections are shown in Fig. 1. The data distribution was normal, did not violate sphericity (Geisser-Greenhouse’s epsilon value > 0.8), and the variances of the groups were equal.
Fig. 1Bar graphs of post-hoc pairwise comparisons with Bonferroni correction for different cephalometric parameters. Groups with significant differences between them are shown by asterisks, with the number asterisks indicating the degree of difference (*p ≤ 0.05, **p ≤ 0.01, *** p ≤ 0.001, ****p ≤ 0.00001)
Significant differences were observed when WebCeph™ (10 of 11 parameters), CEPHIO (7 of 11 parameters), and CEPPRO DDH Inc. (7 of 11 parameters) were compared to the reference standard (Fig. 2).The difference in mean between the reference standard and WebCeph™ was high (2–9 units difference) for all parameters except U1-Apog, whereas the difference was high compared to CEPHIO for three parameters (U1-NA, Interincisal angle, and nasolabial angle) and CEPPRO DDH Inc. for seven parameters (FMA, U1-NA, U1-Apog, L1-MP, Interincisal Angle, and Nasolabial Angle). The time taken by the different entities to produce the lateral cephalometric tracing report is presented in Table 2.
Fig. 2Differences between the human gold standard and AI technique parameter measurements for each parameter. Red cells indicate statistically significant differences between human and AI software measurements, whereas green cells indicate no statistical difference. The numbers in each cell are the mean differences between the human and AI software measurements
For the diagnostic parameters, differences in defining the sagittal and vertical skeletal patterns based on the values of the three AI-based cephalometric software were observed (Table 3). These differences were also noted in the dental and soft tissue measurements. However, there were instances when the values for sensitivity and specificity were 100% (the value that a perfect diagnostic method should have) or close to 100%. Overall, the sensitivity and specificity were low (< 80%) for the parameters in all AI software. Notable exceptions were detection by CEPHIO of FMA (~ 80%), SN-MP (> 80%), WebCeph™ of U1-APog (> 90%), and interincisal angle (≥ 80%). A heatmap of the accuracy of the various parameters and software is presented in Supplementary Table 1.
Table 2Time taken by different modalities for lateral cephalometric tracings and report generation Tracing modalityTime for landmarking and analysis (seconds)Time for report generation (seconds)Human expert 11702Human expert 21652WebCeph™3110Cephio7.5120CEPPRO DDH Inc343
Table 3The distribution of diagnostic outcomes using the three AI-based tracing software programsParameterSoftwareSensitivitySpecificitySNAOrthognathicPrognathicRetrognathicOrthognathicPrognathicRetrognathicWEBCEPH31.7100.09.159.144.2100.0CEPHIO0.0100.00.0100.00.0100.0CEPPRO90.236.472.754.596.296.2SNBOrthognathicPrognathicRetrognathicOrthognathicPrognathicRetrognathicWEBCEPH78.4100.031.842.386.4100.0CEPHIO89.275.086.484.694.997.6CEPPRO81.150.0100.092.398.385.4ANBSkeletal Class ISkeletal Class IISkeletal Class IIISkeletal Class ISkeletal Class IISkeletal Class IIIWEBCEPH29.2100.014.384.645.2100.0CEPHIO79.296.942.987.287.198.2CEPPRO91.784.457.179.593.5100.0Y-axisNormalHorizontalVerticalNormalHorizontalVerticalWEBCEPH83.9100.050.071.485.0100.0CEPHIO92.966.750.057.195.098.3CEPPRO96.466.725.042.996.7100.0FMANormalHorizontalVerticalNormalHorizontalVerticalWEBCEPH75.7100.039.150.083.3100.0CEPHIO86.5100.078.380.898.390.0CEPPRO86.5100.069.673.193.397.5SN-MPNormalHorizontalVerticalNormalHorizontalVerticalWEBCEPH60.4100.060.086.767.996.6CEPHIO83.380.080.080.088.796.6CEPPRO83.360.080.066.794.391.4U1-NANormalProclinedRetroclinedNormalProclinedRetroclinedWEBCEPH85.035.9100.041.9100.094.9CEPHIO85.056.475.058.1100.094.9CEPPRO95.064.1100.067.4100.098.3U1-ApogNormalProtrudedRetrudedNormalProtrudedRetrudedWEBCEPH91.397.4100.097.595.898.4CEPHIO91.389.70.087.5100.096.8CEPPRO65.2100.00.097.566.7100.0L1-MPNormalProclinedRetroclinedNormalProclinedRetroclinedWEBCEPH82.991.770.081.886.3100.0CEPHIO85.475.070.072.790.298.1CEPPRO36.60.00.068.294.128.3Interincisal angleNormalIncreasedDecreasedNormalIncreasedDecreasedWEBCEPH93.880.085.785.198.3100.0CEPHIO87.560.095.291.598.395.2CEPPRO56.3100.066.770.287.9100.0Nasolabial angleNormalIncreasedReducedNormalIncreasedReducedWEBCEPH60.735.5100.048.6100.078.0CEPHIO82.154.8100.060.0100.091.5CEPPRO35.796.850.091.443.8100.0
In this study, cephalometric measurements of three commercially available AI-based cephalometric tracing programs and diagnoses derived from them were compared with semi-automatic digital cephalometrics.The selected linear and angular measurements were representative of the vertical and sagittal measurements that orthodontists frequently rely on for case planning. To better understand and assess the performance of the three AI programs, no human input was used.
For the computer-assisted semi-automatic cephalometric analysis program, landmarks were manually identified, and distances and angles were automatically calculated. Although this helps reduce the inaccuracies of pure manual measurements, errors in landmark identification can still occur because operator experience affects random errors [15]. The human expert reliability values for the different measurements in this study are comparable to those in recent reports [16]. Other than the SNA angle, the reliability scores for both human experts in this study were excellent. This is not surprising, as difficulties in identifying point A have been previously reported and might explain tracing variance [17, 18]. Interestingly, erroneous detection of point A has also been reported by AI-based cephalometric programs, leading to the misclassification of sagittal skeletal patterns [19]. This is also not surprising, as human labelled training data are used as the foundation for validating the performance of AI programs.
The error range for significant clinical differences (> 2° or 2 mm) in semi-automated cephalometric analysis [20] has also been extended to fully automated AI performance studies. Successful detection rates for skeletal landmarks using AI-based methods have been shown to be within a 2 mm range [21, 22]. However, in this study the error size was not within the widely acceptable range for most of the results generated by the different AI-based programs. This could be because some cephalometric landmarks used in this study, particularly the Menton, Gonion, Lower incisor apex, Orbitale, Porion and Nasion, are questionable, unpredictable, and exhibit errors irrespective of the method (manual, semi-automatic, fully automatic) employed for their identification [18, 23, 24]. The detection of the average of bilateral points and the overlapping of cranial base structures are among the factors contributing to difficulties in accurate landmark identification.
WebCeph™ has been reported to produce results similar to those of human experts for both linear and angular measurements [25–27]. High intra-investigator reliability for certain measurements (ANB, FMA, IMPA/L1 to MP, LL to E-line, L1 to NB, L1 to NB, and S-N to Go-Gn) but lower reliability for other measurements (UL to E-line, SNA, SNB, and U1 to NA) have been reported for WebCeph™. In this study, a statistically significant difference (> 2 units) in the mean was observed between the human expert and WebCeph™ measurements for all parameters, with the exception of U1-APog. Inconsistencies in landmark identification and significant errors in measurements have also been reported for WebCeph™ [28, 29]. WebCeph™ measurements for the nasolabial angle in this study demonstrated a greater magnitude of error (9.3) than previously reported (2.7 ± 4.9) [30]. These errors may be attributed to the fact that the nasolabial angle landmarks are situated on a large-radius curve [31]. Substantial discrepancies were also observed in the angular measurements of the L1-MP angle values generated by WebCeph™, potentially attributable to inconsistencies in the mandibular plane identification.
Initial comparisons between the two deep learning algorithms, YOLO and single-shot multibox detector (SSD), indicated that YOLO exhibited fewer errors and greater potential for automated AI-based tracing [32]. The landmarks found to have significant differences (> 2 mm) in the early version of the YOLO algorithm were the U1 apex, L1 apex, Basion, Gonion, and Orbitale. Subsequent versions using the deep learning algorithm YOLOv3 were found to be comparable to those of human examiners [12, 33]. However, in this study, significant differences compared to human expert measurements were observed for seven of the eleven parameters for CEPPRO DDH Inc. A significant difference in the mean (> 2 units) between the human experts and CEPPRO DDH Inc. was noted for the angular measurements of the mandibular plane, U1-NA, L1-MP, interincisal angle, and nasolabial angle, as well as for the linear measurement of U1-APog.
The Cephio cephalometric tracing platform was included in this study because of the scarcity of studies evaluating its performance. No statistically significant differences were observed in a study that compared the Cephio program with WebCeph™ or the manual method. However, the SNA, SNB, and ANB angles were relatively higher with WebCeph™ than with Cephio or manual cephalometric analyses. The mandibular and occlusal plane angles were higher with Cephio than with WebCeph™ or the manual method. In this study, seven of the eleven measurements differed from the human reference standard for Cephio, with a significant difference (> 2 units difference) for three angular measures (U1-NA, interincisal angle, and nasolabial angle) [34].
Overall, the findings of this study contradict the conclusions of previous studies that reported clinically acceptable results for AI-based cephalometrics [14, 27]. This study concurs with reports that AI-assisted cephalometric analysis has significant errors [29], potentially attributable to poor landmark identification and inconsistency of measurements. In this study, the algorithms underlying the cephalometric tracing software were known for two (WebCeph™: Convolutional neural network; CEPPRO DDH Inc.: YOLOv3) software and not known for the third (Cephio). For image identification tasks, differences have been reported between YOLO and Convolutional Neural Network [35] architectures and may explain the variability noted in this study. Establishing clarity on the algorithms employed is key to explainable AI [36, 37], and is necessary to mitigate differences in the future. Furthermore, discrepancies may also be partially attributed to ethnic and racial differences in the datasets used for training and validation in previous studies compared with the present investigation.
Near-instantaneous digitization and analysis were observed in all three AI programs, and most of the time was expended on report generation (Table 2). Again, variations in the speed at which the algorithms identify anatomical landmarks may explain the differences in speed of subsequent report generation. Also, if the dataset used for training failed to encompass a sufficiently wide range of anatomical variations, the AI software may struggle in certain situations, leading to longer report generation times. Commercial software periodically introduce new features or improvements, changing how measurements are taken and therefore the time for report generation as well. The efficiency and time required to complete a task are inversely proportional. In this study, the time taken by all three programs was similar and significantly shorter than that of manual semi-automatic cephalometric analysis, which is consistent with recent reports [26]. However, none of the three programs investigated in this study had an accuracy comparable to the human expert, disproving our hypothesis. Although AI-based programs conduct rapid analyses, their accuracy is compromised, highlighting the necessity for human involvement. Previous studies [16, 38, 39] have emphasized the importance of human input before generating reports, which calls into question the purported time-saving benefit.
The diagnostic agreement between AI-based programs after the creation of confusion matrices for various parameters was investigated. Previous studies have reported misclassification of craniofacial morphology and differences in classification accuracy based on the AI model used [19, 21]. Incorporating a convolutional neural network into a single-step end-to-end system exhibited > 90% sensitivity, specificity, and accuracy for vertical and sagittal skeletal diagnosis, with higher accuracy for vertical diagnosis [40]. In contrast, machine learning algorithms for classifying maxillofacial morphology have reported a higher accuracy in classifying sagittal morphology than vertical morphology [41]. However, the exact values of sensitivity, specificity, and accuracy for vertical and sagittal skeletal diagnoses have not been reported in these studies, making it challenging to draw comparisons with the results of the present study. Hard and soft tissue evaluations for the same subject may also show differing analytical outcomes [42]. Therefore, future AI-based tracing programs need to integrate capabilities to highlight such discrepancies.
It is essential to consider the limitations of the present study. First, all lateral cephalometric radiographs utilized in the study were acquired using the same machine and operator, thereby reducing the generalizability of the study. Next, to facilitate comparisons between the different AI software programs, there were constraints on the number of cephalometric parameters that could be included, resulting in only one linear measurement. Interestingly, tracing errors have been reported to be greater for linear measurements than for angular measurements [43]. Additionally, the amount of variation of cephalometric values about their mean also should be considered in the diagnostic performance of AI-based cephalometric programs in future studies. Finally, information on the underlying algorithms is available for only two of the three programs investigated in this study.
All three AI-based, fully automated lateral cephalometric tracing programs, WebCeph™, Cephio, and Ceppro DDH Inc., showed significant inaccuracies compared to human expert-based tracings of key cephalometric parameters. Commercially available AI-based cephalometric tracing programs lack consistency in diagnosing skeletal, dental, and soft-tissue characteristics.
Below is the link to the electronic supplementary material.
Supplementary Material 1