Authors: Laura C.M. Ndjonko, Nikol N. Kralimarkova, Yashoswini Chakraborty, Zayn S. Bajwa, Jasmine X. Zimmer, Ayomikun A. Taiwo, Imani N. Bah, Sami S. Khan, Eric K. Holder
Categories: Review Article, lumbar foraminal stenosis, grading systems, cross-sectional imaging, Lee et al., systematic review
Source: Spine Surgery and Related Research
Authors: Laura C.M. Ndjonko, Nikol N. Kralimarkova, Yashoswini Chakraborty, Zayn S. Bajwa, Jasmine X. Zimmer, Ayomikun A. Taiwo, Imani N. Bah, Sami S. Khan, Eric K. Holder
Symptomatic lumbar foraminal stenosis (LFS) occurs when the neuroforamen narrows, compressing the exiting spinal nerve, leading to symptoms such as radicular pain, paresthesias, and potentially weakness. Although cross-sectional imaging studies are used for diagnostic purposes, there is no clear consensus as to which grading system best evaluates LFS, predisposing to inconsistencies in care. This systematic review aimed to evaluate and compare existing published grading systems for LFS to identify (1) systems most used within the literature and (2) the most effective and reliable method for classifying anatomic severity and clinical symptom correlation.
This study is a systematic review following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines, analyzing available literature on grading systems for LFS, level of evidence IV. A comprehensive search of PubMed, Embase, and Cochrane Trials was conducted from inception through July 2024. Eligible studies were evaluated for methods, bias, sample size, patient demographics, imaging modalities, and grading systems. Bias was assessed using the Methodological Index for Non-Randomized Studies. Data were synthesized narratively and descriptively.
The review included 35 studies, most using magnetic resonance imaging (88.6%). Seven grading systems have been identified. The original Lee et al. grading system was the most frequently used LFS grading system (69%), followed by Wildermuth et al. (14.3%). Notably, artificial intelligence (AI) grading systems were included in two studies (5.7%). Findings regarding symptom correlation were mixed.
The Lee et al. grading system remains the most used grading system for LFS in the literature and is reliable. Several small studies found an association between the Lee et al. system and clinical symptoms/treatment outcomes; however, this was not universally found. Further investigation is needed to validate the newer grading. The introduction of AI may offer promise for refining the diagnostic and clinical utility of published LFS grading systems.
Lumbar foraminal stenosis (LFS) is a condition in which the neuroforamen is narrowed, with the potential for compression of the exiting spinal nerve root^1-4)^. Although rare congenital cases exist, LFS is typically a manifestation of disk pathology (i.e., lateral disk herniation), osseoligamentous degenerative changes, and/or malalignment (i.e., spondylolisthesis), which are all radiologic findings that become increasingly prevalent with age^1-3)^. When the neuroforamen narrows, it can compress the exiting nerve and lead to radiculopathy, which may present as pain or paresthesias in a concordant dermatomal pattern and weakness in a concordant myotomal pattern potentially associated with pathologically depressed reflexes^1,3)^. These radicular symptoms have the potential to significantly limit a patient's quality of life^3)^. Alternatively, patients are frequently asymptomatic despite imaging findings indicating LFS and nerve root compression^3)^. Predicting when and whether symptoms will develop in an individual is difficult^1)^. There is no clear consensus as to which grading system best evaluates LFS^4)^. Description of the severity of LFS is often qualitatively subjective, dependent on the training and perspective of the imaging reviewer^4)^. Thus, it is important to clarify which validated grading systems exist and their clinical applicability.
Magnetic resonance imaging (MRI) and computed tomography (CT) scans are the most commonly used cross-sectional imaging modalities to assess the presence and extent of LFS^1,3)^. MRI is generally regarded as the standard of care owing to its ability to differentiate soft tissue structures and its lack of ionizing radiation exposure^1,3,4)^. The utility of CT scan in this clinical scenario lies in its ability to differentiate calcified disks, bony osteophytes, facet hypertrophy, and ossification of ligamentous structures, which may have implications for care planning^1-3)^. The sensible incorporation of these imaging modalities is often a necessary diagnostic adjunct to a thorough patient clinical evaluation and paramount to clinical decision-making, such as whether to pursue conservative management or surgical intervention^1-3)^. Even with the available literature on LFS and grading systems, in clinical practice, there is often variable description of the severity of LFS among imaging reviewers based on subjective interreader evaluation^4)^. Thus, the authors believe that it is a worthy pursuit to review and summarize the literature regarding grading systems for LFS across imaging modalities and their reliability. Lack of awareness or lack of consistent utilization of established grading systems allows a greater margin of variation in interpretation among imaging reviewers, with the potential to foster inconsistencies in clinical practice, and this has implications for diagnosis and treatment recommendations^3)^. Although systematic reviews of cervical foraminal stenosis grading systems exist, it is comparatively sparse for the lumbar spine^3-5)^. To our knowledge, limited prior reviews on this subject matter have focused on management/treatment of LFS or are narrowly focused only on validated MRI grading systems^3,4)^.
The purpose of this systematic review was to evaluate and compare the existing grading systems for LFS, inclusive of all imaging modalities, with the goal of identifying the most studied and reliable systems for assessing this condition. Moreover, high-level spinal care is a multidisciplinary endeavor incorporating physicians with diversified training and skillsets (orthopedists and/or neurosurgeons, physiatrists, anesthesiologists, radiologists) in addition to physical and occupational therapists ideally working in a coordinated fashion to promote optimal patient outcomes. Clinician awareness, interpretation, and value ascribed along with nomenclature used to describe imaging findings can vary substantially because of the heterogeneity in the training of spine care specialists. This review aims to foster cohesion by furthering the awareness of published grading systems for LFS and their clinical applicability, diagnostic parameters, and intra/interreader reliability.
This study was not previously registered as a protocol within the International Prospective Register of Systematic Review and follows the guidelines set by the Preferred Reporting Items for Systematic reviews and Meta-Analyses criteria^6)^. The PRISMA 2020 Checklist is available in the supplementary files.
Original studies that incorporated grading systems for LFS were included in this review. Studies were excluded if they lacked full text in English, grading systems, or duplicates, or were not specific to LFS.
A systematic literature search of PubMed, Embase, and Cochrane Trials was conducted from inception through July 2024. The Boolean Query was ([“Lumbar Stenosis”] AND [“Grading system” OR Grading]). This query allowed us to include all possible variations of lumbar stenosis before manual exclusions for studies not specific to LFS. Titles and abstracts were manually screened by two reviewers on the basis of our eligibility criteria, and articles with questionable eligibility were reviewed as full texts. Full-text screening was conducted by the remaining authors, and any conflicts were resolved by the corresponding author.
Data extraction was performed by six reviewers using an Excel spreadsheet. Patient demographic variables collected included study type, sample size, age, sex, imaging evaluated, lumbar spine levels evaluated, and specifics of the LFS grading systems (Table 1). The Methodological Index for Non-Randomized Studies (MINORS) criteria were used to assess bias and the methodologic quality of articles^7)^. The MINORS scale is out of 16 for non-comparative studies and 24 for comparative studies, each item ranges from 0 to 2^7)^.
Descriptive summaries of the study characteristics were accomplished using the reported means, standard deviations, and ranges. The LFS grading systems were reviewed and presented in a narrative manner.
A total of 312 articles were identified using our search criteria. After abstract screening, 92 full texts were retrieved, and 35 LFS articles were included in this study (Fig. 1). The MINORS scores for bias and quality of all the studies are presented in Fig. 2. A total of 6,719 patients across all studies were included in the analysis (Table 1). The average age of participants showed considerable variation, ranging from 42.2 years to 73 years. The sex distribution across studies was relatively balanced, with 2,829 women and 2,330 men reported in total. The reported designs were predominantly retrospective (n=26, 74.3%)^8-34)^, with a few prospective (n=8; 22.9%)^14,35-41)^ and one nonrandomized controlled study (n=1, 2.9%)^42)^.


Most studies used MRI as the primary imaging modality (n=31, 88.6%)^8-11,14-23,25-33,35-42)^ with some incorporating CT scans (n=4, 11.4%)^12,13,24,34)^. The assessed lumbosacral spine levels ranged from L1-L2 to L5-S1, with specific findings listed in Table 1.
Seven grading systems for LFS were identified (Table 2). The original Lee et al. grading system was the most frequently used, appearing in 24 studies (69%)^8,9,11,13-21,24,28,29,32-38,40,42)^. The next most commonly used grading system was Wildermuth et al.^41)^ (n=5, 14.3%)^10,26-28,41)^. Notably, AI was used for grading in two studies (5.7%)^22,23)^.
As indicated in Table 2, the mean inter- and intra-reliability kappa (k) values were collected for all grading systems based on the original studies. Centered on the reliability assessments of two or more imaging reviewers, this scale was often a kappa value of ≤0 indicates no agreement; 0.01-0.20 indicates slight agreement; 0.21-0.40 indicates fair agreement; 0.41-0.60 indicates moderate agreement; 0.61-0.80 indicates substantial agreement; and 0.81 or greater indicates nearly perfect agreement^12)^. In terms of interobserver reliability, the Lee et al.^20)^ grading system showed near-perfect agreement, with an average kappa value of 0.921. This finding was followed by Sartoretti et al.^39)^ and Spinnato et al.^31)^, with interobserver reliability scores with similar near-perfect kappa values of 0.866 to 1.0 and 0.834, respectively. The system of Wildermuth et al.^41)^ showed substantial agreement in interobserver reliability, with a kappa value of 0.62. Haleem-Botchu et al.^12)^ reported moderate agreement, with an interobserver kappa of 0.58. The artificial intelligence (AI) grading system had an average interobserver kappa value of 0.627^22,23)^. Regarding intraobserver reliability, Lee et al.^20)^ showed a near-perfect average kappa value of 0.905, whereas Haleem-Botchu et al.^12)^ and Spinnato et al.^31)^ revealed similar near-perfect average kappa values of 0.935 and 0.928, respectively.
The consistent use of MRI as the primary imaging modality underscores its importance in the diagnosis of LFS^8-11,14-23,25-33,35-42)^. MRI provides a detailed evaluation of the soft tissue structures, which is crucial for assessing the extent of stenosis. Furthermore, recent advances in MRI technology include high-resolution 3-dimensional (3D) thin slice sequences that have the potential to identify the relationship between nerve root and surrounding structures within the neuroforamen more precisely^39)^. The CT scan plays a complementary role in providing additional anatomical details, particularly in complex cases when further delineation of osseous structures is necessary^12,13,24,34)^. Moreover, CT-based evaluation of LFS is useful in scenarios in which MRI evaluation is contraindicated or not easily accessible^12)^. To our knowledge, there is only one systematic review that assessed LFS grading systems, which was published in 2022 and limited its scope to MRI classification systems^4)^. The authors determined that the Lee et al. grading system was the only validated MRI classification system for LFS, with several studies validating the grading system with moderate-to-high reliability^4)^. Our review expands the scope by assessing the frequency at which specific LFS grading systems are used in the literature, incorporates newer grading systems, and evaluates grading systems across imaging modalities.
Interobserver reliability refers to the degree of agreement among different raters, whereas intraobserver reliability indicates the consistency of a single rater's assessments over time^12)^. Both are critical components when assessing the feasibility of a classification system's ability to provide practical and reproducibly consistent results in clinical practice^12)^. Grading systems are generally qualitative given quantitative systems often lack reproducibility^30)^. The Lee et al.^20)^ grading system had the highest interobserver agreement overall, with others indicating comparable near-perfect agreement^31,39)^. The intraobserver reliability of Haleem-Botchu et al.^12)^ (k=0.935) and Spinnato et al.^31)^ (k=0.928) marginally outperformed Lee et al.^20)^ (k=0.905), with all three showing near-perfect agreement. The Lee et al.^20)^ grading system was published in 2010 and is the most studied within the literature, based on our findings. Despite early promising results, additional investigation into newer systems would be helpful to assess their feasibility for use in clinical practice^12,22,23,25,31)^.
Given the existence of several grading systems for LFS and the lack of consensus regarding a preferred system, clinicians involved in imaging evaluation and management for the same patient may rely on an individualized subjective interpretation or use different grading systems to assess the same imaging^23)^.
Wildermuth et al.^41)^ proposed the first MRI grading system for LFS based on qualitative features ranging from grade 1 (normal dorsolateral border of the intervertebral disk and normal form of the foraminal epidural fat―oval or pear-shaped) to grade 4 (advanced stenosis with complete obliteration of the foraminal epidural fat) (Table 2)^41)^. The authors reviewed sagittal T2-weighted fast spin-echo images completed with an open 0.5-T MRI as compared with lumbar myelography for evaluation of the dural sac and the intervertebral foramina^41)^. Interobserver agreement for LFS MRI grade was substantial (k=0.62)^41)^. The authors determined that myelography was not ideal for assessment of the foramen^41)^. They also aimed to assess the effects of various body positions on the intervertebral foramen including supine, upright flexion, and extension positions and found positionally dependent differences in the qualitative foraminal grade to be uncommon but tended to occur in the upright flexion or extension positions and more so in extension (than in neutral) when present in a minority of patients^41)^. The Wildermuth et al.^41)^ LFS grading system has since been used in selected studies, including a 2008 study indicating the non-inferiority of accelerated parallel imaging (GeneRalized Auto-calibrating Partially Parallel Acquisition technique) to conventional MRI sequences for the evaluation of lumbar degenerative spine disease including LFS^26)^. Furthermore, a 2023 study incorporated the Wildermuth et al.^41)^ LFS grading system as one of the assessment tools used to evaluate the correlation between lateral stenosis, posterior disk height/disk degeneration, and clinical symptoms, pain severity, and extent of disability, and found no significant interrelation^27)^. The authors concluded that their findings reinforce symptomatic LFS as a clinical-radiologic syndrome and that comprehensive clinical examination and correlation with imaging are necessary for decision-making^27)^.
Before publication of this grading system in 2010, sparse literature on classification with or without grading systems for LFS existed^20)^. In 1991, Kunogi and Hasue proposed a classification of LFS based on morphology without a stenosis grade, and similarly, in 1998, Wildermuth et al.^41)^ proposed a grading system without consideration of direct nerve root deformity or compression^41,43)^. Lee et al.^20)^ formulated a grading system that considers both perineural fat obliteration and nerve root morphology. Our findings agree with those of Hutchins et al.^4)^, showing the Lee et al.^20)^ grading system has moderate-to-high inter- and intraobserver reliability. In 2010, Lee et al.^20)^ published this qualitative four-grade system for LFS, ranging from grade 0 (absence of foraminal stenosis) to grade 3 (severe foraminal stenosis showing nerve root collapse or morphologic change) primarily based on sagittal T1-weighted MRI images, whereas T2-weighted imaging was used as an additional tool to exclude false-positive findings (Table 2). The average interobserver and intraobserver agreement among all lumbar foramen examined was found to be nearly perfect (k=0.921 and k=0.905 respectively)^20)^. It was determined that a higher incidence of LFS occurred on the left side and involved the lower lumbar segments^20)^. Lee et al.^20)^ acknowledged that a limitation of their grading system was a lack of evaluation for symptom correlation.
Since its publication, the Lee et al.^20)^ grading system has been the most used grading system for LFS in the literature^8,9,11,13-21,24,28,29,32-38,40,42)^, and its reliability has been validated by several studies^8,9,13,16,17,19,28,32)^. Furthermore, several studies have shown a correlation between Lee et al.^20)^ grade and clinical symptoms/outcomes^8,9,16,17,28,29,42)^. In 2012, Park et al.^28)^ published the first study evaluating the correlation of the Lee et al.^20)^ grading system with neurologic signs and symptoms (paresthesia, extremity weakness, funicular or radicular pain). The authors also compared this system with that of Wildermuth et al.^41)^ and determined that the Lee et al.^20)^ system showed slightly better interobserver agreement and good clinical correlation, especially for patients in the younger cohort (<50 years old)^28)^.
Several studies have evaluated the role of the Lee et al.^20)^ grading system in lumbar spinal surgery. A 2017 study evaluated the reliability and utility of the Lee et al.^20)^ grading system in patients with LFS who underwent foraminotomy^16)^. The authors used the grading system to investigate operated neuroforamen versus non-operated asymptomatic neuroforamen for patients with persistent radicular pain refractory to conservative care. The authors determined the mean grade of operated neuroforamens from L3-L4 through the L5-S1 level was more than 2.5, and this was significantly greater than the non-operated neuroforamen (p<0.001)^16)^. They found moderate interobserver agreement for operated neuroforamen (k=0.511) and good agreement for asymptomatic neuroforamen (k=0.696)^16)^. The authors determined that Lee et al.^20)^ grade correlated with clinical features such as pain intensity^16)^. The authors concluded that the Lee et al.^20)^ grading system is a useful tool that allows LFS to be evaluated more objectively and that for grade 3 LFS, surgical treatment can be considered over conservative management^16)^. They also determined the Lee et al.^20)^ grading system was less reliable for symptomatic L5-S1 foraminal stenosis, highlighting that various clinical factors in addition to grade are required for surgical decision-making^16)^.
Similarly, another retrospective study evaluated 1,248 patients with lumbar spinal stenosis for a mean duration of 7.7 years (range: 5.17-9.8 years) using the Lee et al.^20)^ grading system. The authors determined that surgical probabilities in grade 2 or 3 LFS were 22.2%-62.3% and 33.3%-57.9%, respectively, dependent on concomitant central stenosis, and that LFS of grades 2 and 3 (OR: 2.22 and 2.12, respectively) were significant risk factors for surgical management^17)^. Other investigators determined that Lee at al. grade 3 LFS (OR=2.42) moderately increased the risk of second-stage posterior direct decompression after lateral lumbar interbody fusion procedure^24)^.
Similarly, another study evaluating the radiologic efficiency of percutaneous endoscopic lumbar foraminotomy using the Lee et al.^20)^ grading system showed high interobserver agreement for pre- and postoperative foraminal parameters^9)^. The authors favored the Lee et al.^20)^ grading system because of its superiority to that of Wildermuth et al.^41)^ in evaluating the surgical change of the compressed exiting nerve root after the decompression procedure^9)^. Other investigations evaluated tubular microdiscectomy (TMD) versus transforaminal endoscopic lumbar diskectomy (TELD) to compare clinical outcomes of both techniques for treatment of lumbar radiculopathy in relation to preoperative Lee et al.^20)^ LFS grade^29)^. The authors found that patients with preoperative grade 3 LFS showed significantly higher postoperative VAS scores (p<0.01) and worse functional outcomes than did those with grade 1 or 2 LFS, and the severity of preoperative LFS correlated with clinical outcomes in TELD>TMD, suggesting that preoperative LFS grading is a useful tool^29)^.
One study showed a 5-fold reduction in chance of achieving a 30% improvement in Oswestry Disability Index after microsurgical decompression performed for classically symptomatic lumbar central stenosis without radicular symptoms in a setting of Lee et al.^20)^ grade 2-3 LFS as compared with patients with grade 0-1 LFS at the index OR 0.22 (95% CI 0.06-0.83), (p=0.03)^8)^.
Research has evaluated the utility of repeat MRI performed for pre-surgical evaluation of symptomatic lumbar stenosis (central, lateral, or foraminal)^19)^. The authors determined that repeat MRI for lumbar stenosis due to osseoligamentous degenerative changes without disk herniation of any form (central, lateral, or foraminal) is of low value unless the patient presents with new neurological deficits, especially if the repeat study is performed within one year of the previous MRI. They used the Lee et al.^20)^ grading system to evaluate changes in LFS and found high interobserver reliability (k=0.80)^19)^.
In summary, the Lee et al.^20)^ grading system is the most recognized grading system for LFS owing to its straightforward but comprehensive description with high reproducibility^30)^. However, it is also important to note that there are studies that have shown no association between Lee et al.^20)^ grade LFS and clinical symptoms, disability, and/or other health-related quality of life measures^11,36)^. Moreover, this classification system has received critique because it is based on limited MRI sagittal evaluation of the foramen, making it difficult to evaluate the entire nerve root inherent to the 3D nature of the spine^19,42)^. Some authors have advocated that given advances in spinal care, more up-to-date MRI reporting criteria are needed to describe the surgical anatomy in the neuroforamen relevant to procedures such as minimally invasive endoscopic transforaminal decompression^21)^. Miskin et al.^25)^ reported difficulties differentiating between Lee et al.^20)^ moderate and severe stenosis, and as a result proposed a modified grading system (Table 2)^25)^. Similarly, Sartoretti et al.^30,39)^ stated “the original grading system by Lee et al.^20)^ is no longer suitable for the evaluation of high-resolution images as a far more complex relationship between the nerve root and surrounding structures with the intervertebral foramen can be identified” and thus proposed an updated grading system.
In 2020, Haleem-Botchu et al.^12)^ created a novel CT-based classification for LFS, with the goal of creating a common language and consistency when describing LFS on CT imaging.
Before this, no other objective or qualitative grading system using CT scans for the evaluation of LFS existed^9)^. The authors created a four-point grading system for assessing LFS, providing a viable alternative to the Lee et al.^20)^ MRI-based grading system with near-perfect agreement (k=0.81) (Table 2)^12)^. In addition, intraobserver agreement for CT classification was nearly perfect for both readers (k=0.89, k=0.98), with moderate interobserver agreement (k=0.58)^12)^. The authors concluded that based on their results, the novel CT-based classification can accurately replace the MRI grading system in necessary clinical scenarios^12)^.
In 2021, Miskin et al.^25)^ proposed a simplified multidisciplinary (neurosurgery, orthopedic surgery, and physiatry) grading system for the assessment of spinal stenosis, foraminal stenosis, lateral recess stenosis, and facet arthropathy. For LFS, three multidisciplinary readers achieved moderate agreement (k=0.544). The authors advocated that their grading system has the potential to serve as the basis for interdisciplinary communication of spine findings, including LFS severity^25)^. The authors created a modified grading system based on the Lee et at al. grading system and similarly performed their assessment on T1-weighted sagittal images (Table 2)^25)^. Notably, the authors acknowledge that the interobserver agreement for this LFS grading system was less than that previously reported for the original Lee et al.^20)^ grading system^25)^.
Similarly, in 2021, Sartoretti et al.^30)^ proposed an updated 6-point grading system for LFS based on the Lee et al.^20)^ grading system (Table 2). The authors proposed that using high-resolution 3D MRI techniques (which allows reduction in partial volume effects, thin sagittal slices, and multiplanar reconstruction in submillimeter resolution) as compared with original 2D sagittal images on which the Lee et al.^20)^ grading system was based requires an updated grading system^30,39)^. They proposed an updated grading system to better account for the complex relationship between the nerve root and surrounding structures seen with advanced high-resolution imaging techniques^30,39)^. The authors indicated that their new grading system adds detail and more accurately describes even the smallest anatomical changes in the lumbar foramen, not visible with prior 2D sequences^30,39)^. The updated grading system is based on 3D sagittal high-resolution T2-weighted (T2w) sequence and secondary use of 2D T1-weighted sequence^30,39)^. With these high-resolution images, the entire path of the nerve root within the neuroforamen is visualized in greater detail^30)^. Sartoretti et al.^30)^ extended the Lee et al.^20)^ grading system by two categories as it pertains to severity and provides detail regarding positional information of nerve root contact (Table 2). The authors evaluated 966 lumbar foramina and determined that approximately 30% of cases were graded with categories that were not available in the original Lee et al.^20)^ grading system, and in no case was the updated grading system unable to accurately describe the severity of LFS^30)^. An additional study by the authors has illustrated the reproducibility and reliability of the updated 6-point grading system for LFS for both high-resolution 3D and standard-resolution 2D T2w turbo spin echo imaging^39)^. The authors indicate that clarity with future research is necessary to determine the clinical correlation of this grading system^30)^.
In 2024, Spinnato et al.^31)^ published a new grading system for grading all forms of lumbosacral spinal stenosis including LFS, with a semi-qualitative assessment of the degree of the stenosis. This new grading system aims to not only grade the severity of stenosis but also indicate all factors contributing to the stenosis (Table 2)^31)^. The authors composed a grading system of LFS based on sagittal transverse relaxation times in weighted images (T2-WI) by conceptualizing the neural foramen as a quadrilateral and considering the relationship between the exiting nerve roots and the surrounding four walls (superior wall, pedicle; inferior wall, disk/posterior facets/inferior pedicle; posterior wall, pars interarticularis; anterior wall, vertebral body)^31)^. Severity of stenosis is based on the number of foraminal walls with perineural fat obliteration directly in contact with the nerve root^31)^. The classification was modified for each grade of stenosis dependent on the cause of stenosis, which appeared most applicable to central stenosis^31)^. The authors found almost perfect intraobserver agreement (k=0.928) and interobserver (k=0.834) agreement for the LFS grading system^31)^. The authors also found that their new MRI grading systems for the various causes of lumbosacral stenosis showed a strong correlation with clinical symptoms for central stenosis but were less notable for LFS^31)^.
Investigation of the potential benefit of deep-learning algorithms in improving the accuracy and predictive value of spine MRI studies is particularly relevant given recent advances in technology^22,23)^. In recent years, Lewandrowski et al.^22,23)^ published work on the ability of deep learning neural network models to accurately identify MRI findings representative of variable spine degenerative pathology. The authors advocate that such models may be particularly useful in the setting of minimally invasive endoscopic transforaminal decompression techniques when a tailored focus to region of painful pathology is necessary^22,23)^. Multus RadBot AI deep learning network MRI lumbar spine imaging was compared with experienced radiologist review, with high reliability (k=0.627)^23)^. The authors determined that the Multus RadBot reports on painful pathology were as useful as radiologist reports in the successful treatment of LFS due to a herniated disk using endoscopic decompression procedures^22)^. Notably, Lewandrowski et al.^23)^ incorporated a simplistic two-point grading system (normal anatomy or stenosis present) to assess the AI algorithm's performance, and the authors acknowledge that further refinement is necessary and is likely to evolve with advances in technology. In summation, the AI studies by Lewandrowski et al.^22,23)^ indicate a promising use of deep-learning algorithms to potentially match radiologist-level accuracy in diagnosing stenosis and predicting surgical outcomes with high reliability metrics, with the potential to further reduce intra- and interobserver variability.
This systematic review has several limitations. The scarcity of prospective and comparative studies reduces our ability to draw definitive conclusions regarding the clinical applicability of the analyzed LFS grading systems. Most of the studies were retrospective, with inherent limitations. Moreover, variability in sample sizes and demographic characteristics across studies may affect the generalizability of the findings. The use of different imaging modalities and protocols also introduces heterogeneity, making it challenging to compare the results across studies. Notably, all LFS grading systems are based on static imaging without evaluation of dynamic changes in LFS characteristics that may correlate with symptoms^30)^. However, this falls within the standards of current practice because dynamic cross-sectional imaging is not typically used for the evaluation of LFS. Future research should aim to address these limitations by standardizing study designs, sample characteristics, and imaging protocols to compare the reliability of the various LFS grading systems along with symptom and clinical outcome correlations.
Our review of the available LFS grading systems across all imaging modalities indicates that the grading system of Lee et al.^20)^ is used most frequently to assess LFS within the literature with notable reliability. Several grading systems have been proposed since the Lee et al.^20)^ grading system that show promise; however, further evaluation is required. Awareness of the published LFS grading systems is of clinical importance to ensure that clinicians across subspecialties evaluate imaging from a similar lens to foster clear communication and ultimately to minimize ambiguity in patient care. The relationship between LFS and symptoms remains a clinical-radiologic diagnosis, taking both patient factors and imaging into account. Further research into the various LFS grading systems and clinical correlations, particularly with the recent rise in machine-learning technology, is a notable area of investigation for future studies to build on.
**Conflicts of Interest: **The authors declare that there are no relevant conflicts of interest.
**Author Contributions: **Laura C.M. Ndjonko: (Methods, Investigation, Visualization, Writing―Original Draft, Writing―Review and Editing, Supervision); Nikol N. Kralimarkova: (Methods, Investigation); Yashoswini Chakraborty: (Methods, Investigation); Zayn S. Bajwa: (Methods, Investigation); Jasmine X. Zimmer: (Methods, Investigation); Ayomikun A. Taiwo: (Methods, Investigation); Imani N. Bah: (Methods, Investigation); Sami S. Khan: (Methods); and Eric K. Holder: (Conceptualization, Investigation, Validation, Visualization, Supervision, Writing―Review and Editing, Project Administration).
**IRB Ethics Approval: **Not required for this study.