Authors: Reil Vinard S. Espino (Philippines; rsespino@ust.edu.ph), Consuelo G. Suarez (Philippines), Donald G. Manlapaz (Philippines), Jazzmine Gale S. Flores (Philippines)
Categories: Research Paper, Reliability, validity, hand-held dynamometer
Source: Hong Kong Physiotherapy Journal
Authors: Reil Vinard S. Espino, Consuelo G. Suarez, Donald G. Manlapaz, Jazzmine Gale S. Flores
Accurate assessment of muscle strength is crucial for clinical practice. While traditional methods like manual muscle testing (MMT) are accessible, isokinetic dynamometry (ID) is the gold standard due to its reliability, although it is expensive, space-consuming, and requires extensive training. Hand-held dynamometers (HHDs) have demonstrated a strong correlation with ID values, suggesting good to excellent validity. However, factors such as muscle group, evaluator proficiency, and protocol standardization can influence hand-held dynamometer (HHD) measurements.
This systematic review aimed to evaluate the inter- and intra-rater reliability and validity of HHDs for lower extremity strength assessment in healthy adults and to identify common test protocols.
A comprehensive electronic search was conducted in six databases (PubMed, Medline and CINAHL via Ebsco host, ISI Web of Science, ProQuest, and Science Direct) from January 2017 to May 2023. Studies were included if they assessed asymptomatic participants using HHDs for isometric or concentric contractions of the hip, knee, or ankle and focused on psychometric properties. The QAREL and QUADAS-2 checklists were used to assess reliability and validity, respectively. To complement these 2 checklists, GRADE was used to determine the certainty of evidence. A meta-analysis was conducted to quantify the pooled reliability and validity of HHDs.
Eighteen studies were included. Sixteen investigated HHD reliability, with eight being solely reliability studies. Reliability was operationalized through the intraclass correlation coefficient (ICC). Eight studies received high QAREL scores, indicating strong methodological quality. The remaining studies received low QAREL scores, suggesting methodological weaknesses. These reliability studies revealed moderate to very high correlations for inter-rater and intra-rater reliability, indicating HHDs can be a dependable tool for evaluating lower limb muscle strength. Ten studies investigated the validity of HHD for measuring muscle strength in the lower limbs. Pearson correlation coefficients showed moderate to perfect positive correlations between HHD and ID measurements, suggesting alignment. Four studies provided data for meta-analysis. The pooled estimate for internal consistency for all hip and knee movement strength assessments across studies was high to very high, indicating minimal measurement error and reliable measurements.
HHDs are reliable and valid for assessing lower extremity muscle strength in healthy adults. Their ease of use, affordability, and portability make them a valuable asset for clinical practice. This research was funded by DOST-SEI. The systematic review is registered in PROSPERO (CRD42023399215).
Muscle strength assumes a pivotal role in function and movement.^1^ Its significance extends to facilitating activities integral to daily living and maintaining individual autonomy.^2,3,4^ While muscle strength is acknowledged as a dependable marker of functional capacity across the general adult population, inadequacies in strength are associated with physical limitations.^4,5^ Assessing this parameter is a fundamental aspect of physiotherapists’ responsibilities for these purposes. Muscle strength normative values derived from healthy adults enable clinicians to identify muscular deficiencies and to measure and diagnose neuromuscular dysfunction by comparing the results with those of a healthy peer of the same age and gender. This helps in gauging patients’ advancements objectively and evaluating the efficacy of treatments.^1,2,3,4,5,6^
Numerous instruments have been created to acquire precise measurements of muscle strength. Manual muscle testing (MMT) is the most widely available technique. Despite its clinical convenience and speed, this subjective approach exhibits subpar psychometric characteristics and substantial constraints in tracking strength changes across time.^1,7,8,9,10,11^
Isokinetic dynamometry (ID) represents an approach with robust psychometric characteristics and is acknowledged as the gold standard for assessing muscle strength. Nonetheless, the equipment comes with a substantial cost, demands ample space for installation, and mandates comprehensive user training.^1,12^ A noteworthy middle ground between MMT and ID involves quantitative muscle testing employing a hand-held dynamometer (HHD). Hand-held dynamometers (HHDs) offer a quantifiable assessment of force, and they are known for their user-friendly nature, compact size, and affordability.^13^ This might provide additional support for the broader clinical utilization.^12^
HHDs demonstrate a strong correlation with values obtained through ID, suggesting good to excellent validity for both approaches.^14^ Nonetheless, it is important to recognize that the utilization of HHD is associated with varying sources of measurement error, contingent upon factors such as the specific muscle group being assessed, the proficiency and training of the evaluators, and the level of protocol standardization.^1,15,16,17^
Chamorro and colleagues focused only on absolute reliability and concurrent validity of HHD, aggregating data on the psychometric properties of HHD in assessing hip, knee, and ankle joints.^13^ This systematic review offers a more current review of this subject matter. Thus, there is a need to update the evidence base with newer studies that assess inter- and intra-rater reliability. Furthermore, we aim to achieve the following (a) identify recent evidence concerning the inter- and intra-rater reliability and validity of HHD in the evaluation of lower extremity strength, encompassing the hip, knee, and ankle, among healthy adults, and (b) determine the prevailing test protocols commonly employed with HHD in the assessment of lower extremity strength, encompassing the hip, knee, and ankle, among healthy adults.
A comprehensive electronic search was conducted in six PubMed, Medline via Ebsco host, CINAHL via Ebsco host, ISI Web of Science, ProQuest, and Science Direct from January 2017 to May 2023. We incorporated the subsequent terms and their variations into our healthy adults, handheld dynamometer, HHD, hand-held dynamometer, and muscle strength dynamometer, along with the keywords validity, reliability, and psychometric properties. Wildcards and truncations were employed, and words were combined using Boolean operators like “OR” and “AND” as needed. The search strategy was adjusted slightly for searches conducted in other databases. In addition to the database search, we manually reviewed the reference lists in the identified papers, aiming to locate pertinent authors and journals.
The criteria for including studies in this review (a) Participants who were asymptomatic or in good health, (b) Evaluation of participants using HHD for isometric or concentric contractions involving the hip, knee, or ankle, and (c) Study designs focused on psychometric properties.
The review excluded studies (a) involved participants with lower extremity musculoskeletal disorders such as sprains and strains, and (b) participants who had undergone lower extremity surgeries such as ACL reconstruction, open reduction, and internal fixation pinning for fractures, or hip or knee replacement procedures.
Before commencing the screening process, a calibration exercise was carried out to guarantee the precision and dependability of article selection. This involved each team member independently screening and charting data from a random sample of five papers. Following this, a pilot test was performed on all team members. In this pilot testing session, each team member autonomously screened and charted data from an additional set of five articles. Within the pilot testing process, one team member assumed the role of an evaluator, responsible for assessing the precision and reliability of the screening and charting process. Any disparities or inconsistencies between the reviewers were identified and resolved through discussions that included the evaluator.
During the screening process, upon finalizing the search, the gathered studies were imported into the Covidence software (Melbourne, Australia), where duplicates were subsequently eliminated. Two separate reviewers, R.E. and D.M., screened abstracts and titles within the identified publications from the databases. In cases where the title was ambiguous, the abstract was assessed to provide additional clarification. If the abstract lacked clarity, the paper was obtained and thoroughly reviewed for definitive clarification. Following the initial selection, the reviewers conducted a more in-depth screening of each paper, considering the inclusion and exclusion criteria. Each criterion was evaluated as either met (included) or not met (not included). Additionally, the rationale for exclusion was documented. At every phase of the selection procedure, any discrepancies or differences among the reviewers were addressed through discussions that included an extra reviewer (J.F).
The reviewers devised an electronic data extraction form for this systematic review. We conducted a pilot test of the charting form with all reviewers participating. Once all reviewers were satisfied with its usability, two independent reviewers utilized the tool to chart the evidence, and any disagreements were resolved through discussions involving another reviewer. The extracted data included comprehensive details on authorship, year of publication, participant characteristics (age and gender), the brand of the HHD employed, comparator (ID) when relevant, examined joint and movement, position used for assessment, method of instrument administration, nature of muscle contraction, muscle testing procedure employed (e.g., “make” or “break” test, MVIC, etc.), validity results (type of validity), reliability results (type of reliability).
Two reviewers (R.E. and D.M.) evaluated the quality of each eligible study utilizing the Quality Appraisal of Reliability Studies (QAREL) checklist^18^ for the reliability studies included in the review. This scoring system comprises 11 elements that explore seven (a) the inclusiveness of participants and assessors, (b) assessor impartiality, (c) examination sequence, (d) appropriateness of the time gap, (e) proper application and interpretation, and (f) statistical analysis.^18,19^ Each element in the QAREL checklist allows for one of three categorical “Yes”, “No”, or “Unclear”. A “Yes” response indicates a positive quality aspect of the study, while a “No” response indicates a negative quality aspect.^18^ This tool has a maximum score of 11. A score exceeding 60% was regarded as the threshold for high quality. When a specific item did not apply to a study, its value was excluded from the overall percentage calculation.^18^ The categorization of the level of evidence for each primary outcome was as “strong” when consistent findings from at least one high-quality study were evident, and the combined sample size from eligible studies was ≥100; “moderate”— when consistent findings from at least one high-quality study were evident, and the combined sample size was ≥50; “limited” when findings from at least one high-quality study were present, and the combined sample size ranged from 25 to 49; and “unknown” when findings were inconclusive or the studies exhibited poor methodological quality or the sample size was ≤25.^19,20^
Conversely, validity studies were assessed using the QUADAS-2 framework. QUADAS-2 is specifically crafted to evaluate the quality of primary diagnostic accuracy studies. It should be used as a supplementary tool to assess study quality alongside the standard process of extracting primary data, such as study design and results, for inclusion in the review. It comprises four primary patient recruitment, index test, reference standard, and patient flow. Every domain is evaluated concerning the potential for bias, with the first three also considering applicability concerns. Signalling questions are included to aid in forming a bias risk assessment. These questions highlight study design aspects that might introduce bias and assist reviewers in making bias risk assessments.^21^
Bias risk is categorized as “low,” “high,” or “unclear.” If all signalling questions for a domain receive a “yes” answer, bias risk is considered “low.” If any signalling question is answered “no,” it suggests potential bias. The “unclear” category is appropriate when insufficient data is available for judgment.^21^ Conversely, the applicability sections are organized similarly to the bias sections but do not incorporate signalling questions. Review authors must document the basis for their applicability judgment and assess the extent to which they are concerned about the study’s alignment with the review question. Concerns about applicability can be rated as “low,” “high,” or “unclear.” The “unclear” rating should only be applied when there is insufficient data to make an assessment.^21^
Two reviewers independently conducted all critical appraisals, and their results were subsequently compared. A third reviewer (J.F.) was on standby to mediate and address any unresolved disputes following a discussion between the two initial reviewers.
Cohen’s Kappa correlation coefficient (κ) between the two reviewers was 0.81 (SE = 0.029, 95% CI= 0.75–0.86), showing “substantial” agreement on the risk of bias assessment.^22^ Figure 1 visually represents this process, following the preferred reporting items for systematic reviews and meta-analyses flowchart.

Following a search for relevant studies, investigations pertaining to homogenous outcome measures were grouped and subjected to a heterogeneity assessment. Subsequently, a meta-analysis was conducted to quantify the pooled reliability and validity of HHD for muscle force assessment in the hip, knee, and ankle joints.
Reliability was operationalized through the intraclass correlation coefficient (ICC). The ICC analysis employed a two-way random effects model with absolute agreement for repeated measures. Munro’s classification scheme (2005)^23^ was adopted to interpret the level of agreement between the devices, where ICC values with slight, low, moderate, high, and very high correlation with the following 0.0 to 0.25, 0.26 to 0.49, 0.50 to 0.69, 0.70 to 0.89, and 0.90 to 1.0, respectively.
We used the Pearson correlation coefficient to assess how well the measurement tool reflects the intended concept (validity). The strength of the relationship was then interpreted using Hopkins’ extension of Cohen’s guidelines. According to this scale, a correlation of 0.00–0.09 Indicates no relationship between the measures; 0.10–0.29 suggests a small positive association; 0.30–0.49 represents a moderate positive association, 0.50–0.69 reflects a large positive association, 0.70–0.89 indicates a very large positive association, 0.90–0.99 suggests a nearly perfect positive association, and 1.00 represents a perfect positive association.^24^
MedCalc^TM^ software was employed to pool effect sizes and generate forest plots for all comparisons. The analysis adopted a 95% confidence interval (CI) for effect size estimates. Heterogeneity within the study population or intervention effect sizes was assessed using MedCalc^TM^. In the event of statistically significant heterogeneity, a random-effects model was implemented for pooling via MedCalc^TM^. This approach acknowledges the assumption of inter-study variability inherent in the random-effects model and consequently yields more conservative estimates of the true effect size. Finally, to determine the certainty of evidence, the modified GRADE approach was used to complement the QUADAS-2 and QAREL quality assessment.
The comprehensive search from six databases yielded 5728 articles using the specified search terms. Among these, 467 were identified as duplicates. A manual cross-reference search did not yield any additional studies. The remaining 5261 papers underwent initial screening based on titles and abstracts, excluding 5228 articles that did not meet the initial criteria. Subsequently, the eligibility criteria were applied to the remaining 32 articles. Among these, 14 papers were automatically excluded due to various reasons, including seven studies that used different interventions or equipment, three studies that had various outcomes, one study that had a different comparator, one study that had a study population that was not part of the inclusion criteria, one that had a different study design, and one that has no full text available. Ultimately, 18 articles satisfied the criteria and were included in the systematic review. Four of the 18 studies that satisfied the criteria and were included in the systematic review that provided enough data for inclusion in the meta-analysis.^25,26,27,28^
Sixteen studies investigated the reliability of HHD for measuring lower limb muscle strength; among these, eight papers were solely reliability studies,^29,30,31,32,33,34,35,36^ while the remaining papers analysed both the reliability and validity of HHD.^25,26,27,28,37,38,39,40^ These studies considered various factors that could influence the measurements, including the specific type of HHD device used, the particular joint and movement being assessed, and the type of muscle contraction performed (e.g., isometric, concentric). Table 1 provides a detailed breakdown of these factors across the sixteen studies.
Ten studies investigated the validity of HHD for measuring muscle strength in the lower limbs, of which only two papers solely evaluated the validity of HHD compared to ID,^41,42^ and the rest analysed both the reliability and validity of HHD.^25,26,27,28,37,38,39,40^ Table 1 details the characteristics of these studies, including the specific joint and movement assessed, the type of HHD device used, the comparison method used, the body position during the assessment, and the type of muscle contraction performed.
Table 2 summarizes the methodological quality assessment of the sixteen reliability studies included in this analysis. The QAREL tool was used to evaluate each study. Eight studies received high QAREL scores, indicating strong methodological quality.^26,30,34,36,37,38,40^ The remaining studies received low QAREL scores, suggesting potential weaknesses in their methodology.^1,25,27,28,31,32,33,39^ Figure 2 shows which items from the QAREL checklist were well-supported by the reliability studies we reviewed and which items were less explored. Most of the included studies addressed key criteria, including subject demographics, time intervals between repeated measures, test application, and interpretation, and statistical analysis methods. In contrast, the checklist revealed that the spectrum of examiners, examiner blinding, and order of examination were the least frequently explored aspects among the included articles.

Table 3 summarizes the QUADAS-2 risk of bias and applicability assessments for the included validity studies. All studies exhibited some risk of bias based on the QUADAS-2 criteria. Eight out of the ten studies had a high risk of bias for the index test.
To complement the QUADAS rating, a modified GRADE approach was used to determine the certainty of evidence. Following this approach, which incorporates both upgrading and downgrading factors, the 18 observational studies initially classified as low certainty were further evaluated based on risk of bias, inconsistency, imprecision, and indirectness. These factors informed whether the certainty of evidence should be upgraded or downgraded.^43,44^ Table 4 summarizes the final certainty ratings and the specific reasons for each adjustment. For the reliability studies, most were rated as moderate (4/9) while the rest remained low (5/9). For validity, both studies achieved very low (2/2). For studies with validity and reliability, the overall certainty of evidence reflected a distribution of moderate (4/7) and low (3/7), providing a balanced overview of the current evidence.
The results of these reliability studies, presented in Table 5, revealed moderate (r=0.50 to 0.69) to very high (r=0.90 to 1.0) correlations for inter-rater reliability and high (r=0.70 to 0.89) to very high correlations (r=0.90 to 1.0) for intra-rater reliability. This suggests that HHD can be a dependable tool for evaluating lower limb muscle strength.
Table 6 presents the results of these validity studies using Pearson correlation coefficients. These results showed moderate (r=0.3 to 0.49) to perfect positive (r=1.00) correlations, suggesting that HHD measurements generally align with those from Isokinetic Dynamometry.
Table 7 presents a stratification of the evidence level associated with this investigation’s primary outcomes. The primary outcomes encompassed movement assessment across all hip, knee, and ankle joints. Studies evaluating the hip and knee segments yielded a strong level of evidence, whereas investigations focused on the ankle segment yielded a moderate level of evidence.
Four studies were included in the meta-analysis, as illustrated in Fig. 1. All studies employed the ICC and standard error of measurement (SEM) to assess the reliability of HHDs. Due to heterogeneity in SEM units across the included studies, a reliability generalization technique^45^ was implemented to synthesize the findings. Assessment of validity used Pearson’s correlation coefficient.
Four studies examined the reliability of the La Fayette HHD for measuring hip and knee muscle strength. Table 8 summarizes these studies, detailing joint and movement assessed (e.g., hip flexion, knee extension), body position during assessment (e.g., sitting, standing), measurement reliability statistics [(SEM), 95% CI, Random and total effect sizes)]. These statistics help determine how consistent HHD measurements are across different studies.
Forest plots summarizing the results from multiple studies for various hip and knee movements are in Figs. 3 and 4. Importantly, none of the lines in these plots cross zero, indicating a consistent effect across all studies. The pooled estimate for internal consistency for all hip movement strength assessments across studies was high (ICC = 0.86, 95% CI: 0.76 to 0.96). Similarly, the pooled ICC for knee extension strength was very high (ICC = 0.94, 95% CI: 0.76 to 0.96). These results suggest minimal measurement error and reliable measurements for specific hip and knee movements using HHD in the same position.




Performing subgroup analysis for each specific movement of the hip, HHD demonstrated high to very high reliability for measuring strength in most movements, including hip flexion (ICC = 0.94, 95% CI: 0.71 to 1.17), extension (ICC = 0.88, 95% CI: 0.65 to 1.11), abduction (ICC = 0.87, 95% CI: 0.66 to 1.07), ER (ICC = 0.86, 95% CI: 0.63 to 1.09) and knee extension (ICC = 0.94, 95% CI: 0.71 to 1.17) (all in supine or sitting positions). Hip internal rotation (ICC = 0.74, 95% CI: 0.51 to 0.96) showed slightly lower, but still moderate reliability when measured with HHD in a sitting position. Overall, this suggests HHD can be a dependable tool for assessing strength in these lower limb movements. For a more detailed breakdown of reliability for different hip movements and positions, see Figs. 7–11.





Two studies^25,26^ investigated the validity of the La Fayette HHD for measuring hip and knee muscle strength, comparing it with isokinetic dynamometer. Table 9 summarizes these studies, including details about the specific joint and movement assessed (e.g., hip flexion, knee extension), the body position during the assessment (e.g., sitting, standing), measurement validity statistics (correlation coefficients and 95% Confidence Interval, random and total effect sizes).
Figures 5 and 6 show the forest plots for validity results for hip flexion and knee extension movements across multiple studies. Analysis of validity for hip flexion was 0.92 (95% CI: 0.88 to 0.95) and knee extension was 0.94 (95% CI: 0.88 to 0.95). results showed a very large positive correlation between HHD and ID.
This systematic review highlighted an up-to-date assessment of the psychometric properties of the HHD for measuring lower extremity muscle strength. The findings demonstrated that the HHD exhibits high levels of reliability and validity when assessing the hip, knee, and ankle muscle strength of asymptomatic (healthy) individuals. Meta-analysis demonstrated a large positive correlation between HHD and ID in different hip movements and knee extension.
Our systematic review and meta-analysis diverge from the 2017 publication^13^ by concentrating on specific psychometric properties of HHD. While the previous paper may have addressed absolute reliability, our current study delves into inter-rater and intra-rater reliability, assessing the consistency of measurements across different assessors and within the same assessor. Chamorro et al.^13^ investigated the absolute reliability of both HHD and ID, quantifying the degree of variation within an individual’s repeated measurements.^44,45^ Their study measured the degree of consistency or how close a measurement is to a true or known value. The precision of absolute reliability can be shown either directly in the measurement units or as a relative comparison of the measurements.^13^ Their analysis employed the SEM or Limits of Agreement (LOA) to evaluate this aspect. HHD results were expressed as kilograms (kg) weight measurements and percentages. For example, a study in Chamorro et al.’s^13^ review assessed the consistency of measurements from Lafayette and Hoggan Health HHDs at the hip, knee, and ankle joints.^46^ This study also evaluated Newton-meter torque measurements’ consistency and ID percentages at the same joints. The study’s findings indicate that HHD exhibits lower reliability for assessing knee extension and ankle plantar flexion than other movements.
While ID demonstrate superior reliability for hip, knee, and ankle movements, their performance is not without limitations. Additionally, the study revealed varying correlations between different muscle groups.^13^ In contrast, our systematic review and meta-analysis concentrated on HHD’s inter-rater and intra-rater reliability. These metrics assess the extent of variation between testers (inter-rater) and the consistency of a single tester across measurements (intra-rater). The ICC was the primary tool for analysing this type of reliability in our investigation. Our study used the ICC values for inter and intra-rater reliability results from the included studies. We evaluated how reliable the assessors were in terms of using the HHD in measuring lower extremity muscle strength. Our study found that the HHD is a reliable tool for measuring muscle strength in the hip, knee, and ankle of healthy people.
Like Chamorro et al.,^13^ we assessed the concurrent validity of HHD with ID, evaluating their correlation with the established gold standard for muscle strength measurement. Concurrent validity assesses how well a new measurement tool (HHD) correlates with an established gold standard (ID).^13,44,45^ Interestingly, both our review and Chamorro et al.’s found moderate to very large positive correlations between HHD and ID.
Strength assessment results are influenced by individual body position.^13^ Our review highlights the conventional positions used in muscle strength assessments. Standardizing techniques significantly improve inter-rater and intra-rater reliability for HHD assessment in healthy populations.^47,48^ Our study uniquely assessed optimal measurement positions for hip and knee movements to further refine HHD usage. We aimed to identify body postures that enhance reliability and validity. By comprehensively examining HHD reliability and validity, our study offers techniques to improve the accuracy of future HHD research. For example, the supine position was commonly used for hip flexion^25,26^ and hip extension.^26,27^ When restricting analysis to studies using identical equipment and positions, we observed high homogeneity (I2 = 0.00%) and strong reliability coefficients for pooled data on hip and knee movements. A meta-analysis of studies using the same HHD brand (Lafayette HHD) revealed strong reliability for hip flexion (ICC = 0.94) and hip extension (ICC = 0.88) in the supine position. Similar findings were observed using Lafayette HHD for hip abduction, external rotation, and internal rotation. The values for these specific movements were mentioned in the previous section.
Inconsistent positions can compromise HHD accuracy. Comparing our study to Chamorro et al.,^13^ we found different HHD positions, especially for the hip joint. For instance, hip flexion measurements varied due to different positions (standing, lying, sitting) compared to the consistent prone position for hip extension. Interestingly, studies included in our systematic review but not in the meta-analysis demonstrated consistent inter-rater and intra-rater reliability despite variations in assessment positions. Morin et al.^29^ utilized a standing position for hip flexion, while Florencio et al.^33^ opted for a seated position. Similarly, hip internal and external rotation assessments varied [prone in Lipovšek et al.’s]^25^ versus [supine in Jackson et al.].^34^ Despite this heterogeneity in positioning, all aforementioned studies reported strong reliability in their measurements.
For knee flexion and extension, Chamorro et al.^13^ and our study utilized a seated position, but the results differed. Chamorro et al.^13^ reported a moderate correlation for knee extension, while our study revealed a very high correlation. One potential explanation for the differing results could be the use of different HHD brands. Our meta-analysis, focusing on the same movement assessment position and a consistent HHD brand, yielded a higher correlation result.
This observation suggests that standardized assessment techniques, independent of specific body positioning, may be paramount for achieving high inter-rater or intra-rater reliability in healthy populations. Morin et al.’s^1^ scoping review revealed a dearth of standardized protocols for assessing muscle strength utilizing HHD. Despite this, the authors identified several standardized techniques that could ensure high reliability and validity when employing HHDs for muscle strength measurement. Most included studies employed a “make” test protocol in gravity-neutralized positions.^49,50,51,52,53,54,55^ Isometric contraction durations ranged from 3 s to 7 s, with rest periods varying from 10 s to 2 min between trials.^49,50,51,52,53,54,55,56,57^ Most studies conducted between 1 and 5 trials per muscle group, with verbal encouragement provided during testing.^54,55,57^ The included studies also provided sufficient methodological detail to enable protocol replication, outlining participant positioning, limb and joint placement, dynamometer positioning, and segment stabilization.
Age and gender were extracted because both variables are commonly cited as potential sources of variability in lower-extremity strength assessment using HHD. Although subgroup analysis was not possible due to limited available data, the broader evidence base helps clarify the extent to which these variables influence measurement reliability and validity.
The findings indicate that age does not compromise the reliability of HHD measurements. Prior studies involving older adults (65–92 years) consistently report high intrarater and interrater reliability (ICC ≈ 0.90) for hip and knee make-tests, demonstrating that the capacity of raters and devices to produce stable, repeatable measurements is preserved across wide age ranges.^59^ The present results align with this pattern, suggesting that the psychometric properties of the HHD — particularly its resistance to measurement error — are robust regardless of participant age. However, age exerts a substantial influence on measurement validity. Age-related declines in muscle force production are well established, and large normative datasets (20–97 years) show that age explains 12.9–25.3% of the variance in strength values among men and 20.8–24.6% among women.^60^ These trends indicate that while the device measures force consistently (high reliability), the meaning of the strength value changes with age. Thus, without age-specific reference values, the absolute strength generated by older adults may be misinterpreted as pathological weakness when it actually reflects normal age-related decline.
Gender shows a similar distinction between reliability and validity. Evidence indicates that the reliability of HHD measurements is not affected by gender, as both males and females exhibit comparable intrarater (ICC > 0.90) and interrater (ICC > 0.80) reliability across major lower-extremity muscle groups. Studies involving mixed-gender clinical samples also report excellent interrater reliability (ICC ≥ 0.91), reinforcing that the measurement process itself does not differ by sex.^61^
In contrast, gender substantially affects validity because men consistently produce higher absolute strength values than women across all age brackets. These differences reflect physiological variations in muscle mass and force-generating capacity rather than measurement error. Consequently, interpreting absolute HHD values without sex-specific normative data risks misclassifying typical female strength levels as deficits. This mirrors findings in grip-strength literature, where sex is frequently identified as a stronger determinant of force output than age, body mass, or height.^62^
Overall, the evidence demonstrates that while age and gender do not influence the reliability of make-test measurements obtained via HHD, both variables meaningfully affect the validity of the strength values. For this reason, accurate interpretation requires the use of age- and sex-adjusted reference standards to distinguish true impairments from normal demographic variation in muscular strength.
A potential source of bias is the exclusion of publications that do not have accessible full text. Efforts were not made to locate unpublished studies and doctoral theses in this domain.
This paper provides an update to the previously published systematic review and meta-analysis on the absolute reliability and concurrent validity of hand-held and isokinetic dynamometry in the hip, knee, and ankle joints. Our review followed the PRISMA guidelines, ensuring that all included studies had undergone peer review. Nonetheless, certain limitations must be acknowledged. Despite efforts to incorporate a wide range of published articles and relevant grey literature, some pertinent sources may not have been captured. Furthermore, the search was restricted to studies published in the English language, which may have limited the scope of the review.
In this systematic review, we evaluated the reliability and validity of HHD for assessing lower extremity strength in healthy adults, with a specific focus on the hip, knee, and ankle. We systematically gathered recent evidence concerning the reliability and validity of HHD for this purpose. Furthermore, we identified and analysed the typical test protocols commonly utilized with HHD when assessing lower extremity strength in healthy adults, covering the hip, knee, and ankle. The results indicated that the HHD shows adequate reliability and validity when evaluating muscle strength in the hip, knee, and ankle of healthy individuals.
Our review also highlights the overall quality and strength of the evidence found in the included studies. While half demonstrated strong methodological quality based on QAREL criteria, the evidence level varied depending on the body segment. Studies examining the hip and knee yielded strong evidence, suggesting high certainty in their findings. However, the evidence for the ankle segment was rated as moderate, indicating somewhat less certainty.
This variation underscores the importance of considering both methodological quality and potential biases. Using the QUADAS-2 tool, we identified that all studies exhibited some degree of bias, and some might not be directly applicable to our specific population or research question. These limitations should be considered when interpreting the overall findings.
This systematic review is registered in PROSPERO (CRD42023399215) (https://www.crd.york.ac.uk/prospero/#myprospero). The presentation of findings in this systematic review adheres to the guidelines outlined in the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA).^58^ These guidelines encompass a checklist comprising 27 items and a flow diagram organized into four phases.
We conducted searches using both Medical Subject Heading (MeSH) terms and free-text keywords. Here is the syntax used for searching the PubMed/Medline (((((((((healthy adults[MeSH Terms])) AND (hand held dynamometer[MeSH Terms])) OR (HHD[MeSH Terms])) OR (hand-held dynamometer[MeSH Terms])) OR (muscle strength dynamometer[MeSH Terms])) AND (validity and reliability[MeSH Terms])) OR (”psychometrics”[MeSH Terms])) OR (sensitivity[MeSH Terms])) OR (specificity[MeSH Terms]).
The authors declare no conflict of interest relevant to this article.
This research is funded by DOST-SEI.
R.V.E, C.G.S, D.G.M, J.G.F – Conception and Design of the Study; Acquisition of Data; Analysis and/or Interpretation of Data; Drafting and Revising the Manuscript; Approval of the manuscript to be published.
Reil Vinard S. Espino https://orcid.org/0000-0002-9695-1614
Consuelo G. Suarez https://orcid.org/0000-0001-8382-474X
Donald G. Manlapaz https://orcid.org/0000-0002-1041-2303
Jazzmine Gale S. Flores https://orcid.org/0009-0009-9010-8894