Authors: Chittaranjan Andrade (Department of Clinical Psychopharmacology and Neurotoxicology, National Institute of Mental Health and Neurosciences, Bangalore, Karnataka, India)
Categories: Scholar's Corner, Inter-rater reliability, intraclass correlation coefficient, reliability and validity, test-retest reliability
Source: Indian Journal of Psychiatry
Authors: Chittaranjan Andrade
The intraclass correlation coefficient (ICC) is a measure of reliability that is available in many forms, classified as one-way or two-way, random or mixed model, for single measures or as a mean of k measures, and as values for consistency or absolute agreement. Different forms of the ICC are appropriate in different clinical research contexts; each form has its own formula and interpretation. Literature on the ICC is sparse, and so concepts related to this statistic tend to be poorly understood, possibly resulting in a wrong form being used, and commonly resulting in poor reporting of the statistic in research papers. This article therefore explains the ICC and its forms in simple language, with examples where required. Practical guidance is provided for use and interpretation of the ICC. For most studies of test-retest reliability, or inter-rater reliability, the two-way mixed model ICC for single measures is appropriate, and the absolute agreement, along with its 95% confidence interval, should be reported.
In empirical research, we study variables and relationships between these variables. To do this well, we need to measure the variables using instruments that are valid and reliable. Reliability essentially means replicability; that is, repeated use of the instrument should yield the same results; if not, closely similar results. There are different types of reliability [Box 1], among which test-retest reliability and inter-rater reliability are the most reported [Box 2].
For test-retest reliability as well as for inter-rater reliability, the outcome measured can be categorical or continuous [Box 3]. Two points are worthy of note. The first point is that, for both test-retest reliability and inter-rater reliability, it is the same subjects who are being assessed more than once, using the same instrument. This is why the same statistical tests are applied, whether the data come from a test-retest reliability study or an inter-rater reliability exercise. The second point is that these statistical tests can be used even when test-retest reliability is studied across more than two occasions, or when inter-rater reliability is studied for more than two raters.
The statistical tests referred to are the kappa coefficient and the intraclass correlation coefficient (ICC). The former is used when the outcome is categorical, and the latter is used when the outcome is continuous [Box 3].
The ICC is often poorly understood and poorly reported. When poorly reported, it prompts suspicion that it may have been incorrectly computed. This could be because although there are as many as 10 different forms of the ICC,[1] there is little explanation of which form is appropriate for which context. This article therefore explains the ICC with special reference to that forms need to be used in the most encountered situations in clinical research. The explanations provided in this article are detailed yet simple; this is necessary because most of the articles on the subject are technical and, therefore, might discourage the average reader. Those who want more technical discussions can consult the articles by Koo and Li,[1] Trevethan,[2] and Ten Hove et al.[3]
Some readers may Why a correlation coefficient, the ICC, with 10 forms? Why not the r value; that is, the conventional Pearson’s product-moment correlation coefficient? The reasons are explained using the data in Table 1.
In the Table, Column A displays depression rating scores for 10 subjects who had been rated by an expert; so, the scores in Column A can be considered to be gold standard ratings. The ratings in Column B are identical with those in Column A; it is unsurprising that r = 1.00, indicating perfect correlation. Each rating in Column C differs by the same value, +3, from the corresponding rating in Column A; but, although the Column C ratings are obviously systematically different from the gold standard, the correlation is still perfect (r = 1.00). In Column D, as compared with the gold standard, the ratings diverge systematically as scores increase; nevertheless, the correlation remains perfect (r = 1.00).
Some readers may now why not apply a paired t test to demonstrate that the scores in Columns C and D are not the same as the gold standard? There are two reasons. One is that the ratings in Column E are obviously different from the ratings in Column A; yet, a paired t test will show nil difference because the means of Columns A and E are identical (in this dataset, the errors in higher vs lower ratings cancel out exactly). So, a paired t test cannot be trusted to reveal errors in ratings. The more important reason, however, is that the paired t test examines differences in ratings. In reliability, we want to know the agreement between ratings. How meaningful the mean difference is (paired t test) bears little relevance to the question of how similar the paired ratings are (ICC).
The ICC is classified in four ways; these classifications are independent of each other, which is why, when the subtypes are combined in different ways, we can obtain up to 10 different forms of the ICC.[1] The classifications are one-way vs two-way, random vs mixed, single measures vs mean of k measures, and consistency vs absolute agreement. Understanding these subtypes is important because different forms of the ICC are used in different contexts, and because different forms give different results when applied to the same set of data. Each of these classifications is explained below.
The one-way ICC (specifically, the one-way random effects model) tells us how similar two measurements of the same subject are when these measurements are made by two random raters from a larger pool of potential raters. This form of the ICC is appropriate when the same raters do not rate all subjects. As an example, different nurses working in different shifts may record blood pressure of the same patients at different times on different days; how good is the agreement in their recordings? Expressed otherwise, if we select two blood pressure recordings of the same patient, made by two randomly selected nurses, how similar are they? This model therefore presents reliability of rating under heterogeneous rating conditions.
It is called “one-way” because only subjects are modelled (within-subjects variance and between-subjects variance), not raters. The rater variance is included in the error term. The use of this form of the ICC is uncommon in clinical research contexts.
The two-way ICC models rater variance as well as subject variance (hence “two-way”), and so gives us an idea about how the raters perform. This form of the ICC is applied when the same raters assess all subjects.
The two-way random effects ICC examines reliability when the raters are assumed to be a random subset of raters in the population; so, the findings can be generalized to other raters in the population. As an example, if there is good agreement in how a randomly selected group of doctors score a neurological test in a random sample of patients, the two-way random effects ICC findings can be generalized to other doctors in the population. Expressed otherwise, if these doctors can reliably score the test, all doctors can reliably score it.
The two-way mixed effects ICC examines reliability in a random sample of subjects assessed by a fixed group of raters; because random effects and fixed effects are combined in the model, it is called a mixed model. The results of this form of ICC can hence be generalized to other subjects in other samples, but cannot be generalized to other raters. We conclude that these raters, specifically, will do as well in other samples as they do in this sample.
The two-way mixed effects ICC is by far the commonest model used to assess test-retest reliability and inter-rater reliability in clinical research. Readers may note that the formulae for the two-way random and mixed forms of the ICC are identical[1]; so, the difference in use between these two forms of the ICC lies in how the study is designed and how the results are interpreted.
Whether one-way or two-way, whether random or mixed, the ICC form is different for single measures vs mean of k measures.
Single measures are also described as single measurements or single raters; these terms all mean that a single rating is obtained and this single value is recorded as the assessed value. As an example, blood pressure is measured once, and that reading is recorded as the assessed value.
Mean of k measures, also described as mean of k measurements and mean of k raters, means that the rating is obtained k times, or by k raters, where k represents the number of repetitions or the number of raters. As an example, blood pressure is measured thrice, or by three different people, and the average of the three measurements is recorded as the assessed value.
For most assessments in clinical research, measurements are conducted once; there is no averaging. So, the ICC for single measures is the commonest form that is used.
As a final classification, the ICC can be obtained as a value for consistency and as a value for absolute agreement.
The ICC value for consistency tells us how well the ranking of measurements is preserved across occasions of measurement. This, therefore, is like but not identical with the Spearman’s rank order correlation; the difference is that the ICC for consistency also tells us how well subjects are distinguished within the ranking order. This value of the ICC is important only when systematic variation across occasions of measurement is expected, as when there is a learning effect on a task (e.g., repeated administration of the same cognitive test, such as serial subtraction).
The ICC value for absolute agreement tells us how similar the values are across occasions of assessment or across different raters. From the discussion of Table 1 in an earlier section of this article, it should be obvious that absolute agreement is the form of the ICC that is most often required.
For both test-retest reliability and inter-rater reliability (as well as for reliability of parallel forms of a test), researchers will usually want to examine absolute agreement using the two-way mixed effects model of the ICC for single measures. This is true even if the number of occasions of testing, the number of raters, and the number of parallel forms is more than two. The single measures specification does not mean that the ICC is computed for only two raters, or only two occasions; nor does it mean that the ICC was computed for pairwise comparisons. Rather, the single measures specification means that the ICC describes the reliability of any rating by any rater on any occasion, based on the data analyzed.
If reliability is being assessed across several points in time, or if reliability is assessed among multiple raters, and if the outcome of interest is the averaged ratings, then the two-way mixed effects model of the ICC with average of k measures can be used. This form of the ICC tells us about the reliability of the averaged ratings, and not about the reliability of any single occasion of rating, or the reliability of any single rater. This form of the ICC should be used only if the instrument is intended to be scored as an average of k ratings. Readers may note that this is not a usual situation.
The ICC range from negative values to +1.00. An ICC of 1.00 indicates perfect agreement.
A negative ICC is uncommon; if obtained, it indicates that the agreement is worse than could have been expected by chance (for those who are mathematically minded, this happens when within subject variance is greater than between subjects variance. Negative ICCs are sometimes interpreted as though ICC = 0.00. An ICC of 0.00 indicates that there is no agreement.
As a rule of thumb, Koo and Li[1] suggest that an ICC < 0.50 suggests poor reliability; an ICC between 0.50 and 0.75 suggests moderate reliability; an ICC between 0.75 and 0.90 suggests good reliability; and an ICC > 0.90 suggests excellent reliability.
As minor points, for a particular set of scores, the ICC for the mean of k measures is usually higher than the ICC for single measures, and the ICC for consistency is higher than the ICC for absolute agreement. As major points, the ICC can be misleadingly low if the sample size is small or if the sample is homogenous; that is, there is low variation in scores between the subjects in the sample.
Under Statistical Methods in the Material and Methods section of their paper, researchers should describe the complete form of the ICC used. For example, they might state, “We examined test-retest reliability as the absolute agreement (along with 95% confidence interval) between ratings, using the two-way, mixed effects model of the intraclass correlation coefficient for single measures”.
In the Results section, the authors can state, for example, “ICC (absolute agreement) = 0.33; 95% CI, 0.00 −0.78.” The ICC for consistency is usually not reported; it conveys little useful information unless, for any reason, researchers are interested in fidelity of ranking as well as fidelity of absolute scores. As explained in an earlier section, this can happen when there is systematic change in scores across occasions of assessment, for example due to a learning effect.
A sample size of at least 30 subjects is recommended for estimation of the ICC, such as in inter-rater reliability exercises. However, when the ICC is the primary purpose of a study, as in estimation of test-retest reliability during development of a new instrument, sample size needs to be more formally estimated. Parameters that need to be set are the ICC model, the expected value of the ICC, its precision (95% confidence interval), the number of raters (or rating occasions), the power (commonly, 80%), and the value for statistical significance (commonly, 0.05%). A free online calculator for estimating sample size for the ICC is available https://wnarifin.github.io/ssc/ssicc.html (accessed on January 04, 2026).
In a recent study describing the Hindi adaptation and psychometric validation of the Affiliate Stigma Scale, in the context of test-retest reliability, Kumari et al.[4] wrote, “the single measure ICC value was 0.40, with a 95% confidence interval 0.30–0.50.” There was no further explanation. We do not know whether the authors used a one-way or two-way model, whether they were referring to random or mixed effects, and whether the ICC reported was for consistency or absolute agreement. Furthermore, they also reported the use of Pearson’s r as a measure of test–retest reliability.
In another study, Chacko et al.[5] used the ICC to assess inter-rater reliability of an instrument that assessed psychiatric disability in patients with traumatic brain injury; no explanation whatsoever was provided to understand how they used the ICC. Similar limitations were observed in other articles.[67]
It is hoped that the explanations about the ICC, provided in this article, will promote the correct use and reporting of this statistic.
No artificial intelligence tools were used in the preparation of this article.
There are no conflicts of interest.