Several researchers have shown interest in having R functions that can compute several chance-corrected agreement coefficients, their standard errors, confidence interval, and p-values as described in my book Handbook and Inter-Rater Reliability (3rd ed.). I have finally found the time to write these R functions, which can be downloaded from this r-functions page of the agreestat website.
All these R functions can handle missing values without problems, and cover several types of agreement coefficients including Gwet AC1/AC2 (2008, 2012), Kappa coefficients of Cohen (1960), Fleiss (1971), Conger (1980), Brennan & Prediger (1981), Krippendorff (1970), and the percent agreement.
Bibliography:
[1] Brennan, R. L., and Prediger, D. J. (1981). "Coefficient Kappa: some uses, misuses, and alternatives." Educational and Psychological Measurement, 41, 687-699.
[2] Cohen, J. (1960). "A coefficient of agreement for nominal scales." Educational and Psychological Measurement, 20, 37-46.
[3] Conger, A. J. (1980), "Integration and Generalization of Kappas for Multiple Raters," Psychological Bulletin, 88, 322-328.
[4] Fleiss, J. L. (1971). "Measuring nominal scale agreement among many raters", Psychological Bulletin, 76, 378-382
[5] Gwet, K. L. (2008). "Computing inter-rater reliability and its variance in the presence of high agreement." British Journal of Mathematical and Statistical Psychology, 61, 29-48.
[6] Gwet, K.L. (2012). Handbook of Inter-Rater Reliability (3rd Ed.), Advanced Analytics, LLC, Maryland, USA
[7] Krippendorff, K. (1970). "Estimating the reliability, systematic error, and random error of interval data," Educational and Psychological Measurement, 30, 61-70
On this blog, I discuss about some techniques and general issues related to the design and analysis of inter-rater reliability studies. My mission is to help researchers improve how they address inter-rater reliability assessments through the learning of simple and specific statistical techniques that the community of statisticians has left us to discover on our own.
Monday, March 31, 2014
Saturday, March 8, 2014
The Perreault-Leigh Agreement Coefficient is Problematic
Perreault and Leigh (1989) considering that there was a need to have an agreement coefficient "that is more appropriate to the type of data typically encountered in marketing contexts," decided to propose a new agreement coefficient known in the literature with the symbol Ir. I firmly believe that the mathematical derivations that led to this coefficient were wrong, even though the underlying ideas are right. A proper translation of these ideas would inevitably have led to the Brennan-Prediger coefficient or to the percent agreement depending on the assumption made.
The Perreault-Leigh agreement coefficient is formally defined as follows:
where S defined by,
is the agreement coefficient recommended by Bennet et al. (1954) and is a special case of the coefficient recommended by Brennan and Prediger (1981). The symbol Ir used by Perreault and Leigh (1989) appears to stand for “Index of Reliability.”
I carefully reviewed the Perreault and Leigh article. It presents an excellent review of the various agreement coefficients that were current at the time it was written. Perreault and Leigh define Ir as the percent of subjects that a typical judge could code consistent given the nature of the observations. Note that Ir is an attribute of the typical judge, and therefore does not represent any aspect of agreement among the judges. Perreault and Leigh (1989) consider the product NxIr2 (with N representing the number of subjects) to represent the number of reliable judgments on which judges agree. This cannot be true. To see this note that Ir2 is the probability that 2 judges both independently perform a reliable judgment. If both (reliable) judgments must lead to an agreement then they have to refer to the exact same category. However the probability Ir2 does not say which category was chosen and cannot represent any agreement among judges. Even if you decide to assume that any 2 reliable judgments must necessarily result in an agreement, then the judgments will no longer be independent. The probability for two judges to agree will now become equal to the probability for the first rater to perform a reliable judgment times the conditional probability for the second judge to perform a reliability judgment given that the first judge did. This second conditional probability cannot be evaluated unless there are additional assumptions.
What Perreault and Leigh (1989) have proposed is not an agreement coefficient. Their coefficient quantifies something other than the extent of agreement among raters. It should not be compared with the other coefficients available in the literature until someone can tell us what it does.
References:
[1] Bennett, E. M., Alpert, R. & Goldstein, A. C. (1954). Communication through limited response questioning. Public Opinion Quarterly, 18, 303-308.
[2] Brennan, R. L., & Prediger, D. J. (1981). Coefficient kappa: Some uses, misuses, and alternatives. Educational and Psychological Measurement, 41, 687-699.
[3] Perreault, W. D. & Leigh, L. E. (1989). Reliability of nominal data based on qualitative judgments. Journal of Marketing Research, 26, 135-148.
The Perreault-Leigh agreement coefficient is formally defined as follows:
where S defined by,
is the agreement coefficient recommended by Bennet et al. (1954) and is a special case of the coefficient recommended by Brennan and Prediger (1981). The symbol Ir used by Perreault and Leigh (1989) appears to stand for “Index of Reliability.”
I carefully reviewed the Perreault and Leigh article. It presents an excellent review of the various agreement coefficients that were current at the time it was written. Perreault and Leigh define Ir as the percent of subjects that a typical judge could code consistent given the nature of the observations. Note that Ir is an attribute of the typical judge, and therefore does not represent any aspect of agreement among the judges. Perreault and Leigh (1989) consider the product NxIr2 (with N representing the number of subjects) to represent the number of reliable judgments on which judges agree. This cannot be true. To see this note that Ir2 is the probability that 2 judges both independently perform a reliable judgment. If both (reliable) judgments must lead to an agreement then they have to refer to the exact same category. However the probability Ir2 does not say which category was chosen and cannot represent any agreement among judges. Even if you decide to assume that any 2 reliable judgments must necessarily result in an agreement, then the judgments will no longer be independent. The probability for two judges to agree will now become equal to the probability for the first rater to perform a reliable judgment times the conditional probability for the second judge to perform a reliability judgment given that the first judge did. This second conditional probability cannot be evaluated unless there are additional assumptions.
What Perreault and Leigh (1989) have proposed is not an agreement coefficient. Their coefficient quantifies something other than the extent of agreement among raters. It should not be compared with the other coefficients available in the literature until someone can tell us what it does.
References:
[1] Bennett, E. M., Alpert, R. & Goldstein, A. C. (1954). Communication through limited response questioning. Public Opinion Quarterly, 18, 303-308.
[2] Brennan, R. L., & Prediger, D. J. (1981). Coefficient kappa: Some uses, misuses, and alternatives. Educational and Psychological Measurement, 41, 687-699.
[3] Perreault, W. D. & Leigh, L. E. (1989). Reliability of nominal data based on qualitative judgments. Journal of Marketing Research, 26, 135-148.
Tuesday, February 25, 2014
Inter-rater reliability and Many-Facet Rasch Measurement
I just finished reading the book entitled "Introduction to Many-Facet Rasch Measurement" by Thomas Eckes. In this book, Mr. Thomas Eckes argues that the classical approach to inter-rater reliability that consists of training the raters and measuring their extent of agreement until they reach an acceptable level does not really work. It is because no matter how much training the raters received, they will still not be interchangeable. A residual intrinsic disagreement will remain among the raters, some of them being more stringent than others in their approach to rating.
The solution that Mr. Eckes proposes is to develop statistical models that describe the different facets of the inter-rater reliability experiment, such as the rater facet, the subject facet and possibly other facets. These statistical models will then be used to make some adjustments to the ratings so that the subjects supposed to be humans can get a fair test. This adjustment will supposedly not penalize the subjects who were unlucky enough to be rated by the more severe raters.
I must say I did like this book very much in the way the author describes the different issues associated with an inter-rater reliability experiment. The presentation of these issues by the author is very instructive and is done with considerable clarity. That alone justifies the investment in time and money one can make on this book. However, I have always been somehow skeptical about the use of theoretical statistical models for the purpose of making important practical decisions, especially decisions involving human subjects. As a matter of fact, even if the raters introduce some bias in the ratings, two statisticians will probably not recommend the same statistical models either. Using these models to adjust the ratings may only be adding the statistician bias that could compound with the rater bias to produce an outcome that can hardly be seen as more reliable. The statistical models can always help the researcher gain more insight into a reality with powerful modelling tools, but cannot and should not be seen as an expression of that reality. Nevertheless, this book is remarkably well written, and should certainly be useful to anyone interested in the topic of inter-rater reliability.
Bibliography
[1] Eckes, Thomas. (2011). Introduction to Many-Facet Rasch Measurement. Peter Lang, ISBN: 978-3-631-61350-4.
The solution that Mr. Eckes proposes is to develop statistical models that describe the different facets of the inter-rater reliability experiment, such as the rater facet, the subject facet and possibly other facets. These statistical models will then be used to make some adjustments to the ratings so that the subjects supposed to be humans can get a fair test. This adjustment will supposedly not penalize the subjects who were unlucky enough to be rated by the more severe raters.
I must say I did like this book very much in the way the author describes the different issues associated with an inter-rater reliability experiment. The presentation of these issues by the author is very instructive and is done with considerable clarity. That alone justifies the investment in time and money one can make on this book. However, I have always been somehow skeptical about the use of theoretical statistical models for the purpose of making important practical decisions, especially decisions involving human subjects. As a matter of fact, even if the raters introduce some bias in the ratings, two statisticians will probably not recommend the same statistical models either. Using these models to adjust the ratings may only be adding the statistician bias that could compound with the rater bias to produce an outcome that can hardly be seen as more reliable. The statistical models can always help the researcher gain more insight into a reality with powerful modelling tools, but cannot and should not be seen as an expression of that reality. Nevertheless, this book is remarkably well written, and should certainly be useful to anyone interested in the topic of inter-rater reliability.
Bibliography
[1] Eckes, Thomas. (2011). Introduction to Many-Facet Rasch Measurement. Peter Lang, ISBN: 978-3-631-61350-4.
Wednesday, December 18, 2013
The Paradoxes of Agreement Coefficients: An Impossible Justification
Feinstein and Cicchetti (1990) exposed the kappa coefficient of Cohen (1960) as an agreement metric prone to yield unduly low values when the distribution of subjects is skewed towards one category, even when the raters strongly agree about their ratings. This problem is known in the inter-rater reliability literature as the kappa paradox. As it turned out, kappa was not the only agreement coefficient to carry this issue. A few other agreement coefficients - Scott’s (1955) π and Krippendorff’s (1980, 2004a, 2012) α among others - used by some researchers have the same problem. Despite the availability of abundant, well-documented and strong evidence supporting its seriousness, some authors have attempted and are still attempting to present the kappa paradox as a non-issue or a side issue. Presenting one’s viewpoint is always welcome, as it increases knowledge and provides insights into a problem. What irritates me is when some scholars become demagogues in an attempt to defend their past contributions to the literature, instead of using new evidence to improve them.
Kraemer et al. (2002, p. 2114) attempted to defend the kappa coefficient by arguing that what is presented as a kappa paradox is not a paradox. I pointed out in Gwet (2012, p. 38) that Kraemer et al. (2002) only made excuses for the poor kappa performance by blaming the distribution of subjects. More recently, when commenting an article by Zhao, Liu, and Deng (2013), Krippendorff (2013) made another even more demagogic argument to come to the conclusion that the low values of Cohen’s (1960) k , Scott’s (1955) π and Krippendorff’s (1980, 2004a, 2012) α, even in the presence of high raters’ agreement are justified.
To make his case, here is what Krippendorff (2013) says:
“Suppose an instrument manufacturer claims to have developed a test to diagnose a rare disease. Rare means that the probability of that disease in a population is small and to have enough cases in
the sample, a large number of individuals need to be tested. Let us use the authors’ numerical example: Suppose two separate doctors administer the test to the same 1,000 individuals. Suppose each doctor finds one in 1,000 to have the disease and they agree in 998 cases on the outcome of the test. The authors note that Cohen’s (1960) &kappa , Scott’s (1955) π, and Krippendorff’s (1980, 2004a, 2012) α are all below zero (-0.001 or -0.0005)...”
“... I contend that a test which produces 99.8% negatives, 2% disagreements, and not a single case of an agreement on the presence of the disease is totally unreliable indeed. Nobody in her right mind should trust a doctor who would treat patients based on such test results. The inference of zero is perfectly justifiable. The paradox of “high agreement but low reliability” does not characterize any of the reliability indices cited but resides entirely in the authors’ conceptual limitations. How could the authors be so wrong ?”
I am outraged by this point. Here is how I perceive it. If you are going to quantify the extent of agreement among raters who strongly agree in any sense you can think of, then your scoring method must assign a high agreement coefficient to these raters. If it fails to do so, then don’t switch topics by pretending that your low coefficient must instead be associated with the unascertained shortcomings of the measuring instrument. The propensity of a test to detect a rare trait is an entirely different topic, which requires a different experimental design and different quantitative methods with little in common with agreement coefficients.
Let us scrutinize a little further what is said in Krippendorff (2013):
Kraemer et al. (2002, p. 2114) attempted to defend the kappa coefficient by arguing that what is presented as a kappa paradox is not a paradox. I pointed out in Gwet (2012, p. 38) that Kraemer et al. (2002) only made excuses for the poor kappa performance by blaming the distribution of subjects. More recently, when commenting an article by Zhao, Liu, and Deng (2013), Krippendorff (2013) made another even more demagogic argument to come to the conclusion that the low values of Cohen’s (1960) k , Scott’s (1955) π and Krippendorff’s (1980, 2004a, 2012) α, even in the presence of high raters’ agreement are justified.
To make his case, here is what Krippendorff (2013) says:
“Suppose an instrument manufacturer claims to have developed a test to diagnose a rare disease. Rare means that the probability of that disease in a population is small and to have enough cases in
the sample, a large number of individuals need to be tested. Let us use the authors’ numerical example: Suppose two separate doctors administer the test to the same 1,000 individuals. Suppose each doctor finds one in 1,000 to have the disease and they agree in 998 cases on the outcome of the test. The authors note that Cohen’s (1960) &kappa , Scott’s (1955) π, and Krippendorff’s (1980, 2004a, 2012) α are all below zero (-0.001 or -0.0005)...”
“... I contend that a test which produces 99.8% negatives, 2% disagreements, and not a single case of an agreement on the presence of the disease is totally unreliable indeed. Nobody in her right mind should trust a doctor who would treat patients based on such test results. The inference of zero is perfectly justifiable. The paradox of “high agreement but low reliability” does not characterize any of the reliability indices cited but resides entirely in the authors’ conceptual limitations. How could the authors be so wrong ?”
I am outraged by this point. Here is how I perceive it. If you are going to quantify the extent of agreement among raters who strongly agree in any sense you can think of, then your scoring method must assign a high agreement coefficient to these raters. If it fails to do so, then don’t switch topics by pretending that your low coefficient must instead be associated with the unascertained shortcomings of the measuring instrument. The propensity of a test to detect a rare trait is an entirely different topic, which requires a different experimental design and different quantitative methods with little in common with agreement coefficients.
Let us scrutinize a little further what is said in Krippendorff (2013):
- In order to justify the unjustifiable, and explain the inexplicable, Krippendorff (2013) attempts to stay away from the initial goal of agreement coefficients by bringing in some fuzzy notions such as “informational context” or “reliability to be inferred.” If an agreement coefficient is now required to quantify such broad notions, then how do we know what theoretical construct we want to quantify ? There is here an unfortunate demagogic attempt to expand the clear concept of agreement as much as necessary until it can incorporate even the most outlying estimations, which may not be justified otherwise. When the US financial industry lowered the requirements for obtaining a loan, the notion of acceptable mortgage credit risk was artificially expanded. As a result, individuals with bad credit history qualified. We all know what followed.
- In the example above, suppose the instrument manufacturer developed a highly reliable test to diagnose a very common disease (i.e. a disease with high prevalence rate). Suppose also that each doctor finds only one individual in 1,000 without the disease, and both agree in 998 cases (i.e. correctly identify the same 998 patients with the disease), would Cohen’s (1960) &alpha , Scott’s (1955) π, and Krippendorff’s (1980, 2004a, 2012) α tell the correct story? Unfortunately the answer is still no. This proves beyond any doubt that the poor performance of these indices has little to do with the reliability of the measuring instrument.
- Notice the sentence “The inference of zero is perfectly justifiable.” Really ! Can an estimate of 0 be now called “an inference of 0 ?” What does the word “inference” mean in this context ? Is this inference statistical? This is what I refer to as pure demagogy, when someone decides to carry the word inference everywhere for the sole purpose of a conveying a false sense of sophistication.
- Looking at the experiment described above, how does one know whether the instrument itself is reliable or not ? If you want to test the propensity for an instrument to properly detect the presence of a rare trait, the statistical method of choice is the odds ratio, and not the agreement coefficient. Moreover, the use of 2 ordinary raters in an experiment aimed at testing the effectiveness of an instrument is rather odd, unless they are known experts in the use of that device. I am not sure why the effectiveness of the instrument is even part of this discussion.
Bibliography
[1] Cohen, J. (1960). “A coefficient of agreement for nominal scales.” Educational and Psychological Measurement, 20, 37-46.
[2] Feinstein, A. R., and Cicchetti, D. V. (1990), “High agreement but low kappa : I. The problems of two paradoxes,” Journal of Clinical Epidemiology, 43, 543-549.
[3] Gwet, K. L. (2012). Handbook of inter-rater reliability: The definitive to measuring the extent of agreement among multiple raters (3rded.). Gaithersburg, MD: Advanced Analytics, LLC Statistics in Medicine, 21, 2109-2129.
[4] Kraemer, H. C., Peryakoil, V. S., and Noda, A. (2002). “Kappa Coefficients in Medical Research,” Statistics in Medicine, 21, 2109-2129.
[5] Krippendorff, K. (1980). Content Analysis: An Introduction to Its Methodology. Thousand Oaks, Calif, USA.
[6] Krippendorff, K. (2004). Content Analysis: An Introduction to Its Methodology (2nd ed.). Thousand Oaks, Calif, USA.
[7] Krippendorff, K. (2012). Content Analysis: An Introduction to Its Methodology (3rd ed.). Thousand Oaks, Calif, USA.
[8] Krippendorff, K. (2013). Commentary : “A dissenting view on so-called paradoxes of reliability coefficients.” In C. T. Salmon (Ed.), Communication Yearbook, 36, (pp. 481-499).
[9] Scott, W. A. (1955). “Reliability of content analysis : the case of nominal scale coding.” Public Opinion Quarterly, XIX, 321-325.
[10] Xinshu, Zhao, Jun, S. Liu, and Ke, Deng. (2013). “Assumptions behind intercoder reliability indices,” In C. T. Salmon (Ed.), Communication Yearbook, 36, (pp. 419-499).

