Skip to main navigation Skip to search Skip to main content

Using Elo Rating as a Metric for Comparative Judgement in Educational Assessment

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearch

Abstract

Marking and feedback are essential features of teaching and learning, across the overwhelming majority of educational settings and contexts. However, it can take a great deal of time and effort for teachers to mark assessments, and to provide useful feedback to the students. Furthermore, it also creates a significant cognitive load on the assessors, especially in ensuring fairness and equity. Therefore, an alternative approach to marking called comparative judgement (CJ) has been proposed in the educational space. Inspired by the law of comparative judgment (LCJ). The key idea here is that the better submission between a pair will be identified by a suitably qualified or experienced assessor. This pairwise comparison for as many pairs as possible can then be used to rank all submissions. Studies suggest that CJ is highly reliable and accurate while making it quick for the teachers. Alternative studies have questioned this claim suggesting that the process can increase bias in the results as the same submission is shown many times to an assessor for increasing reliability. Additionally, studies have also found that CJ can result in the overall marking process taking longer than a more traditional method of marking as information about many pairs must be collected. There is a clear necessity to investigate the efficacy of alternative rating and ranking systems that do not require extensive data on every pair of submissions, to reduce the temporal and cognitive burden on assessors, and bias from observing the same submission repeatedly. In this paper, we investigate Elo, which has been extensively used in rating players in zero-sum games such as chess - for devising a ranking between submissions in a comparative judgement context. We experimented on a large-scale Twitter dataset on the topic of a recent major UK political event ("Brexit", the UK's political exit from the European Union) to ask users which tweet they found funnier between a pair selected from ten tweets. Our analysis of the data reveals that the Elo rating is statistically significantly similar to the CJ ranking with a Kendall's tau score of 0.96 and a p-value of . We finish with an informed discussion regarding the potential wider application of this approach to a range of educational contexts.

Original languageEnglish
Title of host publicationICEMT 2022 - 2022 6th International Conference on Education and Multimedia Technology
EditorsTomokazu Nakayama, Jianli Jiao, Edwin P. Christmann
Place of PublicationNew York NY USA
PublisherAssociation for Computing Machinery (ACM)
Pages272-278
Number of pages7
ISBN (Electronic)9781450396455
DOIs
Publication statusPublished - 2022
Externally publishedYes
EventInternational Conference on Education and Multimedia Technology 2022 - Hybrid, Guangzhou, China
Duration: 13 Jul 202215 Jul 2022
Conference number: 6th
https://dl.acm.org/doi/proceedings/10.1145/3551708 (Proceedings (ACM digital library))
https://www.icemt.org/icemt2022.html (Website)

Conference

ConferenceInternational Conference on Education and Multimedia Technology 2022
Abbreviated titleICEMT 2022
Country/TerritoryChina
CityGuangzhou
Period13/07/2215/07/22
Internet address

Keywords

  • Assessment
  • Bradley-Terry Model
  • Comparative Judgement
  • Elo Rating
  • Marking
  • Teaching and Learning

Cite this