What makes an effective rubric for assessing interpreting?

Han, C., Jiang, M., & Chen, Q. (2026). Rubricizing the assessment practice: A systematic review and meta-analysis of rubrics in rater-mediated assessment of language interpreting. Language Testing, 43(2), 197–232. DOI

Rubrics are widely used in interpreter education, professional certification, and research, yet their development and effectiveness have rarely been examined systematically. In this study, we conducted the first large-scale systematic review and meta-analysis of rubric-based interpreting assessment, investigating how rubrics are developed, designed, applied, and validated.

Through database searching, citation tracking, and targeted examination of core literature, we analyzed 87 publications and catalogued 80 unique rubrics containing 265 subscales. We also synthesized 114 reliability coefficients from 21 empirical studies.

We found that most rubrics were analytic and designed for related types of interpreting tasks, typically containing four assessment criteria and five performance levels. Informational fidelity received the greatest weighting, averaging 52.3%, followed by target-language expression and delivery. Rubric use centered heavily on spoken consecutive interpreting and involved practitioners, trainers, and university students as raters.

A major concern involved rubric development. Approximately 80% of the reported development sources came from theory, existing rubrics, proficiency frameworks, or expert consultation, while performance samples and rater data were used much less frequently. Score use and policy context were almost entirely overlooked, even in high-stakes certification.

Overall, we found high rater reliability and moderate-to-strong validity. Reliability nevertheless varied by statistical coefficient, assessment criterion, and rubric length, with longer rubrics associated with lower reliability.

Our findings support concise, behaviorally anchored rubrics grounded in real interpreting performances and accompanied by systematic validation and rigorous rater training.

Back to Publication Highlights