As artificial intelligence becomes increasingly capable of assessing complex language performance, we asked a critical question: How robust, valid, and transparent are current AI-based approaches to evaluating human translation and interpreting?
Following PRISMA guidelines, we systematically identified and analyzed 69 publications through a three-pronged search combining academic databases, AI-powered discovery tools, and citation tracking. We examined these studies using the QAM³T&I framework and a validation taxonomy covering assessment design, model architecture, human benchmark construction, score validity, generalizability, and transparency.
We found that research has expanded rapidly since 2020 but remains heavily concentrated on English–Chinese translation and interpreting in educational settings. Translation and consecutive interpreting dominate the literature, while subtitling, signed-language interpreting, post-editing, and other practices remain underexplored. Feature-based machine-learning models and repurposed machine-translation metrics remain the most common approaches, while conversational large language models represent only a small but growing share.
We also identified substantial methodological weaknesses. Many studies inadequately reported rater qualifications, training, reliability, and benchmark construction. Validation typically relied on correlations with human scores, with little attention to agreement, construct validity, or generalizability across populations and conditions. Post-hoc explainability was also rarely implemented.
Our review calls for transparent benchmark development, multi-pronged validation, more diverse languages and contexts, large annotated datasets, and explainable AI as foundations for trustworthy automated T&I assessment.