In this study, we conducted one of the first large-scale examinations of large language models as zero-shot evaluators of spoken-language interpreting. Using 541 pre-scored English–Chinese consecutive and simultaneous interpretations from the Interpreting Quality Evaluation Corpus, we compared GPT-4o and DeepSeek-R1 with trained human raters across three dimensions: scoring reliability, severity, and validity.
We created 16 LLM-based e-raters by systematically varying reference availability, scoring granularity, and model temperature. We then evaluated their performance using correlation analysis, linear mixed-effects modelling, and Many-Facet Rasch Measurement.
We found that both models generated more internally consistent scores than human raters and achieved moderately strong alignment with human judgments. Important differences nevertheless emerged. GPT-4o displayed scoring severity broadly comparable to that of human raters and achieved greater overall accuracy, whereas DeepSeek-R1 was considerably harsher. On average, GPT-4o assigned scores approximately 1.2 points higher than DeepSeek-R1 on an eight-point scale.
Assessment design also mattered. Segment-level scoring without reference interpretations generally produced the most accurate results, while temperature had little influence. Performance further varied across interpreting modes, directions, and quality criteria, with the strongest human–LLM alignment observed for information completeness.
Our findings demonstrate the potential of LLMs for scalable interpreting assessment while showing that model choice and scoring configuration critically shape the reliability and validity of automated judgments.