Han, C., & Lu, X. (2026). Large language models as zero-shot evaluators of English–Chinese interpreting: A comparison of GPT-4o and DeepSeek-R1. Language Testing, 43(3), 263–289.
Related Publications
Han, C., Lu, X., Wang, W., & Chen, S. (2026). Applying n-gram-based evaluation metrics to assess human interpreting: A battery of replications with internal meta-analysis. Interpreting, 28(1), 58–90.
Jiang, Z., & Zhang, Z. (2025). From black box to transparency: Enhancing automated interpreting assessment with explainable AI in college classrooms. Research Methods in Applied Linguistics, 4(3), 100237.
Han, C. (2025). Quality assessment in multilingual, multimodal, and multiagent translation and interpreting (QAM³T&I): Proposing a unifying framework for research. Interpreting and Society: An Interdisciplinary Journal, 5(1), 27–55.
Han, C., Xiao, R., & Su, W. (2021). Assessing the fidelity of consecutive interpreting: The effects of using source versus target text as the reference material. Interpreting, 23(2), 245–268.