Interpreting and Communication in HealthcareNatural Language Processing TechniquesSpeech and dialogue systems

Anja Rütten

2026.4.22Lebende Sprachen

DOI: 10.1515/les-2025-0046

Abstract

Abstract This article explores the benefits of using automatic quality scores designed for machine translation (MT) to obtain an indicative quality estimation for individual segments of both automatic speech translation (AST) and human simultaneous interpretation (HSI). In a first step, a set of assessment metrics for interpreting (AMI) is set up using MQM as a starting point and completing and adapting it based on quality criteria from interpreting studies and practice. A sample human simultaneous interpretation and automatic speech translation are then assessed segment by segment using AMI and compared to the COMET scores calculated for these segments. A comparative analysis of the results explores potential correlations between the human quality assessment and the COMET scores, the focus being semantic deviations of the target from the source text. The study shows higher amounts of semantic deviations and grammar issues for lower COMET scores in both HSI and AST, suggesting that using automatic quality estimation scores as a pre-screening instrument for human experts to single out critical segments of a speech when assessing AST or HSI might be an avenue worth exploring.

Citation format

RÜTTEN, Anja. Can automatic quality estimation help to pre-assess the quality of human simultaneous interpretation and automatic speech translation? An exploratory case study based on MQM-inspired assessment metrics for interpreting (AMI) and COMET. Lebende Sprachen, 2026, 0.