Translation Studies and PracticesText Readability and SimplificationNatural Language Processing Techniques
DOI: 10.25145/j.refiull.2026.52.14

Abstract

This paper contributes to research on machine translation quality assessment through a study focused on the translation of academic discourse between Italian and Spanish. Drawing on the relevance theory (Sperber and Wilson, 1986), the study compares two neural machine translation systems (NMT) (DeepL and Google Translate) with two LLMs (ChatGPT-4o and Deepseek), through both quantitative and qualitative analysis based on the MQM framework. The analysis is conducted on a medium-sized corpus consisting of ten abstracts in Italian and ten abstracts in Spanish. The results confirm the initial hypotheses. First, LLMs exhibit superior performance compared to NMT systems, producing fewer serious errors. Second, most errors affect the reconstruction of explicature, specifically within the categories of terminology and accuracy. Third, a high degree of stylistic variation is observed, with LLMs standing out for their ability to rephrase and enhance texts in terms of clarity. Fourth, system performance proves to be better when translating from Italian into Spanish. Ultimately, the human translator is reaffirmed as the key agent responsible for ensuring the overall textual quality.

Citation format

SAINZ, E.; BOVE, Antonella. Discurso académico, traducción automática y evaluación de la calidad: Comparación entre sistemas. REVISTA DE FILOLOGIA DE LA UNIVERSIDAD DE LA LAGUNA, 2026: 379–416.