Natural Language Processing TechniquesTopic ModelingText Readability and Simplification

D. Suleymanov, A. Gatiatullin, N. A. Prokopyev

2026.4.14Uchenye Zapiski Kazanskogo Universiteta-Seriya Fiziko-Matematicheskie Nauki

DOI: 10.26907/2541-7746.2026.1.152-166

Abstract

This article focuses on methods of developing machine translation systems for low-resource languages, such as Turkic languages. A hybrid approach combining data augmentation and transfer learning was proposed. Data augmentation was achieved via morphotranslation of parallel corpora from closely related languages and paraphrasing based on the “Turklang” linguistic knowledge graph, enabling the generation of synthetic parallel data through an intermediate ontological representation. The BERT architecture was selected as both encoder and decoder. The training data comprised parallel corpora of approximately 30 million sentence pairs. The experimental part included training and evaluation of both one-way and multilingual models for six of the most low-resource Turkic–Russian language pairs: Crimean Tatar, Khakas, Tuvan, Altaic, Kumyk, and Yakut. The training was performed on a small portion of the collected data to ensure data volume balance between different languages. The results show that the multilingual model, trained using data augmentation and transfer learning, significantly improves BLEU, SacreBLEU, and ChrF scores compared to the one-way models.

Citation format

SULEYMANOV, D.; GATIATULLIN, A.; PROKOPYEV, N. A. Methods of developing machine translation systems for low-resource languages. Uchenye Zapiski Kazanskogo Universiteta-Seriya Fiziko-Matematicheskie Nauki, 2026, 168(1): 152–166.