M. Bolatbek, Shynar Mussiraliyeva, Kymbat Baisylbayeva

2025.12.30Research in Language

DOI: 10.18778/1731-7533.23.21

tlooto Summary

It is demonstrated that hybrid neural network models (CNN+BiLSTM) are the most effective solution, while DistilBERT performs best among transformer models.

Abstract

Modern information technologies enable the automatic analysis of textual data to detect extremist and propagandistic content. This paper examines deep learning methods and transformers models for the automatic classification of ideologically charged texts in the Kazakh language. A comparison was conducted between neural network models (CNN, BiLSTM, GRU, Hybrid CNN+BiLSTM) and modern transformers (DistilBERT). The performance evaluation of the models was based on accuracy, recall, precision, and F1-score metrics, as well as error analysis. Experimental results showed that hybrid CNN+BiLSTM demonstrated the highest accuracy (95.11%), outperforming other models. CNN, BiLSTM and GRU also achieved high results (92-93%), making them effective for this task. Among transformers, DistilBERT proved to be the most balanced (85.74%). This study demonstrates that hybrid neural network models (CNN+BiLSTM) are the most effective solution, while DistilBERT performs best among transformer models. The findings can be utilized for developing automatic monitoring and filtering systems for Kazakh-language texts, capable of efficiently identifying ideologically charged content.

Citation format

BOLATBEK, M.; MUSSIRALIYEVA, Shynar; BAISYLBAYEVA, Kymbat. Detection and classification of ideological texts in the kazakh language using machine learning and transformers. Research in Language, 2025.