Natural Language Processing TechniquesText Readability and SimplificationSpeech Recognition and Synthesis

Xinfeng Liao, Jiahang Xu, Jie Li, Xixuan Huang

2026.6.1Data Intelligence

DOI: 10.3724/2096-7004.di.2026.0092

Abstract

Spelling Correction (SC) in low-resource languages is significantly challenged by scarce labeled data and unbalanced error distributions. To address these limitations, this paper presents a multistage framework integrating data augmentation and prompt learning. First, we establish a hybrid rule-LLM augmentation mechanism that generates large-scale pseudo-error corpora from diverse sources, ensuring broad error coverage. Second, we implement a progressive training paradigm—transitioning from augmented data to human annotations—to effectively mitigate distribution biases during error detection. Finally, we introduce a prompt-based fine-tuning strategy with hierarchical templates, guiding lightweight Large Language Models (LLMs) to perform precise correction via error localization signals. Our framework ranked first in all four tracks of the IMLIP 2025 Low-Resource SC evaluation. This work provides a practical data-centric solution and demonstrates that lightweight models, when equipped with effective data strategies, can address data scarcity in low-resource spelling correction.

Citation format

LIAO, Xinfeng, et al. A multi-stage framework for spelling correction in low-resource languages with data augmentation and prompt learning. Data Intelligence, 2026: 20260092.