C. Kuo, Kuan-Yu Chen
2026.4.16APSIPA Transactions on Signal and Information Processing
Abstract
In recent years, automatic speech recognition (ASR) systems have become essential components of various computer applications, enabling seamless communication between humans and machines. N-best reranking and error correction are widely used post-processing techniques, each contributing to improving ASR output with distinct strengths. Building on these advantages, they propose a comprehensive framework designed to achieve superior performance. The framework begins with a text correction module to rectify and expand the original N-best list. This augmented list is then evaluated and reranked jointly by a text rescoring module and a text–speech matching module. The text rescoring module assesses hypotheses from a textual perspective, while the text–speech matching module evaluates each hypothesis from the alignment between speech and text. The authors explore three distinct strategies for implementing the text–speech matching module. Using the publicly available HypR benchmark, they fairly evaluated the proposed framework against other methods. Experimental results demonstrate the feasibility of the proposed framework and highlight the contributions of this study.
Citation format
KUO, C.; CHEN, Kuan-Yu. An integrated framework for ASR post-processing with error correction, rescoring, and text-speech matching. APSIPA Transactions on Signal and Information Processing, 2026, 15(1): 223–248.