Guoqing Hu, Hao Wang, S. Yau

2026Communications in Information and Systems

DOI: 10.4310/cis.260612020426

Abstract

With the rapid development of genome sequencing technology, genomic sequence analysis has become an important field in modern biological research. However, sequencing errors, repetitive regions, and complex biological processes often lead to missing or ambiguous bases in genomic sequences, which are typically represented by non-standard symbols (such as R, Y, S, W, K, etc.). These issues severely affect the accuracy of genomic data, especially in tasks such as gene assembly and variant detection. To address this issue, this study proposes an encoding method based on asymmetric covariance natural vectors to characterize genomic sequences and predict ambiguous bases using Gated Recurrent Unit (GRU). Experimental results demonstrate that, compared with traditional encoding methods (such as one-hot encoding), the asymmetric covariance natural vector can more effectively utilize the information surrounding missing nucleotides for prediction, showing significant advantages in recovering nucleotides missing from intermediate positions. Additionally, this method also performs well on the SARS-CoV-2 Alpha variant dataset, with an error rate of only 1.09% in predicting non-standard bases during the encoding recovery process, further validating its effectiveness and potential for practical genomic data analysis.

Citation format

HU, Guoqing; WANG, Hao; YAU, S. Asymmetric natural vector method for predicting ambiguous nonstandard base codes. Communications in Information and Systems, 2026.