S. Palanisamy, Chitra Duraisamy
2026.1.7JOURNAL OF BIOLOGICAL REGULATORS AND HOMEOSTATIC AGENTS
tlooto Summary
This research work solves class imbalance issue by reducing the quantity of majority class samples while trying to restore the original data distribution when the dataset is acquired and efficiency was compared to those of cutting-edge machine learning techniques like XG boost and random forest.
Abstract
<p class="MsoNormal" style="text-align: justify; text-justify: inter-ideograph; line-height: 120%; mso-pagination: none; layout-grid-mode: char; mso-layout-grid-align: none; text-autospace: none; margin: 12.0pt 0cm 6.0pt 147.4pt;"><span lang="EN-US" style="font-size: 10.0pt; mso-bidi-font-size: 11.0pt; line-height: 120%; font-family: 'Times New Roman',serif; color: black; mso-themecolor: text1;">When faced with imbalanced data, classification techniques in the area of artificial intelligence have a tendency to Favor the majority class samples, which lowers the recognition rates of minority class samples. This problem is solved by undersampling, which reduces the quantity of majority class samples while trying to restore the original data distribution when the dataset is acquired. The initial imbalanced dataset and its classification accuracy as a whole are strongly impacted by the constraints of the clustering-based undersampling techniques utilized today. To solve these issues, in this research work, initially the highly imbalanced dataset is pre-processed using Non-Negative Matrix Factorization (NMF) Algorithm. Next, Hybrid Extremely Randomized Trees (HERT), an efficient ensemble learning-based method, is employed to quickly choose the features. Afterwards, to solve class imbalance issue, Generative Adversarial Network (GAN)-based oversampling is suggested. This method has shown exceptional capacity to solve class imbalance as it may detect the genuine data distribution of minority class samples and produce new samples. By selecting useful instances from each cluster and avoiding information loss, the Fuzzy C means (FCM) clustering system is suggested for the undersampling method. Here Combined form of Fuzzy C means clustering for majority class and Adasyn-GAN centred over sampling for minority class are together to produce better results. Finally, the sampled dataset has undergone classification using Adaptive Weight Bi-Directional Long Short-Term Memory (AWBi-LSTM) classifier. Three huge, unbalanced data sets are applied to assess the suggested algorithm. The suggested system</span><span lang="EN-US" style="font-size: 10.0pt; mso-bidi-font-size: 11.0pt; line-height: 120%; font-family: 'Times New Roman',serif; color: black; mso-themecolor: text1; mso-fareast-language: ZH-CN;">’</span><span lang="EN-US" style="font-size: 10.0pt; mso-bidi-font-size: 11.0pt; line-height: 120%; font-family: 'Times New Roman',serif; color: black; mso-themecolor: text1;">s efficiency was compared to those of cutting-edge machine learning (ML) techniques like XG boost and random forest. The suggested method</span><span lang="EN-US" style="font-size: 10.0pt; mso-bidi-font-size: 11.0pt; line-height: 120%; font-family: 'Times New Roman',serif; color: black; mso-themecolor: text1; mso-fareast-language: ZH-CN;">’</span><span lang="EN-US" style="font-size: 10.0pt; mso-bidi-font-size: 11.0pt; line-height: 120%; font-family: 'Times New Roman',serif; color: black; mso-themecolor: text1;">s effectiveness is demonstrated by the performance assessment with regard to accuracy, recall, precision, and F1-score. Furthermore, the suggested plan requires less training time than cutting-edge methods.</span></p>
Citation format
PALANISAMY, S.; DURAISAMY, Chitra. AWBI-LSTM classifier with hybrid ADASYN-GAN oversampling and optimized FCM undersampling for imbalanced data. JOURNAL OF BIOLOGICAL REGULATORS AND HOMEOSTATIC AGENTS, 2026, 39(4): 8283.