Simardeep Kaur, Siddhant Ranjan Padhi, T. Mithraa, Naseeb Singh, M. Tomar, Racheal John, Amit Kumar, Veerendra Kumar Verma, Mohar Singh, K. Tripathy, Gayacharan, Vinod Kumar, R. Kalia, A. Singh, D. P. Wankhede, J. Rana, Rakesh Bhardwaj, A. Riar
2026.1.1Future Foods
Abstract
Rapid, eco-friendly, and non-destructive estimation of protein content is crucial for efficient nutritional phenotyping and large-scale germplasm screening in legumes. Traditional biochemical methods are time-consuming, costly, and labor-intensive, posing challenges to breeders and the food industry. This study aimed to develop and validate universal near-infrared spectroscopy (NIRS)-based predictive models for protein quantification across multiple legume species. A genetically diverse dataset comprising 1,169 grain samples from cowpea, mung bean, horse gram, pea, lentil, faba bean, winged bean, adzuki bean, rice bean, lablab bean, and chickpea was utilized. Spectral data (1100–2498 nm) were preprocessed using Standard Normal Variate, detrending, derivatives, and smoothing techniques. Two models; Modified Partial Least Squares (MPLS) and one-dimensional Convolutional Neural Network (1D CNN) were developed and validated on an independent set of 351 samples. The 1D CNN model outperformed MPLS, achieving R² = 0.883 and RPD = 2.932, compared to MPLS (R² = 0.814; RPD = 2.320), demonstrating greater accuracy and robustness. This is the first report of a universal NIRS-based deep learning model for protein prediction across diverse legumes. Its integration into portable NIR sensors can accelerate field-based protein screening, enhancing breeding efficiency, gene bank evaluations, food quality control, and the development of functional foods.
Citation format
KAUR, Simardeep, et al. Protein informatics: Development and validation of a universal NIR spectroscopy-based deep learning and chemometric models for protein quantification in legume crops - a high-throughput approach for large germplasm screening. Future Foods, 2026, 13: 100909.