Deependra Singh, Saksham Agarwal, Subhankar Mishra
Abstract
Retinal diseases impact a significant portion of the global population, yet specialized care is predominantly available in urban areas. We have developed AI-based methods for diagnosing various retinal conditions using fundus images to address this disparity. We aim to create a highly accurate automatic detection system to bridge the gap between patients and the limited number of retinal specialists. Critical challenges in multi-label classification tasks include limited sample sizes per label and imbalanced class distributions. To overcome these issues and enhance data diversity, we created three composite datasets (MRID) by aggregating multiple open source datasets. We developed hybrid models integrating deep Convolutional Neural Networks (CNNs), Transformer encoders, and ensemble architectures to classify retinal fundus images into 21 distinct labels. Our models demonstrated superior performance compared to baseline methods and the existing state-of-the-art model, with significant improvements in evaluation metrics. Performance was further enhanced by incorporating domain knowledge into the patch extraction step of Vision Transformers used in specific hybrid models. To address user trust and model interpretability, we implemented SHAP-based explanations and used these alongside evaluation metrics for comparative analysis of model performance across datasets. This research aims to advance retinal disease diagnosis, offering accessible healthcare solutions in underserved regions while ensuring accurate and comprehensive disease prediction.
Citation format
SINGH, Deependra; AGARWAL, Saksham; MISHRA, Subhankar. Interpretable AI models for detecting and classifying multiple retinal conditions using hybrid CNN-Transformer-Ensemble architectures. ACM Transactions on Computing for Healthcare, 2026, 7(2): 1–29.