Nan Sun, Mengcen Guan, Piyu Zhou, S. Yau
2026.6.4BMC BIOINFORMATICS
Abstract
Single-cell RNA sequencing (scRNA-seq) enables high resolution characterization of cellular heterogeneity but poses significant challenges for cross institutional collaboration due to privacy constraints and distributional heterogeneity. To address this problem, we propose a Federated Distillation framework with Knowledge Sharing (scKSFD) for privacy-preserving cell type classification. Unlike conventional federated learning approaches that exchange model parameters, scKSFD performs knowledge aggregation in prediction space by sharing probability-level soft label outputs on a reference dataset, thereby reducing privacy risks. To better accommodate domain specific characteristics of scRNA-seq data, scKSFD integrates stratified proxy sampling to preserve rare cell populations and employs probability-level aggregation to mitigate batch specific expression shifts without explicit feature level correction. Comprehensive evaluations across 42 clinical single-cell transcriptome datasets demonstrate that scKSFD achieves higher or comparable F1 scores relative to centralized and existing federated baselines under heterogeneous settings, with statistically significant improvements in paired comparisons. In a multiple hospital COVID-19 case study, federated collaboration using scKSFD improved classification performance compared with local-only training while avoiding direct sharing of patient level expression data. Overall, scKSFD provides a federated distillation framework that balances predictive performance, robustness, and data-sharing constraints for multiple institutional single-cell transcriptomic analysis.
Citation format
SUN, Nan, et al. Scksfd: Federated distillation model with knowledge sharing for cell type classification of clinical transcriptome data. BMC BIOINFORMATICS, 2026.