Topic ModelingSpeech and dialogue systemsMultimodal Machine Learning Applications

Padegal Amit, Omkar Mahesh Kashyap, Namitha Rayasam, Nidhi Shekhar, Surabhi Narayan

2026.5.6Proceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS

DOI: 10.32473/flairs.39.1.141869

Abstract

Quantifying modality contributions in Vision-Language Models (VLMs) remains challenging. Existing approaches rely on perturbation or gradient-based methods, which conflate inherent modality informativeness with model-specific biases and fail to capture complex cross-modal interactions. We address this gap by introducing an information-theoretic framework based on Partial Information Decomposition (PID) that decomposes internal representations into unique, redundant, and synergistic components. Our method operates directly on internal embeddings and derives an inference-only modality contribution metric from unique information scores. Applying our framework to six modern VLMs across six benchmarks, we uncover a persistent imbalance in modality contributions driven by low cross-modal synergy. Analysis reveals that fusion architecture significantly impacts the distribution of unique, redundant, and synergistic information. Our framework provides a scalable diagnostic tool for understanding and improving multimodal integration in vision-language systems.

Citation format

AMIT, Padegal, et al. Quantifying modality contributions via disentangling multimodal representations. Proceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS, 2026, 39(1).