Computer ScienceLinguistics

Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, Lizhuang Ma

2024.5.17Visual Intelligence

DOI: 10.1007/s44267-025-00099-6

tlooto Summary

This survey summarizes the timeline of representative efficient MLLMs, the current state of research in structures and strategies, and the applications, and the limitations of current efficient MLLM research and promising future directions are discussed.

Abstract

In the past years, multimodal large language models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering and visual understanding and reasoning. However, the extensive model size and high training and inference costs have hindered the widespread application of MLLMs in academia and industry. Thus, studying efficient and lightweight MLLMs has enormous potential, especially in edge computing scenarios. In this survey, we provide a comprehensive and systematic review of the current state of efficient MLLMs. Specifically, this survey summarizes the timeline of representative efficient MLLMs, the current state of research in structures and strategies, and the applications. Finally, the limitations of current efficient MLLM research and promising future directions are discussed.

Citation format

JIN, Yizhang, et al. Efficient multimodal large language models: A survey [preprint]. arXiv, 2024. arXiv:2405.10739.