Academic Journal

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.
Συγγραφείς: Liang, Chia Xin, Tian, Pu, Yin, Caitlyn Heqi, Yua, Yao, Wei, An-Hou, Li, Ming, Song, Xinyuan, Wang, Tianyang, Bi, Ziqian, Liu, Ming, Bao, Riyang, Feng, Pengbin
Πηγή: Computation; Jun2026, Vol. 14 Issue 6, p125, 91p
Θεματικοί όροι: Language models, Photograph captions
Περίληψη: This survey provides a comprehensive guide to Multimodal Large Language Models (MLLMs) with a focus on vision–language tasks, including image captioning, visual question answering, cross-modal retrieval, visual grounding, multi-image reasoning, long-video understanding, and embodied AI. We examine architectures, training pipelines, and practical applications, covering visual encoders, language model backbones, connector modules, contrastive pre-training, instruction tuning, and preference alignment. We also foreground first-principles constraints—information bottlenecks, data-processing limits, and statistical co-occurrence bias—that shape architecture, robustness, and evaluation. This survey centers on vision–language systems and does not cover audio-only models or code-generation tools without visual inputs. Through task-level analysis and system-level case studies, we examine prominent MLLM implementations while addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning. We present a unified taxonomy of the MLLM design space, a comparative overview of representative models and evaluation benchmarks, and a discussion of open problems. Concluding with ethical considerations and responsible AI development, this survey offers theoretical frameworks and practical insights for researchers, practitioners, and students working at the intersection of natural language processing and computer vision. [ABSTRACT FROM AUTHOR]
Copyright of Computation is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index