Academic Journal

A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.
Συγγραφείς: Liang, Chia Xin, Tian, Pu, Yin, Caitlyn Heqi, Yua, Yao, Wei, An-Hou, Li, Ming, Song, Xinyuan, Wang, Tianyang, Bi, Ziqian, Liu, Ming, Bao, Riyang, Feng, Pengbin
Πηγή: Computation; Jun2026, Vol. 14 Issue 6, p125, 91p
Θεματικοί όροι: Language models, Photograph captions
Περίληψη: This survey provides a comprehensive guide to Multimodal Large Language Models (MLLMs) with a focus on vision–language tasks, including image captioning, visual question answering, cross-modal retrieval, visual grounding, multi-image reasoning, long-video understanding, and embodied AI. We examine architectures, training pipelines, and practical applications, covering visual encoders, language model backbones, connector modules, contrastive pre-training, instruction tuning, and preference alignment. We also foreground first-principles constraints—information bottlenecks, data-processing limits, and statistical co-occurrence bias—that shape architecture, robustness, and evaluation. This survey centers on vision–language systems and does not cover audio-only models or code-generation tools without visual inputs. Through task-level analysis and system-level case studies, we examine prominent MLLM implementations while addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning. We present a unified taxonomy of the MLLM design space, a comparative overview of representative models and evaluation benchmarks, and a discussion of open problems. Concluding with ethical considerations and responsible AI development, this survey offers theoretical frameworks and practical insights for researchers, practitioners, and students working at the intersection of natural language processing and computer vision. [ABSTRACT FROM AUTHOR]
Copyright of Computation is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edb&genre=article&issn=20793197&ISBN=&volume=14&issue=6&date=20260601&spage=125&pages=125-215&title=Computation&atitle=A%20Comprehensive%20Survey%20and%20Guide%20to%20Multimodal%20Large%20Language%20Models%20in%20Vision%E2%80%93Language%20Tasks.&aulast=Liang%2C%20Chia%20Xin&id=DOI:10.3390/computation14060125
    Name: Full Text Finder (for New FTF UI) (ns324271)
    Category: fullText
    Text: Full Text Finder
    MouseOverText: Full Text Finder
Header DbId: edb
DbLabel: Complementary Index
An: 194907768
RelevancyScore: 1082
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1082.4189453125
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Liang%2C+Chia+Xin%22">Liang, Chia Xin</searchLink><br /><searchLink fieldCode="AR" term="%22Tian%2C+Pu%22">Tian, Pu</searchLink><br /><searchLink fieldCode="AR" term="%22Yin%2C+Caitlyn+Heqi%22">Yin, Caitlyn Heqi</searchLink><br /><searchLink fieldCode="AR" term="%22Yua%2C+Yao%22">Yua, Yao</searchLink><br /><searchLink fieldCode="AR" term="%22Wei%2C+An-Hou%22">Wei, An-Hou</searchLink><br /><searchLink fieldCode="AR" term="%22Li%2C+Ming%22">Li, Ming</searchLink><br /><searchLink fieldCode="AR" term="%22Song%2C+Xinyuan%22">Song, Xinyuan</searchLink><br /><searchLink fieldCode="AR" term="%22Wang%2C+Tianyang%22">Wang, Tianyang</searchLink><br /><searchLink fieldCode="AR" term="%22Bi%2C+Ziqian%22">Bi, Ziqian</searchLink><br /><searchLink fieldCode="AR" term="%22Liu%2C+Ming%22">Liu, Ming</searchLink><br /><searchLink fieldCode="AR" term="%22Bao%2C+Riyang%22">Bao, Riyang</searchLink><br /><searchLink fieldCode="AR" term="%22Feng%2C+Pengbin%22">Feng, Pengbin</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Computation; Jun2026, Vol. 14 Issue 6, p125, 91p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Photograph+captions%22">Photograph captions</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This survey provides a comprehensive guide to Multimodal Large Language Models (MLLMs) with a focus on vision–language tasks, including image captioning, visual question answering, cross-modal retrieval, visual grounding, multi-image reasoning, long-video understanding, and embodied AI. We examine architectures, training pipelines, and practical applications, covering visual encoders, language model backbones, connector modules, contrastive pre-training, instruction tuning, and preference alignment. We also foreground first-principles constraints—information bottlenecks, data-processing limits, and statistical co-occurrence bias—that shape architecture, robustness, and evaluation. This survey centers on vision–language systems and does not cover audio-only models or code-generation tools without visual inputs. Through task-level analysis and system-level case studies, we examine prominent MLLM implementations while addressing key challenges in scalability, memory, energy use, inference cost, robustness, and cross-modal learning. We present a unified taxonomy of the MLLM design space, a comparative overview of representative models and evaluation benchmarks, and a discussion of open problems. Concluding with ethical considerations and responsible AI development, this survey offers theoretical frameworks and practical insights for researchers, practitioners, and students working at the intersection of natural language processing and computer vision. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Computation is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=194907768
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3390/computation14060125
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 91
        StartPage: 125
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Photograph captions
        Type: general
    Titles:
      – TitleFull: A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Liang, Chia Xin
      – PersonEntity:
          Name:
            NameFull: Tian, Pu
      – PersonEntity:
          Name:
            NameFull: Yin, Caitlyn Heqi
      – PersonEntity:
          Name:
            NameFull: Yua, Yao
      – PersonEntity:
          Name:
            NameFull: Wei, An-Hou
      – PersonEntity:
          Name:
            NameFull: Li, Ming
      – PersonEntity:
          Name:
            NameFull: Song, Xinyuan
      – PersonEntity:
          Name:
            NameFull: Wang, Tianyang
      – PersonEntity:
          Name:
            NameFull: Bi, Ziqian
      – PersonEntity:
          Name:
            NameFull: Liu, Ming
      – PersonEntity:
          Name:
            NameFull: Bao, Riyang
      – PersonEntity:
          Name:
            NameFull: Feng, Pengbin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 20793197
          Numbering:
            – Type: volume
              Value: 14
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: Computation
              Type: main
ResultId 1