Zero-shot medical image classification via large multimodal models and knowledge graphs-driven processing.

Bibliographic Details
Title: Zero-shot medical image classification via large multimodal models and knowledge graphs-driven processing.
Authors: Liu X; Key Laboratory of Water Big Data Technology of Ministry of Water Resources, Hohai University, Nanjing 211100, China; College of Computer Science and Software Engineering, Hohai University, Nanjing 211100, China. Electronic address: liuxinfu@hhu.edu.cn., Wu Y; Key Laboratory of Water Big Data Technology of Ministry of Water Resources, Hohai University, Nanjing 211100, China; College of Computer Science and Software Engineering, Hohai University, Nanjing 211100, China. Electronic address: wuyirui@hhu.edu.cn., Zhou Y; Key Laboratory of Water Big Data Technology of Ministry of Water Resources, Hohai University, Nanjing 211100, China; College of Computer Science and Software Engineering, Hohai University, Nanjing 211100, China. Electronic address: zhouyuting@hhu.edu.cn.
Source: Methods (San Diego, Calif.) [Methods] 2026 Jan; Vol. 245, pp. 25-34. Date of Electronic Publication: 2025 Oct 13.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: Academic Press Country of Publication: United States NLM ID: 9426302 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1095-9130 (Electronic) Linking ISSN: 10462023 NLM ISO Abbreviation: Methods Subsets: MEDLINE
Imprint Name(s): Publication: Duluth, MN : Academic Press
Original Publication: San Diego : Academic Press, c1990-
MeSH Terms: Image Processing, Computer-Assisted*/methods , Diagnostic Imaging*/methods , Diagnostic Imaging*/classification , Natural Language Processing*, Humans ; Algorithms ; Knowledge Bases
Abstract: With the continuous advancement of medical enterprise, intelligent medical technologies supported by natural language processing and knowledge representation have made significant progress. However, with the continuous generation of vast amounts of medical data, the current methods still perform poorly in handling specialized medical data, particularly unlabeled medical diagnostic data. Inspired by the outstanding performance of large language models in various downstream expert tasks in recent years, this article leverages large language models to handle the massive unlabelled medical data, aiming to provide more accurate technical solutions for medical image classification tasks. Specifically, we propose a novel Cross-Modal Knowledge Representation framework (CMKR) to handle vast unlabeled medical data, which utilizes large language models to extract implicit knowledge from medical images, while also extracting explicit textual knowledge with the aid of knowledge graphs. To better utilize the associative information between medical images and textual records, we have designed a cross-modal alignment strategy that enhances knowledge representation capabilities both intra- and inter-modal. We conducted extensive experiments on public datasets, demonstrating that our method outperforms most mainstream approaches.
(Copyright © 2025. Published by Elsevier Inc.)
Competing Interests: Declaration of competing interest The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Yirui Wu reports that financial support was provided by National Key R&D Program of China. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Cross-modal alignment; Knowledge graph; Knowledge representation; Large language model; Medical image classification
Entry Date(s): Date Created: 20251015 Date Completed: 20251126 Latest Revision: 20251126
Update Code: 20260130
DOI: 10.1016/j.ymeth.2025.09.006
PMID: 41093082
Database: MEDLINE
Be the first to leave a comment!
You must be logged in first