VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.
Συγγραφείς: Yan C; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: yancaixia@xjtu.edu.cn., Jiao M; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: jmy0406@stu.xjtu.edu.cn., Xue N; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: xuenuohan@stu.xjtu.edu.cn., Zhang W; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: zhangwzh@xjtu.edu.cn., Wang J; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: uguisu@stu.xjtu.edu.cn., Chang X; Department of Electronic Engineering and Information Science, University of Science and Technology of China, China. Electronic address: cxj273@gmail.com., Tian F; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: fengtian@mail.xjtu.edu.cn.
Πηγή: Neural networks : the official journal of the International Neural Network Society [Neural Netw] 2026 Sep; Vol. 201, pp. 108899. Date of Electronic Publication: 2026 Mar 25.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Pergamon Press Country of Publication: United States NLM ID: 8805018 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1879-2782 (Electronic) Linking ISSN: 08936080 NLM ISO Abbreviation: Neural Netw Subsets: MEDLINE
Imprint Name(s): Original Publication: New York : Pergamon Press, [c1988-
Ιατρικοί όροι (MeSH): Detection Algorithms* , Generative Adversarial Networks*
Περίληψη: Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator's embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.
(Copyright © 2026 Elsevier Ltd. All rights reserved.)
Competing Interests: Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Generative adversarial networks; Vision-language distillation; Zero-shot object detection
Entry Date(s): Date Created: 20260403 Date Completed: 20260613 Latest Revision: 20260619
Update Code: 20260620
DOI: 10.1016/j.neunet.2026.108899
PMID: 41932121
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1879-2782
DOI:10.1016/j.neunet.2026.108899