VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.
Συγγραφείς: Yan C; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: yancaixia@xjtu.edu.cn., Jiao M; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: jmy0406@stu.xjtu.edu.cn., Xue N; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: xuenuohan@stu.xjtu.edu.cn., Zhang W; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: zhangwzh@xjtu.edu.cn., Wang J; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: uguisu@stu.xjtu.edu.cn., Chang X; Department of Electronic Engineering and Information Science, University of Science and Technology of China, China. Electronic address: cxj273@gmail.com., Tian F; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: fengtian@mail.xjtu.edu.cn.
Πηγή: Neural networks : the official journal of the International Neural Network Society [Neural Netw] 2026 Sep; Vol. 201, pp. 108899. Date of Electronic Publication: 2026 Mar 25.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Pergamon Press Country of Publication: United States NLM ID: 8805018 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1879-2782 (Electronic) Linking ISSN: 08936080 NLM ISO Abbreviation: Neural Netw Subsets: MEDLINE
Imprint Name(s): Original Publication: New York : Pergamon Press, [c1988-
Ιατρικοί όροι (MeSH): Detection Algorithms* , Generative Adversarial Networks*
Περίληψη: Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator's embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.
(Copyright © 2026 Elsevier Ltd. All rights reserved.)
Competing Interests: Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Generative adversarial networks; Vision-language distillation; Zero-shot object detection
Entry Date(s): Date Created: 20260403 Date Completed: 20260613 Latest Revision: 20260619
Update Code: 20260620
DOI: 10.1016/j.neunet.2026.108899
PMID: 41932121
Βάση Δεδομένων: MEDLINE
FullText Links:
  – Type: other
    Url: https://resolver.ebsco.com:443/public/rma-ftfapi/ejs/direct?AccessToken=4257A606EBB9F6E1B945&Show=Object
Text:
  Availability: 0
CustomLinks:
  – Url: https://www.doi.org/10.1016/j.neunet.2026.108899?
    Name: ScienceDirect (all content) (s7799221)
    Category: fullText
    Text: View record from ScienceDirect
    MouseOverText: View record from ScienceDirect
Header DbId: cmedm
DbLabel: MEDLINE
An: 41932121
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AU" term="%22Yan+C%22">Yan C</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: yancaixia@xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Jiao+M%22">Jiao M</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: jmy0406@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Xue+N%22">Xue N</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: xuenuohan@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Zhang+W%22">Zhang W</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: zhangwzh@xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Wang+J%22">Wang J</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: uguisu@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Chang+X%22">Chang X</searchLink>; Department of Electronic Engineering and Information Science, University of Science and Technology of China, China. Electronic address: cxj273@gmail.com.<br /><searchLink fieldCode="AU" term="%22Tian+F%22">Tian F</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: fengtian@mail.xjtu.edu.cn.
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%228805018%22">Neural networks : the official journal of the International Neural Network Society</searchLink> [Neural Netw] 2026 Sep; Vol. 201, pp. 108899. <i>Date of Electronic Publication: </i>2026 Mar 25.
– Name: TypePub
  Label: Publication Type
  Group: TypPub
  Data: Journal Article
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: TitleSource
  Label: Journal Info
  Group: Src
  Data: <i>Publisher: </i><searchLink fieldCode="PB" term="%22Pergamon+Press%22">Pergamon Press </searchLink><i>Country of Publication: </i>United States <i>NLM ID: </i>8805018 <i>Publication Model: </i>Print-Electronic <i>Cited Medium: </i>Internet <i>ISSN: </i>1879-2782 (Electronic) <i>Linking ISSN: </i><searchLink fieldCode="IS" term="%2208936080%22">08936080 </searchLink><i>NLM ISO Abbreviation: </i>Neural Netw <i>Subsets: </i>MEDLINE
– Name: PublisherInfo
  Label: Imprint Name(s)
  Group: PubInfo
  Data: <i>Original Publication</i>: New York : Pergamon Press, [c1988-
– Name: SubjectMESH
  Label: MeSH Terms
  Group: Su
  Data: <searchLink fieldCode="MM" term="%22Detection+Algorithms%22">Detection Algorithms*</searchLink> <br /><searchLink fieldCode="MM" term="%22Generative+Adversarial+Networks%22">Generative Adversarial Networks*</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator's embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.<br /> (Copyright © 2026 Elsevier Ltd. All rights reserved.)
– Name: Abstract
  Label: Competing Interests
  Group: Ab
  Data: Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
– Name: SubjectMinor
  Label: Contributed Indexing
  Group:
  Data: <i>Keywords: </i>Generative adversarial networks; Vision-language distillation; Zero-shot object detection
– Name: DateEntry
  Label: Entry Date(s)
  Group: Date
  Data: <i>Date Created: </i>20260403 <i>Date Completed: </i>20260613 <i>Latest Revision: </i>20260619
– Name: DateUpdate
  Label: Update Code
  Group: Date
  Data: 20260620
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1016/j.neunet.2026.108899
– Name: AN
  Label: PMID
  Group: ID
  Data: 41932121
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=cmedm&AN=41932121
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.neunet.2026.108899
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        StartPage: 108899
    Subjects:
      – SubjectFull: Detection Algorithms
        Type: general
      – SubjectFull: Generative Adversarial Networks
        Type: general
    Titles:
      – TitleFull: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yan C
      – PersonEntity:
          Name:
            NameFull: Jiao M
      – PersonEntity:
          Name:
            NameFull: Xue N
      – PersonEntity:
          Name:
            NameFull: Zhang W
      – PersonEntity:
          Name:
            NameFull: Wang J
      – PersonEntity:
          Name:
            NameFull: Chang X
      – PersonEntity:
          Name:
            NameFull: Tian F
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: 2026 Sep
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-electronic
              Value: 1879-2782
          Numbering:
            – Type: volume
              Value: 201
          Titles:
            – TitleFull: Neural networks : the official journal of the International Neural Network Society
              Type: main
ResultId 1