Academic Journal
VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection.
| Τίτλος: | VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection. |
|---|---|
| Συγγραφείς: | Yan C; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: yancaixia@xjtu.edu.cn., Jiao M; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: jmy0406@stu.xjtu.edu.cn., Xue N; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: xuenuohan@stu.xjtu.edu.cn., Zhang W; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: zhangwzh@xjtu.edu.cn., Wang J; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: uguisu@stu.xjtu.edu.cn., Chang X; Department of Electronic Engineering and Information Science, University of Science and Technology of China, China. Electronic address: cxj273@gmail.com., Tian F; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: fengtian@mail.xjtu.edu.cn. |
| Πηγή: | Neural networks : the official journal of the International Neural Network Society [Neural Netw] 2026 Sep; Vol. 201, pp. 108899. Date of Electronic Publication: 2026 Mar 25. |
| Τύπος έκδοσης: | Journal Article |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: Pergamon Press Country of Publication: United States NLM ID: 8805018 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1879-2782 (Electronic) Linking ISSN: 08936080 NLM ISO Abbreviation: Neural Netw Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: New York : Pergamon Press, [c1988- |
| Ιατρικοί όροι (MeSH): | Detection Algorithms* , Generative Adversarial Networks* |
| Περίληψη: | Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator's embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS. (Copyright © 2026 Elsevier Ltd. All rights reserved.) |
| Competing Interests: | Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. |
| Contributed Indexing: | Keywords: Generative adversarial networks; Vision-language distillation; Zero-shot object detection |
| Entry Date(s): | Date Created: 20260403 Date Completed: 20260613 Latest Revision: 20260619 |
| Update Code: | 20260620 |
| DOI: | 10.1016/j.neunet.2026.108899 |
| PMID: | 41932121 |
| Βάση Δεδομένων: | MEDLINE |
| FullText | Links: – Type: other Url: https://resolver.ebsco.com:443/public/rma-ftfapi/ejs/direct?AccessToken=4257A606EBB9F6E1B945&Show=Object Text: Availability: 0 CustomLinks: – Url: https://www.doi.org/10.1016/j.neunet.2026.108899? Name: ScienceDirect (all content) (s7799221) Category: fullText Text: View record from ScienceDirect MouseOverText: View record from ScienceDirect |
|---|---|
| Header | DbId: cmedm DbLabel: MEDLINE An: 41932121 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AU" term="%22Yan+C%22">Yan C</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: yancaixia@xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Jiao+M%22">Jiao M</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: jmy0406@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Xue+N%22">Xue N</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: xuenuohan@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Zhang+W%22">Zhang W</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: zhangwzh@xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Wang+J%22">Wang J</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: uguisu@stu.xjtu.edu.cn.<br /><searchLink fieldCode="AU" term="%22Chang+X%22">Chang X</searchLink>; Department of Electronic Engineering and Information Science, University of Science and Technology of China, China. Electronic address: cxj273@gmail.com.<br /><searchLink fieldCode="AU" term="%22Tian+F%22">Tian F</searchLink>; School of Computer Science and Technology, Xi'an Jiaotong University, China; Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, Xi'an Jiaotong University, China. Electronic address: fengtian@mail.xjtu.edu.cn. – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%228805018%22">Neural networks : the official journal of the International Neural Network Society</searchLink> [Neural Netw] 2026 Sep; Vol. 201, pp. 108899. <i>Date of Electronic Publication: </i>2026 Mar 25. – Name: TypePub Label: Publication Type Group: TypPub Data: Journal Article – Name: Language Label: Language Group: Lang Data: English – Name: TitleSource Label: Journal Info Group: Src Data: <i>Publisher: </i><searchLink fieldCode="PB" term="%22Pergamon+Press%22">Pergamon Press </searchLink><i>Country of Publication: </i>United States <i>NLM ID: </i>8805018 <i>Publication Model: </i>Print-Electronic <i>Cited Medium: </i>Internet <i>ISSN: </i>1879-2782 (Electronic) <i>Linking ISSN: </i><searchLink fieldCode="IS" term="%2208936080%22">08936080 </searchLink><i>NLM ISO Abbreviation: </i>Neural Netw <i>Subsets: </i>MEDLINE – Name: PublisherInfo Label: Imprint Name(s) Group: PubInfo Data: <i>Original Publication</i>: New York : Pergamon Press, [c1988- – Name: SubjectMESH Label: MeSH Terms Group: Su Data: <searchLink fieldCode="MM" term="%22Detection+Algorithms%22">Detection Algorithms*</searchLink> <br /><searchLink fieldCode="MM" term="%22Generative+Adversarial+Networks%22">Generative Adversarial Networks*</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator's embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.<br /> (Copyright © 2026 Elsevier Ltd. All rights reserved.) – Name: Abstract Label: Competing Interests Group: Ab Data: Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. – Name: SubjectMinor Label: Contributed Indexing Group: Data: <i>Keywords: </i>Generative adversarial networks; Vision-language distillation; Zero-shot object detection – Name: DateEntry Label: Entry Date(s) Group: Date Data: <i>Date Created: </i>20260403 <i>Date Completed: </i>20260613 <i>Latest Revision: </i>20260619 – Name: DateUpdate Label: Update Code Group: Date Data: 20260620 – Name: DOI Label: DOI Group: ID Data: 10.1016/j.neunet.2026.108899 – Name: AN Label: PMID Group: ID Data: 41932121 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=cmedm&AN=41932121 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.neunet.2026.108899 Languages: – Code: eng Text: English PhysicalDescription: Pagination: StartPage: 108899 Subjects: – SubjectFull: Detection Algorithms Type: general – SubjectFull: Generative Adversarial Networks Type: general Titles: – TitleFull: VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Yan C – PersonEntity: Name: NameFull: Jiao M – PersonEntity: Name: NameFull: Xue N – PersonEntity: Name: NameFull: Zhang W – PersonEntity: Name: NameFull: Wang J – PersonEntity: Name: NameFull: Chang X – PersonEntity: Name: NameFull: Tian F IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: 2026 Sep Type: published Y: 2026 Identifiers: – Type: issn-electronic Value: 1879-2782 Numbering: – Type: volume Value: 201 Titles: – TitleFull: Neural networks : the official journal of the International Neural Network Society Type: main |
| ResultId | 1 |