Dissertation/ Thesis
Unified Approaches for Multi-Task Vision-Language Interactions
| Title: | Unified Approaches for Multi-Task Vision-Language Interactions |
|---|---|
| Authors: | You, Haoxuan |
| Publication Year: | 2024 |
| Collection: | Columbia University: Academic Commons |
| Subject Terms: | Artificial intelligence, Computer multitasking, Computer vision--Computer programs, Questions and answers--Computer programs |
| Description: | Vision and Language are two major modalities that humans rely on to perceive the environment and understand the world. Recent advances in Artificial Intelligence (AI) facilitate the development of a variety of vision-language tasks derived from diverse multimodal interactions in daily life, such as image captioning, image-text matching, visual question answering (VQA), text-to-image generation, etc. Despite the remarkable performance, most previous state-of-the-art models are merely specialized for a single vision-language task, which lack generalizability across multiple tasks. Additionally, those specialized models sophisticate the algorithm designs and bring redundancy to model deployment when dealing with complex scenes. In this study, we investigate developing unified approaches capable of solving various vision-language interactions in a multi-task manner. We argue that unified multi-task methods could enjoy several potential advantages: (1) A unified framework for multiple tasks can reduce human efforts in designing different models for different tasks; (2) Reusing and sharing parameters across tasks can improve efficiency; (3) Some tasks may be complementary to other tasks so that multi-tasking can boost the performance; (4) They can deal with the complex tasks that need a joint collaborating of multiple basic tasks and enable new applications. In the first part of this thesis, we explore unified multi-task models with the goal of sharing and reusing as many parameters as possible between different tasks. We started with unifying many vision-language question-answering tasks, such as visual entailment, outside-knowledge VQA, and visual commonsense reasoning, in a simple iterative divide-and-conquer framework. Specifically, it iteratively decomposes the original text question into sub-question, solves each sub-question, and derives the answer to the original question, which can uniformly handle reasoning of various types and semantics levels within one framework. In the next work, we take one step further ... |
| Document Type: | thesis |
| Language: | English |
| DOI: | 10.7916/n044-ra77 |
| Availability: | https://doi.org/10.7916/n044-ra77 |
| Accession Number: | edsbas.2B83F9CB |
| Database: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://doi.org/10.7916/n044-ra77# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.2B83F9CB RelevancyScore: 872 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 871.7080078125 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Unified Approaches for Multi-Task Vision-Language Interactions – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22You%2C+Haoxuan%22">You, Haoxuan</searchLink> – Name: DatePubCY Label: Publication Year Group: Date Data: 2024 – Name: Subset Label: Collection Group: HoldingsInfo Data: Columbia University: Academic Commons – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+multitasking%22">Computer multitasking</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+vision--Computer+programs%22">Computer vision--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Questions+and+answers--Computer+programs%22">Questions and answers--Computer programs</searchLink> – Name: Abstract Label: Description Group: Ab Data: Vision and Language are two major modalities that humans rely on to perceive the environment and understand the world. Recent advances in Artificial Intelligence (AI) facilitate the development of a variety of vision-language tasks derived from diverse multimodal interactions in daily life, such as image captioning, image-text matching, visual question answering (VQA), text-to-image generation, etc. Despite the remarkable performance, most previous state-of-the-art models are merely specialized for a single vision-language task, which lack generalizability across multiple tasks. Additionally, those specialized models sophisticate the algorithm designs and bring redundancy to model deployment when dealing with complex scenes. In this study, we investigate developing unified approaches capable of solving various vision-language interactions in a multi-task manner. We argue that unified multi-task methods could enjoy several potential advantages: (1) A unified framework for multiple tasks can reduce human efforts in designing different models for different tasks; (2) Reusing and sharing parameters across tasks can improve efficiency; (3) Some tasks may be complementary to other tasks so that multi-tasking can boost the performance; (4) They can deal with the complex tasks that need a joint collaborating of multiple basic tasks and enable new applications. In the first part of this thesis, we explore unified multi-task models with the goal of sharing and reusing as many parameters as possible between different tasks. We started with unifying many vision-language question-answering tasks, such as visual entailment, outside-knowledge VQA, and visual commonsense reasoning, in a simple iterative divide-and-conquer framework. Specifically, it iteratively decomposes the original text question into sub-question, solves each sub-question, and derives the answer to the original question, which can uniformly handle reasoning of various types and semantics levels within one framework. In the next work, we take one step further ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: thesis – Name: Language Label: Language Group: Lang Data: English – Name: DOI Label: DOI Group: ID Data: 10.7916/n044-ra77 – Name: URL Label: Availability Group: URL Data: https://doi.org/10.7916/n044-ra77 – Name: AN Label: Accession Number Group: ID Data: edsbas.2B83F9CB |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.2B83F9CB |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.7916/n044-ra77 Languages: – Text: English Subjects: – SubjectFull: Artificial intelligence Type: general – SubjectFull: Computer multitasking Type: general – SubjectFull: Computer vision--Computer programs Type: general – SubjectFull: Questions and answers--Computer programs Type: general Titles: – TitleFull: Unified Approaches for Multi-Task Vision-Language Interactions Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: You, Haoxuan IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2024 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |