Dissertation/ Thesis

Unified Approaches for Multi-Task Vision-Language Interactions

Bibliographic Details
Title: Unified Approaches for Multi-Task Vision-Language Interactions
Authors: You, Haoxuan
Publication Year: 2024
Collection: Columbia University: Academic Commons
Subject Terms: Artificial intelligence, Computer multitasking, Computer vision--Computer programs, Questions and answers--Computer programs
Description: Vision and Language are two major modalities that humans rely on to perceive the environment and understand the world. Recent advances in Artificial Intelligence (AI) facilitate the development of a variety of vision-language tasks derived from diverse multimodal interactions in daily life, such as image captioning, image-text matching, visual question answering (VQA), text-to-image generation, etc. Despite the remarkable performance, most previous state-of-the-art models are merely specialized for a single vision-language task, which lack generalizability across multiple tasks. Additionally, those specialized models sophisticate the algorithm designs and bring redundancy to model deployment when dealing with complex scenes. In this study, we investigate developing unified approaches capable of solving various vision-language interactions in a multi-task manner. We argue that unified multi-task methods could enjoy several potential advantages: (1) A unified framework for multiple tasks can reduce human efforts in designing different models for different tasks; (2) Reusing and sharing parameters across tasks can improve efficiency; (3) Some tasks may be complementary to other tasks so that multi-tasking can boost the performance; (4) They can deal with the complex tasks that need a joint collaborating of multiple basic tasks and enable new applications. In the first part of this thesis, we explore unified multi-task models with the goal of sharing and reusing as many parameters as possible between different tasks. We started with unifying many vision-language question-answering tasks, such as visual entailment, outside-knowledge VQA, and visual commonsense reasoning, in a simple iterative divide-and-conquer framework. Specifically, it iteratively decomposes the original text question into sub-question, solves each sub-question, and derives the answer to the original question, which can uniformly handle reasoning of various types and semantics levels within one framework. In the next work, we take one step further ...
Document Type: thesis
Language: English
DOI: 10.7916/n044-ra77
Availability: https://doi.org/10.7916/n044-ra77
Accession Number: edsbas.2B83F9CB
Database: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://doi.org/10.7916/n044-ra77#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.2B83F9CB
RelevancyScore: 872
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 871.7080078125
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Unified Approaches for Multi-Task Vision-Language Interactions
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22You%2C+Haoxuan%22">You, Haoxuan</searchLink>
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2024
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Columbia University: Academic Commons
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+multitasking%22">Computer multitasking</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+vision--Computer+programs%22">Computer vision--Computer programs</searchLink><br /><searchLink fieldCode="DE" term="%22Questions+and+answers--Computer+programs%22">Questions and answers--Computer programs</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Vision and Language are two major modalities that humans rely on to perceive the environment and understand the world. Recent advances in Artificial Intelligence (AI) facilitate the development of a variety of vision-language tasks derived from diverse multimodal interactions in daily life, such as image captioning, image-text matching, visual question answering (VQA), text-to-image generation, etc. Despite the remarkable performance, most previous state-of-the-art models are merely specialized for a single vision-language task, which lack generalizability across multiple tasks. Additionally, those specialized models sophisticate the algorithm designs and bring redundancy to model deployment when dealing with complex scenes. In this study, we investigate developing unified approaches capable of solving various vision-language interactions in a multi-task manner. We argue that unified multi-task methods could enjoy several potential advantages: (1) A unified framework for multiple tasks can reduce human efforts in designing different models for different tasks; (2) Reusing and sharing parameters across tasks can improve efficiency; (3) Some tasks may be complementary to other tasks so that multi-tasking can boost the performance; (4) They can deal with the complex tasks that need a joint collaborating of multiple basic tasks and enable new applications. In the first part of this thesis, we explore unified multi-task models with the goal of sharing and reusing as many parameters as possible between different tasks. We started with unifying many vision-language question-answering tasks, such as visual entailment, outside-knowledge VQA, and visual commonsense reasoning, in a simple iterative divide-and-conquer framework. Specifically, it iteratively decomposes the original text question into sub-question, solves each sub-question, and derives the answer to the original question, which can uniformly handle reasoning of various types and semantics levels within one framework. In the next work, we take one step further ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.7916/n044-ra77
– Name: URL
  Label: Availability
  Group: URL
  Data: https://doi.org/10.7916/n044-ra77
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.2B83F9CB
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.2B83F9CB
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.7916/n044-ra77
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Computer multitasking
        Type: general
      – SubjectFull: Computer vision--Computer programs
        Type: general
      – SubjectFull: Questions and answers--Computer programs
        Type: general
    Titles:
      – TitleFull: Unified Approaches for Multi-Task Vision-Language Interactions
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: You, Haoxuan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1