Academic Journal

Pixel level image understanding with deep learning

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Pixel level image understanding with deep learning
Συνεισφορές: Qi, Xiaojuan (author.), Jia, Jiaya (thesis advisor.), Chinese University of Hong Kong Graduate School. Division of Computer Science and Engineering. (degree granting institution.)
Έτος έκδοσης: 2018
Συλλογή: The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏
Θεματικοί όροι: Image processing--Data processing, Computer vision, Machine learning, TA1637 .Q52 2018eb
Περιγραφή: Ph.D. ; Pixel-level image understanding and synthesis are very important problems in computer vision since they provide the most comprehensive and detailed understanding of the visual world. In this thesis, we study these problems in RGB image domain, semantic domain and geometric domain. ; The first problem we address is, how to extract pixel-level semantic knowledge from RGB images. We address it by developing an object clique potential for semantic segmentation. Our object clique potential addresses the misclassified object-part issues arising in solutions based on fully-convolutional networks. Our object clique set, compared to that yielded from segmentproposal based approaches, is with a significantly smaller size, making our method consume notably less computation. Regarding system design and model formation, our object clique potential can be regarded as a functional complement to local-appearance-based CRF models and works in synergy with these effective approaches for further performance improvement. ; However, training neural network for pixel-level semantic understanding is data hungry. So the second problem we tackle is to learn pixel-level semantic information with only image-level supervision. Our method unifies semantic segmentation and object localization with important proposal aggregation and selection modules. They greatly reduce the notorious error accumulation problem that commonly arises in weakly supervised learning. Our proposed training algorithm progressively improves segmentation performance with augmented feedback in iterations. ; When human infer the semantics of the visual world, we not only use the RGB information, but also take the geometry into consideration. Thus, the third problem we handle is to incorporate geometry data for semantic segmentation. To be more specific, we study RGBD semantic segmentation. RGBD semantic segmentation requires joint reasoning about 2D appearance and 3D geometric information. We propose a 3D graph neural network (3DGNN) that builds a k-nearest ...
Τύπος εγγράφου: text
Περιγραφή αρχείου: electronic resource; remote; 1 online resource (xv, 124 leaves) : illustrations (some color); computer; online resource
Γλώσσα: English
Chinese
Relation: cuhk:2187981; local: ETD920200132; local: AAI13837880; local: 991039750259003407
Διαθεσιμότητα: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039750259003407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US
https://repository.lib.cuhk.edu.hk/en/item/cuhk-2187981
Rights: Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/)
Αριθμός Καταχώρησης: edsbas.90C7699
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039750259003407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.90C7699
RelevancyScore: 872
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 872.436645507813
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Pixel level image understanding with deep learning
– Name: Author
  Label: Contributors
  Group: Au
  Data: Qi, Xiaojuan (author.)<br />Jia, Jiaya (thesis advisor.)<br />Chinese University of Hong Kong Graduate School. Division of Computer Science and Engineering. (degree granting institution.)
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2018
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Image+processing--Data+processing%22">Image processing--Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+vision%22">Computer vision</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22TA1637+%2EQ52+2018eb%22">TA1637 .Q52 2018eb</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Ph.D. ; Pixel-level image understanding and synthesis are very important problems in computer vision since they provide the most comprehensive and detailed understanding of the visual world. In this thesis, we study these problems in RGB image domain, semantic domain and geometric domain. ; The first problem we address is, how to extract pixel-level semantic knowledge from RGB images. We address it by developing an object clique potential for semantic segmentation. Our object clique potential addresses the misclassified object-part issues arising in solutions based on fully-convolutional networks. Our object clique set, compared to that yielded from segmentproposal based approaches, is with a significantly smaller size, making our method consume notably less computation. Regarding system design and model formation, our object clique potential can be regarded as a functional complement to local-appearance-based CRF models and works in synergy with these effective approaches for further performance improvement. ; However, training neural network for pixel-level semantic understanding is data hungry. So the second problem we tackle is to learn pixel-level semantic information with only image-level supervision. Our method unifies semantic segmentation and object localization with important proposal aggregation and selection modules. They greatly reduce the notorious error accumulation problem that commonly arises in weakly supervised learning. Our proposed training algorithm progressively improves segmentation performance with augmented feedback in iterations. ; When human infer the semantics of the visual world, we not only use the RGB information, but also take the geometry into consideration. Thus, the third problem we handle is to incorporate geometry data for semantic segmentation. To be more specific, we study RGBD semantic segmentation. RGBD semantic segmentation requires joint reasoning about 2D appearance and 3D geometric information. We propose a 3D graph neural network (3DGNN) that builds a k-nearest ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: text
– Name: Format
  Label: File Description
  Group: SrcInfo
  Data: electronic resource; remote; 1 online resource (xv, 124 leaves) : illustrations (some color); computer; online resource
– Name: Language
  Label: Language
  Group: Lang
  Data: English<br />Chinese
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: cuhk:2187981; local: ETD920200132; local: AAI13837880; local: 991039750259003407
– Name: URL
  Label: Availability
  Group: URL
  Data: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039750259003407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US<br />https://repository.lib.cuhk.edu.hk/en/item/cuhk-2187981
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/)
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.90C7699
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.90C7699
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
      – Text: Chinese
    Subjects:
      – SubjectFull: Image processing--Data processing
        Type: general
      – SubjectFull: Computer vision
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: TA1637 .Q52 2018eb
        Type: general
    Titles:
      – TitleFull: Pixel level image understanding with deep learning
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Qi, Xiaojuan (author.)
      – PersonEntity:
          Name:
            NameFull: Jia, Jiaya (thesis advisor.)
      – PersonEntity:
          Name:
            NameFull: Chinese University of Hong Kong Graduate School. Division of Computer Science and Engineering. (degree granting institution.)
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2018
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1