Dissertation/ Thesis

Integration and processing of large-scale biomedical data

Bibliographic Details
Title: Integration and processing of large-scale biomedical data
Authors: Zhang, Wenhua, 张闻华
Contributors: Pan, J, Wang, WP
Publisher Information: The University of Hong Kong (Pokfulam, Hong Kong)
Publication Year: 2023
Collection: University of Hong Kong: HKU Scholars Hub
Subject Terms: Medical informatics - Data processing, Biomedical engineering - Data processing
Description: With the improvements in data collection methods, high-quality data abounds in medical imaging and other fields. While many algorithms have emerged to detect, segment, or classify the data, few have proposed methods to re-organize or mine them. It is important to develop approaches to dive into the data and fully exploit them. This thesis tackles three dataset integration and processing problems of large-scale biomedical data: labeled dataset merging, unlabeled dataset self-supervised training, and scalable volumetric data mesh generating. The first part of this thesis addresses the problem of integrating inconsistent datasets. A large number of labeled data is required to train effective nucleus classification models. However, it is challenging to label a large-scale nucleus classification dataset, considering that high-quality labeling requires specific domain knowledge and tremendous efforts. In addition, existing public datasets are often inconsistently labeled. Due to this inconsistency, conventional models tend to work independently to infer their classification results, thus limiting the classification performance. To fully utilize all annotated datasets, we propose a method to integrate all the available annotated datasets. Specifically, we formulate the problem as a multi-label problem with missing labels. Thus, we can utilize all the datasets in a unified framework. Besides the substantial improvement compared to other methods, our result dataset also has a uniform format which can help future research on nucleus classification. The second part of this thesis addresses the problem of representation learning for nucleus instance classification. Unlike the limited scale of annotated data, unlabeled data is usually of large scale. Thus, we aim to design a self-supervised method for representation learning on unlabeled datasets to alleviate the burden of data annotation. Moreover, previous methods often downplay the contextual information that is critical for classification. To explicitly provide the ...
Document Type: doctoral or postdoctoral thesis
Language: English
Relation: HKU Theses Online (HKUTO); 991044705906303414; https://hub.hku.hk/handle/10722/328917
Availability: https://hub.hku.hk/handle/10722/328917
Rights: The author retains all proprietary rights, (such as patent rights) and the right to use in future works. ; This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Accession Number: edsbas.3FAC5B27
Database: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://hub.hku.hk/handle/10722/328917#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.3FAC5B27
RelevancyScore: 851
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 851.480346679688
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Integration and processing of large-scale biomedical data
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Wenhua%22">Zhang, Wenhua</searchLink><br /><searchLink fieldCode="AR" term="%22张闻华%22">张闻华</searchLink>
– Name: Author
  Label: Contributors
  Group: Au
  Data: Pan, J<br />Wang, WP
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: The University of Hong Kong (Pokfulam, Hong Kong)
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2023
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: University of Hong Kong: HKU Scholars Hub
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Medical+informatics+-+Data+processing%22">Medical informatics - Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Biomedical+engineering+-+Data+processing%22">Biomedical engineering - Data processing</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: With the improvements in data collection methods, high-quality data abounds in medical imaging and other fields. While many algorithms have emerged to detect, segment, or classify the data, few have proposed methods to re-organize or mine them. It is important to develop approaches to dive into the data and fully exploit them. This thesis tackles three dataset integration and processing problems of large-scale biomedical data: labeled dataset merging, unlabeled dataset self-supervised training, and scalable volumetric data mesh generating. The first part of this thesis addresses the problem of integrating inconsistent datasets. A large number of labeled data is required to train effective nucleus classification models. However, it is challenging to label a large-scale nucleus classification dataset, considering that high-quality labeling requires specific domain knowledge and tremendous efforts. In addition, existing public datasets are often inconsistently labeled. Due to this inconsistency, conventional models tend to work independently to infer their classification results, thus limiting the classification performance. To fully utilize all annotated datasets, we propose a method to integrate all the available annotated datasets. Specifically, we formulate the problem as a multi-label problem with missing labels. Thus, we can utilize all the datasets in a unified framework. Besides the substantial improvement compared to other methods, our result dataset also has a uniform format which can help future research on nucleus classification. The second part of this thesis addresses the problem of representation learning for nucleus instance classification. Unlike the limited scale of annotated data, unlabeled data is usually of large scale. Thus, we aim to design a self-supervised method for representation learning on unlabeled datasets to alleviate the burden of data annotation. Moreover, previous methods often downplay the contextual information that is critical for classification. To explicitly provide the ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: doctoral or postdoctoral thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: HKU Theses Online (HKUTO); 991044705906303414; https://hub.hku.hk/handle/10722/328917
– Name: URL
  Label: Availability
  Group: URL
  Data: https://hub.hku.hk/handle/10722/328917
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: The author retains all proprietary rights, (such as patent rights) and the right to use in future works. ; This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.3FAC5B27
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.3FAC5B27
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Medical informatics - Data processing
        Type: general
      – SubjectFull: Biomedical engineering - Data processing
        Type: general
    Titles:
      – TitleFull: Integration and processing of large-scale biomedical data
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhang, Wenhua
      – PersonEntity:
          Name:
            NameFull: 张闻华
      – PersonEntity:
          Name:
            NameFull: Pan, J
      – PersonEntity:
          Name:
            NameFull: Wang, WP
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1