Academic Journal

ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data

Bibliographic Details
Title: ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data
Authors: Yang, Zhiliang, Xu, Yuyang, Yao, Minhao, Wang, Gao, Liu, Zhonghua
Publication Year: 2023
Collection: Columbia University: Academic Commons
Subject Terms: Nucleotide sequence--Data processing, Genomes--Data processing, Python (Computer program language)
Description: Background Large-scale multi-ethnic DNA sequencing data is increasingly available owing to decreasing cost of modern sequencing technologies. Inference of the population structure with such sequencing data is fundamentally important. However, the ultra-dimensionality and complicated linkage disequilibrium patterns across the whole genome make it challenging to infer population structure using traditional principal component analysis based methods and software. Results We present the ERStruct Python Package, which enables the inference of population structure using whole-genome sequencing data. By leveraging parallel computing and GPU acceleration, our package achieves significant improvements in the speed of matrix operations for large-scale data. Additionally, our package features adaptive data splitting capabilities to facilitate computation on GPUs with limited memory. Conclusion Our Python package ERStruct is an efficient and user-friendly tool for estimating the number of top informative principal components that capture population structure from whole genome sequencing data.
Document Type: article in journal/newspaper
Language: English
DOI: 10.7916/kx4x-e654
Availability: https://doi.org/10.7916/kx4x-e654
Accession Number: edsbas.ECF4CBE0
Database: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://doi.org/10.7916/kx4x-e654#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.ECF4CBE0
RelevancyScore: 929
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 929.480346679688
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Yang%2C+Zhiliang%22">Yang, Zhiliang</searchLink><br /><searchLink fieldCode="AR" term="%22Xu%2C+Yuyang%22">Xu, Yuyang</searchLink><br /><searchLink fieldCode="AR" term="%22Yao%2C+Minhao%22">Yao, Minhao</searchLink><br /><searchLink fieldCode="AR" term="%22Wang%2C+Gao%22">Wang, Gao</searchLink><br /><searchLink fieldCode="AR" term="%22Liu%2C+Zhonghua%22">Liu, Zhonghua</searchLink>
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2023
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Columbia University: Academic Commons
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Nucleotide+sequence--Data+processing%22">Nucleotide sequence--Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Genomes--Data+processing%22">Genomes--Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Python+%28Computer+program+language%29%22">Python (Computer program language)</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Background Large-scale multi-ethnic DNA sequencing data is increasingly available owing to decreasing cost of modern sequencing technologies. Inference of the population structure with such sequencing data is fundamentally important. However, the ultra-dimensionality and complicated linkage disequilibrium patterns across the whole genome make it challenging to infer population structure using traditional principal component analysis based methods and software. Results We present the ERStruct Python Package, which enables the inference of population structure using whole-genome sequencing data. By leveraging parallel computing and GPU acceleration, our package achieves significant improvements in the speed of matrix operations for large-scale data. Additionally, our package features adaptive data splitting capabilities to facilitate computation on GPUs with limited memory. Conclusion Our Python package ERStruct is an efficient and user-friendly tool for estimating the number of top informative principal components that capture population structure from whole genome sequencing data.
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: article in journal/newspaper
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.7916/kx4x-e654
– Name: URL
  Label: Availability
  Group: URL
  Data: https://doi.org/10.7916/kx4x-e654
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.ECF4CBE0
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.ECF4CBE0
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.7916/kx4x-e654
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Nucleotide sequence--Data processing
        Type: general
      – SubjectFull: Genomes--Data processing
        Type: general
      – SubjectFull: Python (Computer program language)
        Type: general
    Titles:
      – TitleFull: ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Yang, Zhiliang
      – PersonEntity:
          Name:
            NameFull: Xu, Yuyang
      – PersonEntity:
          Name:
            NameFull: Yao, Minhao
      – PersonEntity:
          Name:
            NameFull: Wang, Gao
      – PersonEntity:
          Name:
            NameFull: Liu, Zhonghua
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1