Academic Journal

Primitive operations in computing massive data revisited

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Primitive operations in computing massive data revisited
Συνεισφορές: Wei, Hao (author.), Yu, Jeffrey Xu (thesis advisor.), Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.)
Έτος έκδοσης: 2017
Συλλογή: The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏
Θεματικοί όροι: Graph theory--Data processing, Computer algorithms, QA166 .W425 2017eb
Περιγραφή: Ph.D. ; With the massive size of data from various sources, achieving high efficiency for even the primitive operations in big data can be challenging. In this thesis, we investigate the primitive operations in processing two widely used data types, graph and string. Graph can model the online social network, hyperlink graph, and knowledge graph etc., while string is widely used to represent textual data including document, message, and DNA sequence. ; We first investigate how to optimize the graph computing in a general way. During our research on various kinds of graph algorithms, we observe that CPU cache perfor mance is the key to the efficiency of graph computing. Motivated by this observation, we analyze the common access pattern of various primitive graph operations, such as Breadth First Search, Depth First Search, PageRank, etc., and formulate the graph ordering problem that aims to utilize the graph data locality to reduce the CPU cache miss ratio based on the common access pattern. We design efficient algorithms to compute the graph ordering and the experiments show that the optimized graph ordering can improve the CPU cache performance for most graph operations. Note that the graph algorithms can get speed up from the graph ordering without the requirement of modifying its implementation and data structures. ; We further study the probabilistic data structure and design hash-based labeling schemes for efficiently answering reachability query in graph data and similarity search query in string data. Graph reachability query is a primitive graph operation that checks whether a node can reach another node over the graph. We propose IP+ labeling, a novel randomized labeling approach designed based on min-wise independent permutation technique. The index size and index construction time of the proposed IP+ label are kept to be linear with the graph size. The IP+ label can guarantee that two unreachable nodes can be answered directly by their IP+ labels with a high probability and no graph traversal is ...
Τύπος εγγράφου: text
Περιγραφή αρχείου: electronic resource; remote; 1 online resource (xiii, 174 leaves) : illustrations (some color); computer; online resource
Γλώσσα: English
Chinese
Διαθεσιμότητα: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US
https://repository.lib.cuhk.edu.hk/en/item/cuhk-1839285
Rights: Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/)
Αριθμός Καταχώρησης: edsbas.83DC5016
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.83DC5016
RelevancyScore: 865
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 864.807434082031
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Primitive operations in computing massive data revisited
– Name: Author
  Label: Contributors
  Group: Au
  Data: Wei, Hao (author.)<br />Yu, Jeffrey Xu (thesis advisor.)<br />Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.)
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2017
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Graph+theory--Data+processing%22">Graph theory--Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+algorithms%22">Computer algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22QA166+%2EW425+2017eb%22">QA166 .W425 2017eb</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: Ph.D. ; With the massive size of data from various sources, achieving high efficiency for even the primitive operations in big data can be challenging. In this thesis, we investigate the primitive operations in processing two widely used data types, graph and string. Graph can model the online social network, hyperlink graph, and knowledge graph etc., while string is widely used to represent textual data including document, message, and DNA sequence. ; We first investigate how to optimize the graph computing in a general way. During our research on various kinds of graph algorithms, we observe that CPU cache perfor mance is the key to the efficiency of graph computing. Motivated by this observation, we analyze the common access pattern of various primitive graph operations, such as Breadth First Search, Depth First Search, PageRank, etc., and formulate the graph ordering problem that aims to utilize the graph data locality to reduce the CPU cache miss ratio based on the common access pattern. We design efficient algorithms to compute the graph ordering and the experiments show that the optimized graph ordering can improve the CPU cache performance for most graph operations. Note that the graph algorithms can get speed up from the graph ordering without the requirement of modifying its implementation and data structures. ; We further study the probabilistic data structure and design hash-based labeling schemes for efficiently answering reachability query in graph data and similarity search query in string data. Graph reachability query is a primitive graph operation that checks whether a node can reach another node over the graph. We propose IP+ labeling, a novel randomized labeling approach designed based on min-wise independent permutation technique. The index size and index construction time of the proposed IP+ label are kept to be linear with the graph size. The IP+ label can guarantee that two unreachable nodes can be answered directly by their IP+ labels with a high probability and no graph traversal is ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: text
– Name: Format
  Label: File Description
  Group: SrcInfo
  Data: electronic resource; remote; 1 online resource (xiii, 174 leaves) : illustrations (some color); computer; online resource
– Name: Language
  Label: Language
  Group: Lang
  Data: English<br />Chinese
– Name: URL
  Label: Availability
  Group: URL
  Data: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US<br />https://repository.lib.cuhk.edu.hk/en/item/cuhk-1839285
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/)
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.83DC5016
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.83DC5016
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
      – Text: Chinese
    Subjects:
      – SubjectFull: Graph theory--Data processing
        Type: general
      – SubjectFull: Computer algorithms
        Type: general
      – SubjectFull: QA166 .W425 2017eb
        Type: general
    Titles:
      – TitleFull: Primitive operations in computing massive data revisited
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Wei, Hao (author.)
      – PersonEntity:
          Name:
            NameFull: Yu, Jeffrey Xu (thesis advisor.)
      – PersonEntity:
          Name:
            NameFull: Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.)
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1