Academic Journal
Primitive operations in computing massive data revisited
| Τίτλος: | Primitive operations in computing massive data revisited |
|---|---|
| Συνεισφορές: | Wei, Hao (author.), Yu, Jeffrey Xu (thesis advisor.), Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.) |
| Έτος έκδοσης: | 2017 |
| Συλλογή: | The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏 |
| Θεματικοί όροι: | Graph theory--Data processing, Computer algorithms, QA166 .W425 2017eb |
| Περιγραφή: | Ph.D. ; With the massive size of data from various sources, achieving high efficiency for even the primitive operations in big data can be challenging. In this thesis, we investigate the primitive operations in processing two widely used data types, graph and string. Graph can model the online social network, hyperlink graph, and knowledge graph etc., while string is widely used to represent textual data including document, message, and DNA sequence. ; We first investigate how to optimize the graph computing in a general way. During our research on various kinds of graph algorithms, we observe that CPU cache perfor mance is the key to the efficiency of graph computing. Motivated by this observation, we analyze the common access pattern of various primitive graph operations, such as Breadth First Search, Depth First Search, PageRank, etc., and formulate the graph ordering problem that aims to utilize the graph data locality to reduce the CPU cache miss ratio based on the common access pattern. We design efficient algorithms to compute the graph ordering and the experiments show that the optimized graph ordering can improve the CPU cache performance for most graph operations. Note that the graph algorithms can get speed up from the graph ordering without the requirement of modifying its implementation and data structures. ; We further study the probabilistic data structure and design hash-based labeling schemes for efficiently answering reachability query in graph data and similarity search query in string data. Graph reachability query is a primitive graph operation that checks whether a node can reach another node over the graph. We propose IP+ labeling, a novel randomized labeling approach designed based on min-wise independent permutation technique. The index size and index construction time of the proposed IP+ label are kept to be linear with the graph size. The IP+ label can guarantee that two unreachable nodes can be answered directly by their IP+ labels with a high probability and no graph traversal is ... |
| Τύπος εγγράφου: | text |
| Περιγραφή αρχείου: | electronic resource; remote; 1 online resource (xiii, 174 leaves) : illustrations (some color); computer; online resource |
| Γλώσσα: | English Chinese |
| Διαθεσιμότητα: | https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US https://repository.lib.cuhk.edu.hk/en/item/cuhk-1839285 |
| Rights: | Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/) |
| Αριθμός Καταχώρησης: | edsbas.83DC5016 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.83DC5016 RelevancyScore: 865 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 864.807434082031 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Primitive operations in computing massive data revisited – Name: Author Label: Contributors Group: Au Data: Wei, Hao (author.)<br />Yu, Jeffrey Xu (thesis advisor.)<br />Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.) – Name: DatePubCY Label: Publication Year Group: Date Data: 2017 – Name: Subset Label: Collection Group: HoldingsInfo Data: The Chinese University of Hong Kong: CUHK Digital Repository / 香港中文大學數碼典藏 – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Graph+theory--Data+processing%22">Graph theory--Data processing</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+algorithms%22">Computer algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22QA166+%2EW425+2017eb%22">QA166 .W425 2017eb</searchLink> – Name: Abstract Label: Description Group: Ab Data: Ph.D. ; With the massive size of data from various sources, achieving high efficiency for even the primitive operations in big data can be challenging. In this thesis, we investigate the primitive operations in processing two widely used data types, graph and string. Graph can model the online social network, hyperlink graph, and knowledge graph etc., while string is widely used to represent textual data including document, message, and DNA sequence. ; We first investigate how to optimize the graph computing in a general way. During our research on various kinds of graph algorithms, we observe that CPU cache perfor mance is the key to the efficiency of graph computing. Motivated by this observation, we analyze the common access pattern of various primitive graph operations, such as Breadth First Search, Depth First Search, PageRank, etc., and formulate the graph ordering problem that aims to utilize the graph data locality to reduce the CPU cache miss ratio based on the common access pattern. We design efficient algorithms to compute the graph ordering and the experiments show that the optimized graph ordering can improve the CPU cache performance for most graph operations. Note that the graph algorithms can get speed up from the graph ordering without the requirement of modifying its implementation and data structures. ; We further study the probabilistic data structure and design hash-based labeling schemes for efficiently answering reachability query in graph data and similarity search query in string data. Graph reachability query is a primitive graph operation that checks whether a node can reach another node over the graph. We propose IP+ labeling, a novel randomized labeling approach designed based on min-wise independent permutation technique. The index size and index construction time of the proposed IP+ label are kept to be linear with the graph size. The IP+ label can guarantee that two unreachable nodes can be answered directly by their IP+ labels with a high probability and no graph traversal is ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: text – Name: Format Label: File Description Group: SrcInfo Data: electronic resource; remote; 1 online resource (xiii, 174 leaves) : illustrations (some color); computer; online resource – Name: Language Label: Language Group: Lang Data: English<br />Chinese – Name: URL Label: Availability Group: URL Data: https://julac.hosted.exlibrisgroup.com/primo-explore/search?query=addsrcrid,exact,991039528240203407,AND&tab=default_tab&search_scope=All&vid=CUHK&mode=advanced&lang=en_US<br />https://repository.lib.cuhk.edu.hk/en/item/cuhk-1839285 – Name: Copyright Label: Rights Group: Cpyrght Data: Use of this resource is governed by the terms and conditions of the Creative Commons "Attribution-NonCommercial-NoDerivatives 4.0 International" License (http://creativecommons.org/licenses/by-nc-nd/4.0/) – Name: AN Label: Accession Number Group: ID Data: edsbas.83DC5016 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.83DC5016 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English – Text: Chinese Subjects: – SubjectFull: Graph theory--Data processing Type: general – SubjectFull: Computer algorithms Type: general – SubjectFull: QA166 .W425 2017eb Type: general Titles: – TitleFull: Primitive operations in computing massive data revisited Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Wei, Hao (author.) – PersonEntity: Name: NameFull: Yu, Jeffrey Xu (thesis advisor.) – PersonEntity: Name: NameFull: Chinese University of Hong Kong Graduate School. Division of Systems Engineering and Engineering Management. (degree granting institution.) IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2017 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |