Dissertation/ Thesis
Cost-effective batch-based migration strategies for NewSQL-based big data systems
| Title: | Cost-effective batch-based migration strategies for NewSQL-based big data systems |
|---|---|
| Authors: | Vadlamudi, Naveen Kumar, University of Lethbridge. Faculty of Arts and Science |
| Contributors: | Osborn, Wendy |
| Publisher Information: | University of Lethbridge, Dept. of Mathematics and Computer Science Department of Mathematics and Computer Science Arts and Science |
| Publication Year: | 2024 |
| Collection: | University of Lethbridge Institutional Repository |
| Subject Terms: | batch-based migration algorithms, cloud computing, NewSQL systems, data migration, data pipelines, documentation, Dissertations, Academic, SQL (Computer program language), Big data, Algorithms, Electronic data processing--Batch processing--Documentation, Data mining |
| Description: | Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics. |
| Document Type: | thesis |
| File Description: | application/pdf |
| Language: | English |
| Relation: | Thesis (University of Lethbridge. Faculty of Arts and Science); https://hdl.handle.net/10133/6939 |
| Availability: | https://hdl.handle.net/10133/6939 |
| Accession Number: | edsbas.EE476C9B |
| Database: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://hdl.handle.net/10133/6939# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.EE476C9B RelevancyScore: 802 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 801.7080078125 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Cost-effective batch-based migration strategies for NewSQL-based big data systems – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Vadlamudi%2C+Naveen+Kumar%22">Vadlamudi, Naveen Kumar</searchLink><br /><searchLink fieldCode="AR" term="%22University+of+Lethbridge%2E+Faculty+of+Arts+and+Science%22">University of Lethbridge. Faculty of Arts and Science</searchLink> – Name: Author Label: Contributors Group: Au Data: Osborn, Wendy – Name: Publisher Label: Publisher Information Group: PubInfo Data: University of Lethbridge, Dept. of Mathematics and Computer Science<br />Department of Mathematics and Computer Science<br />Arts and Science – Name: DatePubCY Label: Publication Year Group: Date Data: 2024 – Name: Subset Label: Collection Group: HoldingsInfo Data: University of Lethbridge Institutional Repository – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22batch-based+migration+algorithms%22">batch-based migration algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22cloud+computing%22">cloud computing</searchLink><br /><searchLink fieldCode="DE" term="%22NewSQL+systems%22">NewSQL systems</searchLink><br /><searchLink fieldCode="DE" term="%22data+migration%22">data migration</searchLink><br /><searchLink fieldCode="DE" term="%22data+pipelines%22">data pipelines</searchLink><br /><searchLink fieldCode="DE" term="%22documentation%22">documentation</searchLink><br /><searchLink fieldCode="DE" term="%22Dissertations%22">Dissertations</searchLink><br /><searchLink fieldCode="DE" term="%22Academic%22">Academic</searchLink><br /><searchLink fieldCode="DE" term="%22SQL+%28Computer+program+language%29%22">SQL (Computer program language)</searchLink><br /><searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+data+processing--Batch+processing--Documentation%22">Electronic data processing--Batch processing--Documentation</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink> – Name: Abstract Label: Description Group: Ab Data: Modern, high-performance applications demand scalable and efficient databases, leading to the evolution of NewSQL systems. The challenge lies in migrating data from Shardingsphere with PostgreSQL to AWS (AmazonWeb Services) cloud object storage. Implementing batch migration algorithms in Apache Spark, specifically targeting Delta Lake format, introduces complexities to ensure seamless data integration and storage within AWS environments. This thesis explores tailored batch-based migration algorithms for transferring data from Shardingsphere with PostgreSQL to AWS cloud object storage, emphasizing performance optimization by transferring the data faster. The study evaluates various batch loading techniques in Apache Spark, including sequential and concurrent strategies for shard-by-shard and aggregated-shards based algorithms. These techniques aim to maximize efficiency in storing data in Delta Lake format within AWS cloud storage, facilitating effective data management, visualization, and utilization for modern applications, business intelligence, AI and ML. Leveraging the Lakehouse architecture for integrated data processing and analytics. – Name: TypeDocument Label: Document Type Group: TypDoc Data: thesis – Name: Format Label: File Description Group: SrcInfo Data: application/pdf – Name: Language Label: Language Group: Lang Data: English – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: Thesis (University of Lethbridge. Faculty of Arts and Science); https://hdl.handle.net/10133/6939 – Name: URL Label: Availability Group: URL Data: https://hdl.handle.net/10133/6939 – Name: AN Label: Accession Number Group: ID Data: edsbas.EE476C9B |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.EE476C9B |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English Subjects: – SubjectFull: batch-based migration algorithms Type: general – SubjectFull: cloud computing Type: general – SubjectFull: NewSQL systems Type: general – SubjectFull: data migration Type: general – SubjectFull: data pipelines Type: general – SubjectFull: documentation Type: general – SubjectFull: Dissertations Type: general – SubjectFull: Academic Type: general – SubjectFull: SQL (Computer program language) Type: general – SubjectFull: Big data Type: general – SubjectFull: Algorithms Type: general – SubjectFull: Electronic data processing--Batch processing--Documentation Type: general – SubjectFull: Data mining Type: general Titles: – TitleFull: Cost-effective batch-based migration strategies for NewSQL-based big data systems Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Vadlamudi, Naveen Kumar – PersonEntity: Name: NameFull: University of Lethbridge. Faculty of Arts and Science – PersonEntity: Name: NameFull: Osborn, Wendy IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2024 Identifiers: – Type: issn-locals Value: edsbas |
| ResultId | 1 |