Dissertation/ Thesis
Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches
| Τίτλος: | Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches |
|---|---|
| Συγγραφείς: | Lion, Romain |
| Συνεισφορές: | Laboratoire Bordelais de Recherche en Informatique (LaBRI), Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS), Université de Bordeaux, Samuel Thibault |
| Πηγή: | https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩. |
| Στοιχεία εκδότη: | HAL CCSD |
| Έτος έκδοσης: | 2022 |
| Συλλογή: | Archive ouverte HAL (Hyper Article en Ligne, CCSD - Centre pour la Communication Scientifique Directe) |
| Θεματικοί όροι: | Failure Tolerance, Checkpoints, Task-based runtime system, Tolérance aux pannes, Checkpoint, Support d’exécution à base de tâches, [INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF] |
| Περιγραφή: | While computing power of systems grows, their reliabity decreases inevitably. Indeed, performance is achieved by leveraging components quantity and complexity, therefore computing systems are subject to failures on a daily basis. The problem is to use a mechanism to tolerate failures while having the least impact on performance. Moreover, supercomputers archichecture become more and more complex, and so becomes their coding. Data-based runtime systems such as StarPU are responding to this problematic. This thesis proposes a dedicated failure tolerance protocol to StarPU. The STF programming model used in StarPU allows to create consistent coordinated non-blocking asynchronous checkpoints very simply, by inserting checkpoint requests statically in the source code, like an application-based checkpoint solution. Furthermore, managing the checkpoints inside StarPU allows to use the synergy between computing data and checkpoint data, allowing to significantly reduce the amount of data that needs to be saved. We exploit this effect by choosing to save checkpoints on the other computing nodes, and by performing local rollback using message logging. The efficiency of our proposal is evaluated with a Cholesky decomposition application. We also show that with a particular setting for this application, our approach allows to tolerate the failure corresponding to our hypothesis without having any data actually replicated on other nodes. This is done by using the fact that the application already replicates enough data due to the computation needs, while with our approach we are able to exploit these data as checkpoint data. ; À mesure que la puissance de calcul des nouveaux supercalculateurs augmente, leur fiabilitédécroît inexorablement. En effet les limites sont repoussées en augmentant le nombre de composantsainsi que leur complexité, et les systèmes de calcul expérimentent des défaillances au quotidien. Laproblématique est donc de pouvoir se prémunir des pannes, tout en limitant l’impact sur les performancesqu’impose un ... |
| Τύπος εγγράφου: | doctoral or postdoctoral thesis |
| Γλώσσα: | French |
| Relation: | NNT: 2022BORD0393; tel-04213186; https://theses.hal.science/tel-04213186; https://theses.hal.science/tel-04213186/document; https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf |
| Διαθεσιμότητα: | https://theses.hal.science/tel-04213186 https://theses.hal.science/tel-04213186/document https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf |
| Rights: | info:eu-repo/semantics/OpenAccess |
| Αριθμός Καταχώρησης: | edsbas.62A19680 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://theses.hal.science/tel-04213186# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.62A19680 RelevancyScore: 842 AccessLevel: 3 PubType: Dissertation/ Thesis PubTypeId: dissertation PreciseRelevancyScore: 842.276306152344 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Lion%2C+Romain%22">Lion, Romain</searchLink> – Name: Author Label: Contributors Group: Au Data: Laboratoire Bordelais de Recherche en Informatique (LaBRI)<br />Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS)<br />Université de Bordeaux<br />Samuel Thibault – Name: TitleSource Label: Source Group: Src Data: <i>https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩</i>. – Name: Publisher Label: Publisher Information Group: PubInfo Data: HAL CCSD – Name: DatePubCY Label: Publication Year Group: Date Data: 2022 – Name: Subset Label: Collection Group: HoldingsInfo Data: Archive ouverte HAL (Hyper Article en Ligne, CCSD - Centre pour la Communication Scientifique Directe) – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Failure+Tolerance%22">Failure Tolerance</searchLink><br /><searchLink fieldCode="DE" term="%22Checkpoints%22">Checkpoints</searchLink><br /><searchLink fieldCode="DE" term="%22Task-based+runtime+system%22">Task-based runtime system</searchLink><br /><searchLink fieldCode="DE" term="%22Tolérance+aux+pannes%22">Tolérance aux pannes</searchLink><br /><searchLink fieldCode="DE" term="%22Checkpoint%22">Checkpoint</searchLink><br /><searchLink fieldCode="DE" term="%22Support+d%27exécution+à+base+de+tâches%22">Support d’exécution à base de tâches</searchLink><br /><searchLink fieldCode="DE" term="%22[INFO%2EINFO-PF]Computer+Science+[cs]%2FPerformance+[cs%2EPF]%22">[INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF]</searchLink> – Name: Abstract Label: Description Group: Ab Data: While computing power of systems grows, their reliabity decreases inevitably. Indeed, performance is achieved by leveraging components quantity and complexity, therefore computing systems are subject to failures on a daily basis. The problem is to use a mechanism to tolerate failures while having the least impact on performance. Moreover, supercomputers archichecture become more and more complex, and so becomes their coding. Data-based runtime systems such as StarPU are responding to this problematic. This thesis proposes a dedicated failure tolerance protocol to StarPU. The STF programming model used in StarPU allows to create consistent coordinated non-blocking asynchronous checkpoints very simply, by inserting checkpoint requests statically in the source code, like an application-based checkpoint solution. Furthermore, managing the checkpoints inside StarPU allows to use the synergy between computing data and checkpoint data, allowing to significantly reduce the amount of data that needs to be saved. We exploit this effect by choosing to save checkpoints on the other computing nodes, and by performing local rollback using message logging. The efficiency of our proposal is evaluated with a Cholesky decomposition application. We also show that with a particular setting for this application, our approach allows to tolerate the failure corresponding to our hypothesis without having any data actually replicated on other nodes. This is done by using the fact that the application already replicates enough data due to the computation needs, while with our approach we are able to exploit these data as checkpoint data. ; À mesure que la puissance de calcul des nouveaux supercalculateurs augmente, leur fiabilitédécroît inexorablement. En effet les limites sont repoussées en augmentant le nombre de composantsainsi que leur complexité, et les systèmes de calcul expérimentent des défaillances au quotidien. Laproblématique est donc de pouvoir se prémunir des pannes, tout en limitant l’impact sur les performancesqu’impose un ... – Name: TypeDocument Label: Document Type Group: TypDoc Data: doctoral or postdoctoral thesis – Name: Language Label: Language Group: Lang Data: French – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: NNT: 2022BORD0393; tel-04213186; https://theses.hal.science/tel-04213186; https://theses.hal.science/tel-04213186/document; https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf – Name: URL Label: Availability Group: URL Data: https://theses.hal.science/tel-04213186<br />https://theses.hal.science/tel-04213186/document<br />https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf – Name: Copyright Label: Rights Group: Cpyrght Data: info:eu-repo/semantics/OpenAccess – Name: AN Label: Accession Number Group: ID Data: edsbas.62A19680 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.62A19680 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: French Subjects: – SubjectFull: Failure Tolerance Type: general – SubjectFull: Checkpoints Type: general – SubjectFull: Task-based runtime system Type: general – SubjectFull: Tolérance aux pannes Type: general – SubjectFull: Checkpoint Type: general – SubjectFull: Support d’exécution à base de tâches Type: general – SubjectFull: [INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF] Type: general Titles: – TitleFull: Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Lion, Romain – PersonEntity: Name: NameFull: Laboratoire Bordelais de Recherche en Informatique (LaBRI) – PersonEntity: Name: NameFull: Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS) – PersonEntity: Name: NameFull: Université de Bordeaux – PersonEntity: Name: NameFull: Samuel Thibault IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2022 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa Titles: – TitleFull: https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩ Type: main |
| ResultId | 1 |