Dissertation/ Thesis
Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches
| Τίτλος: | Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches |
|---|---|
| Συγγραφείς: | Lion, Romain |
| Συνεισφορές: | Laboratoire Bordelais de Recherche en Informatique (LaBRI), Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS), Université de Bordeaux, Samuel Thibault |
| Πηγή: | https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩. |
| Στοιχεία εκδότη: | HAL CCSD |
| Έτος έκδοσης: | 2022 |
| Συλλογή: | Archive ouverte HAL (Hyper Article en Ligne, CCSD - Centre pour la Communication Scientifique Directe) |
| Θεματικοί όροι: | Failure Tolerance, Checkpoints, Task-based runtime system, Tolérance aux pannes, Checkpoint, Support d’exécution à base de tâches, [INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF] |
| Περιγραφή: | While computing power of systems grows, their reliabity decreases inevitably. Indeed, performance is achieved by leveraging components quantity and complexity, therefore computing systems are subject to failures on a daily basis. The problem is to use a mechanism to tolerate failures while having the least impact on performance. Moreover, supercomputers archichecture become more and more complex, and so becomes their coding. Data-based runtime systems such as StarPU are responding to this problematic. This thesis proposes a dedicated failure tolerance protocol to StarPU. The STF programming model used in StarPU allows to create consistent coordinated non-blocking asynchronous checkpoints very simply, by inserting checkpoint requests statically in the source code, like an application-based checkpoint solution. Furthermore, managing the checkpoints inside StarPU allows to use the synergy between computing data and checkpoint data, allowing to significantly reduce the amount of data that needs to be saved. We exploit this effect by choosing to save checkpoints on the other computing nodes, and by performing local rollback using message logging. The efficiency of our proposal is evaluated with a Cholesky decomposition application. We also show that with a particular setting for this application, our approach allows to tolerate the failure corresponding to our hypothesis without having any data actually replicated on other nodes. This is done by using the fact that the application already replicates enough data due to the computation needs, while with our approach we are able to exploit these data as checkpoint data. ; À mesure que la puissance de calcul des nouveaux supercalculateurs augmente, leur fiabilitédécroît inexorablement. En effet les limites sont repoussées en augmentant le nombre de composantsainsi que leur complexité, et les systèmes de calcul expérimentent des défaillances au quotidien. Laproblématique est donc de pouvoir se prémunir des pannes, tout en limitant l’impact sur les performancesqu’impose un ... |
| Τύπος εγγράφου: | doctoral or postdoctoral thesis |
| Γλώσσα: | French |
| Relation: | NNT: 2022BORD0393; tel-04213186; https://theses.hal.science/tel-04213186; https://theses.hal.science/tel-04213186/document; https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf |
| Διαθεσιμότητα: | https://theses.hal.science/tel-04213186 https://theses.hal.science/tel-04213186/document https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf |
| Rights: | info:eu-repo/semantics/OpenAccess |
| Αριθμός Καταχώρησης: | edsbas.62A19680 |
| Βάση Δεδομένων: | BASE |
| Η περιγραφή δεν είναι διαθέσιμη |