Dissertation/ Thesis

Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches
Συγγραφείς: Lion, Romain
Συνεισφορές: Laboratoire Bordelais de Recherche en Informatique (LaBRI), Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS), Université de Bordeaux, Samuel Thibault
Πηγή: https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩.
Στοιχεία εκδότη: HAL CCSD
Έτος έκδοσης: 2022
Συλλογή: Archive ouverte HAL (Hyper Article en Ligne, CCSD - Centre pour la Communication Scientifique Directe)
Θεματικοί όροι: Failure Tolerance, Checkpoints, Task-based runtime system, Tolérance aux pannes, Checkpoint, Support d’exécution à base de tâches, [INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF]
Περιγραφή: While computing power of systems grows, their reliabity decreases inevitably. Indeed, performance is achieved by leveraging components quantity and complexity, therefore computing systems are subject to failures on a daily basis. The problem is to use a mechanism to tolerate failures while having the least impact on performance. Moreover, supercomputers archichecture become more and more complex, and so becomes their coding. Data-based runtime systems such as StarPU are responding to this problematic. This thesis proposes a dedicated failure tolerance protocol to StarPU. The STF programming model used in StarPU allows to create consistent coordinated non-blocking asynchronous checkpoints very simply, by inserting checkpoint requests statically in the source code, like an application-based checkpoint solution. Furthermore, managing the checkpoints inside StarPU allows to use the synergy between computing data and checkpoint data, allowing to significantly reduce the amount of data that needs to be saved. We exploit this effect by choosing to save checkpoints on the other computing nodes, and by performing local rollback using message logging. The efficiency of our proposal is evaluated with a Cholesky decomposition application. We also show that with a particular setting for this application, our approach allows to tolerate the failure corresponding to our hypothesis without having any data actually replicated on other nodes. This is done by using the fact that the application already replicates enough data due to the computation needs, while with our approach we are able to exploit these data as checkpoint data. ; À mesure que la puissance de calcul des nouveaux supercalculateurs augmente, leur fiabilitédécroît inexorablement. En effet les limites sont repoussées en augmentant le nombre de composantsainsi que leur complexité, et les systèmes de calcul expérimentent des défaillances au quotidien. Laproblématique est donc de pouvoir se prémunir des pannes, tout en limitant l’impact sur les performancesqu’impose un ...
Τύπος εγγράφου: doctoral or postdoctoral thesis
Γλώσσα: French
Relation: NNT: 2022BORD0393; tel-04213186; https://theses.hal.science/tel-04213186; https://theses.hal.science/tel-04213186/document; https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf
Διαθεσιμότητα: https://theses.hal.science/tel-04213186
https://theses.hal.science/tel-04213186/document
https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf
Rights: info:eu-repo/semantics/OpenAccess
Αριθμός Καταχώρησης: edsbas.62A19680
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://theses.hal.science/tel-04213186#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.62A19680
RelevancyScore: 842
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 842.276306152344
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Lion%2C+Romain%22">Lion, Romain</searchLink>
– Name: Author
  Label: Contributors
  Group: Au
  Data: Laboratoire Bordelais de Recherche en Informatique (LaBRI)<br />Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS)<br />Université de Bordeaux<br />Samuel Thibault
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <i>https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩</i>.
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: HAL CCSD
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2022
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: Archive ouverte HAL (Hyper Article en Ligne, CCSD - Centre pour la Communication Scientifique Directe)
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Failure+Tolerance%22">Failure Tolerance</searchLink><br /><searchLink fieldCode="DE" term="%22Checkpoints%22">Checkpoints</searchLink><br /><searchLink fieldCode="DE" term="%22Task-based+runtime+system%22">Task-based runtime system</searchLink><br /><searchLink fieldCode="DE" term="%22Tolérance+aux+pannes%22">Tolérance aux pannes</searchLink><br /><searchLink fieldCode="DE" term="%22Checkpoint%22">Checkpoint</searchLink><br /><searchLink fieldCode="DE" term="%22Support+d%27exécution+à+base+de+tâches%22">Support d’exécution à base de tâches</searchLink><br /><searchLink fieldCode="DE" term="%22[INFO%2EINFO-PF]Computer+Science+[cs]%2FPerformance+[cs%2EPF]%22">[INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF]</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: While computing power of systems grows, their reliabity decreases inevitably. Indeed, performance is achieved by leveraging components quantity and complexity, therefore computing systems are subject to failures on a daily basis. The problem is to use a mechanism to tolerate failures while having the least impact on performance. Moreover, supercomputers archichecture become more and more complex, and so becomes their coding. Data-based runtime systems such as StarPU are responding to this problematic. This thesis proposes a dedicated failure tolerance protocol to StarPU. The STF programming model used in StarPU allows to create consistent coordinated non-blocking asynchronous checkpoints very simply, by inserting checkpoint requests statically in the source code, like an application-based checkpoint solution. Furthermore, managing the checkpoints inside StarPU allows to use the synergy between computing data and checkpoint data, allowing to significantly reduce the amount of data that needs to be saved. We exploit this effect by choosing to save checkpoints on the other computing nodes, and by performing local rollback using message logging. The efficiency of our proposal is evaluated with a Cholesky decomposition application. We also show that with a particular setting for this application, our approach allows to tolerate the failure corresponding to our hypothesis without having any data actually replicated on other nodes. This is done by using the fact that the application already replicates enough data due to the computation needs, while with our approach we are able to exploit these data as checkpoint data. ; À mesure que la puissance de calcul des nouveaux supercalculateurs augmente, leur fiabilitédécroît inexorablement. En effet les limites sont repoussées en augmentant le nombre de composantsainsi que leur complexité, et les systèmes de calcul expérimentent des défaillances au quotidien. Laproblématique est donc de pouvoir se prémunir des pannes, tout en limitant l’impact sur les performancesqu’impose un ...
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: doctoral or postdoctoral thesis
– Name: Language
  Label: Language
  Group: Lang
  Data: French
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: NNT: 2022BORD0393; tel-04213186; https://theses.hal.science/tel-04213186; https://theses.hal.science/tel-04213186/document; https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf
– Name: URL
  Label: Availability
  Group: URL
  Data: https://theses.hal.science/tel-04213186<br />https://theses.hal.science/tel-04213186/document<br />https://theses.hal.science/tel-04213186/file/LION_ROMAIN_2022.pdf
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: info:eu-repo/semantics/OpenAccess
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.62A19680
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.62A19680
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: French
    Subjects:
      – SubjectFull: Failure Tolerance
        Type: general
      – SubjectFull: Checkpoints
        Type: general
      – SubjectFull: Task-based runtime system
        Type: general
      – SubjectFull: Tolérance aux pannes
        Type: general
      – SubjectFull: Checkpoint
        Type: general
      – SubjectFull: Support d’exécution à base de tâches
        Type: general
      – SubjectFull: [INFO.INFO-PF]Computer Science [cs]/Performance [cs.PF]
        Type: general
    Titles:
      – TitleFull: Data replication for failure-tolerance in a distributed task-based runtime system ; Réplication de données pour la tolérance aux pannes dans un support d'exécution distribué à base de tâches
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Lion, Romain
      – PersonEntity:
          Name:
            NameFull: Laboratoire Bordelais de Recherche en Informatique (LaBRI)
      – PersonEntity:
          Name:
            NameFull: Université de Bordeaux (UB)-École Nationale Supérieure d'Électronique, Informatique et Radiocommunications de Bordeaux (ENSEIRB)-Centre National de la Recherche Scientifique (CNRS)
      – PersonEntity:
          Name:
            NameFull: Université de Bordeaux
      – PersonEntity:
          Name:
            NameFull: Samuel Thibault
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
          Titles:
            – TitleFull: https://theses.hal.science/tel-04213186 ; Performance et fiabilité [cs.PF]. Université de Bordeaux, 2022. Français. ⟨NNT : 2022BORD0393⟩
              Type: main
ResultId 1