نتایج جستجو برای: checkpointing

تعداد نتایج: 2665  

Journal: :Concurrency and Computation: Practice and Experience 2019

Journal: :Science China Information Sciences 2012

Journal: :Scalable Computing: Practice and Experience 2021

Application checkpointing is a widely used recovery mechanism that consists of saving an application's state periodically to be in case failure. In this study we investigate the utilisation distributed for replicated machines. Conventionally, machines, information stored way each replicas or separately single instance. Applying provides means adjust level fault tolerance approach by giving away...

Journal: :International Journal of High Performance Computing Applications 2021

Progress in numerical weather and climate prediction accuracy greatly depends on the growth of available computing power. As number cores top facilities pushes into millions, increased average frequency hardware software failures forces users to review their algorithms systems order protect simulations from breakdown. This report surveys hardware, application-level algorithm-level resilience ap...

Journal: :IEEE Trans. Computers 1992
E. N. Elnozahy Willy Zwaenepoel

Manetho is a new transparent rollback recovery protocol for long running distributed computations It uses a novel combination of antecedence graph maintenance unco ordinated checkpointing and sender based message logging Manetho simultaneously achieves the advantages of pessimistic message logging namely limited rollback and fast output commit and the advantage of optimistic message logging nam...

1992
Bruno R. Preiss Ian D. MacIntyre Wayne M. Loucks

Optimistically synchronized parallel discrete-event simulation is based on the use of communicating sequential processes. Optimistic synchronization means that the processes execute under the assumption that synchronization is fortuitous. Periodic checkpointing of the state of a process allows the process to roll back to an earlier state when synchronization errors occur. This paper examines th...

2005
Cosimo Anglano Massimo Canonico

In this paper we propose a fault-tolerant scheduler for Bagof-Tasks Grid applications, calledWorkQueue with Replication Fault Tolerant (WQR-FT), obtained by adding checkpointing and replication to the WorkQueue with Replication (WQR) scheduling algorithm. By using discrete-event simulation, we show that WQR-FT not only ensures the successful completion of all the tasks in a bag, but also achiev...

نمودار تعداد نتایج جستجو در هر سال

با کلیک روی نمودار نتایج را به سال انتشار فیلتر کنید