Jd
J.C. de Haan
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
Uncoordinated Checkpointing in Stateful Transactional Systems
Decoupling Fault Tolerance from Coordination in Styx
Styx is a distributed runtime for stateful transactional functions. It executes transactions in deterministic, lockstep epochs and periodically writes state to
stable storage for recovery. In the original design, each checkpoint depends on both the workers’ state files and a coordinator-written global sequencer file.
Consequently, a delayed or failed coordinator write can make otherwise complete worker checkpoints unusable. This thesis investigates whether checkpoint
persistence can be moved from the coordinator to the workers without harming system performance or recovery.
The proposed design uses Styx’s shared epoch boundaries as consistent recovery points. Workers persist their partition checkpoints independently, while
recovery selects the newest complete epoch across all partitions. Transaction execution remains coordinated; only checkpoint persistence is decentralized. The
evaluation considers steady-state performance, checkpoint-file compaction, and partition rebalancing during degraded operation after a worker failure.
Across approximately 1,670 single-node runs and 334 configurations, coordinated and uncoordinated checkpoint persistence show no systematic performance
difference. Their saturation points differ by -5.3% to +6.3%, with no consistent winner. Recovery performance depends more strongly on how stored checkpoints
are managed. Compacting every 100 checkpoints reduces state restoration time from 4,971 ms to 474 ms and provides the best tested balance between recovery
speed and foreground latency. More aggressive compaction restores state faster but becomes disruptive as load increases.
Rebalancing is similarly workload-dependent. After a worker failure, the surviving workers operate with an uneven partition distribution. Redistributing those
partitions reduces P95 latency from 288 ms to 146 ms at 40% utilization without affecting throughput. At loads of 50% and above, however, the newly restored
worker limits epoch progress and lowers throughput during the 90-second observation window.
These results show that checkpoint persistence can be decentralized without measurable steady-state cost under the tested conditions. Compaction should balance
recovery objectives against foreground load, while post-failure rebalancing should be applied selectively rather than as a universal recovery rule. The
findings are limited to one workload and a single-node deployment.
Related dataset 4TU.ResearchData: https://doi.org/10.4121/a18e4390-8d18-4a2a-883c-0fe6a103bda1 ...
stable storage for recovery. In the original design, each checkpoint depends on both the workers’ state files and a coordinator-written global sequencer file.
Consequently, a delayed or failed coordinator write can make otherwise complete worker checkpoints unusable. This thesis investigates whether checkpoint
persistence can be moved from the coordinator to the workers without harming system performance or recovery.
The proposed design uses Styx’s shared epoch boundaries as consistent recovery points. Workers persist their partition checkpoints independently, while
recovery selects the newest complete epoch across all partitions. Transaction execution remains coordinated; only checkpoint persistence is decentralized. The
evaluation considers steady-state performance, checkpoint-file compaction, and partition rebalancing during degraded operation after a worker failure.
Across approximately 1,670 single-node runs and 334 configurations, coordinated and uncoordinated checkpoint persistence show no systematic performance
difference. Their saturation points differ by -5.3% to +6.3%, with no consistent winner. Recovery performance depends more strongly on how stored checkpoints
are managed. Compacting every 100 checkpoints reduces state restoration time from 4,971 ms to 474 ms and provides the best tested balance between recovery
speed and foreground latency. More aggressive compaction restores state faster but becomes disruptive as load increases.
Rebalancing is similarly workload-dependent. After a worker failure, the surviving workers operate with an uneven partition distribution. Redistributing those
partitions reduces P95 latency from 288 ms to 146 ms at 40% utilization without affecting throughput. At loads of 50% and above, however, the newly restored
worker limits epoch progress and lowers throughput during the 90-second observation window.
These results show that checkpoint persistence can be decentralized without measurable steady-state cost under the tested conditions. Compaction should balance
recovery objectives against foreground load, while post-failure rebalancing should be applied selectively rather than as a universal recovery rule. The
findings are limited to one workload and a single-node deployment.
Related dataset 4TU.ResearchData: https://doi.org/10.4121/a18e4390-8d18-4a2a-883c-0fe6a103bda1 ...
Styx is a distributed runtime for stateful transactional functions. It executes transactions in deterministic, lockstep epochs and periodically writes state to
stable storage for recovery. In the original design, each checkpoint depends on both the workers’ state files and a coordinator-written global sequencer file.
Consequently, a delayed or failed coordinator write can make otherwise complete worker checkpoints unusable. This thesis investigates whether checkpoint
persistence can be moved from the coordinator to the workers without harming system performance or recovery.
The proposed design uses Styx’s shared epoch boundaries as consistent recovery points. Workers persist their partition checkpoints independently, while
recovery selects the newest complete epoch across all partitions. Transaction execution remains coordinated; only checkpoint persistence is decentralized. The
evaluation considers steady-state performance, checkpoint-file compaction, and partition rebalancing during degraded operation after a worker failure.
Across approximately 1,670 single-node runs and 334 configurations, coordinated and uncoordinated checkpoint persistence show no systematic performance
difference. Their saturation points differ by -5.3% to +6.3%, with no consistent winner. Recovery performance depends more strongly on how stored checkpoints
are managed. Compacting every 100 checkpoints reduces state restoration time from 4,971 ms to 474 ms and provides the best tested balance between recovery
speed and foreground latency. More aggressive compaction restores state faster but becomes disruptive as load increases.
Rebalancing is similarly workload-dependent. After a worker failure, the surviving workers operate with an uneven partition distribution. Redistributing those
partitions reduces P95 latency from 288 ms to 146 ms at 40% utilization without affecting throughput. At loads of 50% and above, however, the newly restored
worker limits epoch progress and lowers throughput during the 90-second observation window.
These results show that checkpoint persistence can be decentralized without measurable steady-state cost under the tested conditions. Compaction should balance
recovery objectives against foreground load, while post-failure rebalancing should be applied selectively rather than as a universal recovery rule. The
findings are limited to one workload and a single-node deployment.
Related dataset 4TU.ResearchData: https://doi.org/10.4121/a18e4390-8d18-4a2a-883c-0fe6a103bda1
stable storage for recovery. In the original design, each checkpoint depends on both the workers’ state files and a coordinator-written global sequencer file.
Consequently, a delayed or failed coordinator write can make otherwise complete worker checkpoints unusable. This thesis investigates whether checkpoint
persistence can be moved from the coordinator to the workers without harming system performance or recovery.
The proposed design uses Styx’s shared epoch boundaries as consistent recovery points. Workers persist their partition checkpoints independently, while
recovery selects the newest complete epoch across all partitions. Transaction execution remains coordinated; only checkpoint persistence is decentralized. The
evaluation considers steady-state performance, checkpoint-file compaction, and partition rebalancing during degraded operation after a worker failure.
Across approximately 1,670 single-node runs and 334 configurations, coordinated and uncoordinated checkpoint persistence show no systematic performance
difference. Their saturation points differ by -5.3% to +6.3%, with no consistent winner. Recovery performance depends more strongly on how stored checkpoints
are managed. Compacting every 100 checkpoints reduces state restoration time from 4,971 ms to 474 ms and provides the best tested balance between recovery
speed and foreground latency. More aggressive compaction restores state faster but becomes disruptive as load increases.
Rebalancing is similarly workload-dependent. After a worker failure, the surviving workers operate with an uneven partition distribution. Redistributing those
partitions reduces P95 latency from 288 ms to 146 ms at 40% utilization without affecting throughput. At loads of 50% and above, however, the newly restored
worker limits epoch progress and lowers throughput during the 90-second observation window.
These results show that checkpoint persistence can be decentralized without measurable steady-state cost under the tested conditions. Compaction should balance
recovery objectives against foreground load, while post-failure rebalancing should be applied selectively rather than as a universal recovery rule. The
findings are limited to one workload and a single-node deployment.
Related dataset 4TU.ResearchData: https://doi.org/10.4121/a18e4390-8d18-4a2a-883c-0fe6a103bda1