Fault tolerant distributed training on Amazon EKS using NVRx
Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.
What happened
Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.
Why it matters
The proceeding may create legal precedent, financial exposure or operating constraints for NVIDIA, Amazon.
Affected entities
View evidence
1 reports · 1 original report · 1 independent
- AWS Machine Learning BlogPrimary source · Supports · EN · 100%Fault tolerant distributed training on Amazon EKS using NVRx ↗
Claims
- Fault tolerant distributed training on Amazon EKS using NVRx Observed
Conflicts
No material conflict detected in the available evidence.
Timeline
- First reported
Market move following event
Market reaction is not yet available for this asset and time window.