Picture for Gaohong Liu

Gaohong Liu

Minder: Faulty Machine Detection for Large-scale Distributed Model Training

Add code
Nov 04, 2024
Viaarxiv icon