Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Mar 07, 2020

Qinqing Zheng, Bor-Yiing Su, Jiyan Yang, Alisson Azzolini, Qiang Wu, Ou Jin, Shri Karandikar, Hagay Lupesko, Liang Xiong, Eric Zhou

Figure 1 for ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Figure 2 for ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Figure 3 for ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Figure 4 for ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Share this with someone who'll enjoy it:

Abstract:Distributed training is useful to train complicated models to shorten the training time. As each of the workers only sees a small fraction of data, workers need to synchronize on the parameter updates. One of the central questions in distributed training is how to parsimoniously synchronize parameters while preserving model quality. To address this problem, we propose the \textbf{ShadowSync} framework, in which we isolate synchronization from training and run it in the background. In contrast to common strategies including synchronous stochastic gradient descent (SGD), asynchronous SGD, and model averaging on independently trained sub-models, where synchronization happens in the foreground, ShadowSync synchronization is neither part of the backward pass, nor happens every $k$ iterations. Our framework is generic to host various types of synchronization algorithms, and we propose 3 approaches under this theme. The superiority of ShadowSync is confirmed by experiments on training deep neural networks for click-through-rate prediction. Our methods all succeed in making the training throughput linearly scale with the number of trainers. Comparing to their foreground counterparts, our methods exhibit neutral to better model quality and better scalability when we keep the number of parameter servers the same. In our training system which expresses both replication and Hogwild parallelism, ShadowSync also accomplishes the highest example level parallelism number comparing to the prior arts.

View paper on

Share this with someone who'll enjoy it:

Title:ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training

Paper and Code