This paper optimizes the training of massive neural networks on a special-purpose hardware, the Ascend SuperPOD, to improve performance and stability. Practitioners in AI/ML model training might care about the techniques and results presented here for large-scale model training on non-GPU hardware.
Firehose
Filtered to Papers, tagged “large-scale distributed training” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News