This paper investigates where the optimizer state should be allocated in the context of mixture-of-experts training to reduce memory usage without sacrificing accuracy. Practitioners might care because optimizing memory usage can be crucial for large-scale language models.
Firehose
Filtered to Papers, tagged “AdamW” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News