Firehose

Filtered to Papers, tagged “optimizers” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

21 JUL 2026 · Paper

This paper investigates where the optimizer state should be allocated in the context of mixture-of-experts training to reduce memory usage without sacrificing accuracy. Practitioners might care because optimizing memory usage can be crucial for large-scale language models.