Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
This paper investigates where the optimizer state should be allocated in the context of mixture-of-experts training to reduce memory usage without sacrificing accuracy. Practitioners might care because optimizing memory usage can be crucial for large-scale language models.