This paper proposes a new architecture for Mixture-of-Experts (MoE) models that balances participation, execution, and materialization costs. Practitioners might care because it can lead to significant performance gains in applications where memory and computational resources are limited.
Firehose
Filtered to Papers, tagged “mixture-of-experts” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper presents a new technique to reduce memory usage and speed up inference for large neural networks, allowing them to run on consumer hardware with limited memory. A practitioner might care about this because it enables the deployment of large models in edge devices and reduces the need for expensive storage.