This paper proposes a new architecture for Mixture-of-Experts (MoE) models that balances participation, execution, and materialization costs. Practitioners might care because it can lead to significant performance gains in applications where memory and computational resources are limited.
Firehose
Filtered to Papers, tagged “dense output-mixing” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives