IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
This paper proposes a new architecture for Mixture-of-Experts (MoE) models that balances participation, execution, and materialization costs. Practitioners might care because it can lead to significant performance gains in applications where memory and computational resources are limited.