Firehose

Filtered to tagged “mixture-of-experts” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

21 JUL 2026 · Vercel

Laguna S 2.1 from Poolside is now available on AI Gateway. There are 2 versions of the model available: Free version (256K context window): poolside/laguna-s-2.1-free Paid version (1M context window): poolside/laguna-s-2.1 Laguna S 2.1 is a…

21 JUL 2026 · Paper

This paper investigates where the optimizer state should be allocated in the context of mixture-of-experts training to reduce memory usage without sacrificing accuracy. Practitioners might care because optimizing memory usage can be crucial for large-scale language models.