Firehose

Filtered to tagged “latency optimization” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

21 JUL 2026 · Vercel

AI Gateway now supports service tiering. Service tiers let you optimize for latency, throughput, and cost per request to match your use case. Pick a faster tier for interactive workloads (less queueing, higher token throughput), or a lower …

29 APR 2026 · Podcast · Dwarkesh Podcast

Reiner Pope delivers a blackboard lecture on the mathematical and hardware principles behind training and serving large language models. He explains how batch size, sparsity, and various parallelism strategies (expert, pipeline) impact late…