In this episode, Akshat Bubna, CTO of Modal, discusses the evolution of AI infrastructure tailored for agent experience, highlighting Modal's journey from a runtime platform to a specialized cloud for AI workloads. They explore challenges i…
Firehose
Filtered to Podcasts, tagged “LLM inference” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
Reiner Pope delivers a blackboard lecture on the mathematical and hardware principles behind training and serving large language models. He explains how batch size, sparsity, and various parallelism strategies (expert, pipeline) impact late…
LLM trainingLLM inferenceBatch sizeLatency optimizationCost analysisRoofline analysisMemory bandwidthCompute performanceKV CacheSparsityMixture of ExpertsExpert parallelismData center architectureScale up networkScale out networkPipeline parallelismMicrobatchingMemory capacityChinchilla scalingRL generationAPI pricingContext lengthCryptographic ciphersNeural network architectureReversible networks