AI Gateway now supports service tiering. Service tiers let you optimize for latency, throughput, and cost per request to match your use case. Pick a faster tier for interactive workloads (less queueing, higher token throughput), or a lower …
Firehose
Filtered to tagged “Latency optimization” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
artificial intelligence 43open-weight models 26agentic coding 20AI 14reinforcement learning 14AI safety 13continual learning 13AI agents 10cybersecurity 7large language models 7open-source 6Reinforcement learning 6benchmarking 5finance 5tech 5web development 5Agentic AI 4AI ethics 4autoregressive models 4Diffusion models 4language models 4Recursive self-improvement 4security 4vision-language models 4AI infrastructure 3AI security 3Code generation 3diffusion models 3diffusion transformers 3general 3
Reiner Pope delivers a blackboard lecture on the mathematical and hardware principles behind training and serving large language models. He explains how batch size, sparsity, and various parallelism strategies (expert, pipeline) impact late…
LLM trainingLLM inferenceBatch sizeLatency optimizationCost analysisRoofline analysisMemory bandwidthCompute performanceKV CacheSparsityMixture of ExpertsExpert parallelismData center architectureScale up networkScale out networkPipeline parallelismMicrobatchingMemory capacityChinchilla scalingRL generationAPI pricingContext lengthCryptographic ciphersNeural network architectureReversible networks