Firehose

Filtered to tagged “large language models” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

24 JUL 2026 · Hacker News · 138 pts

Hetzner is launching an experimental Large Language Model (LLM) inference API, compatible with OpenAI models, allowing users to run inference on Hetzner's infrastructure without the need for custom hardware. The initial model is Qwen/Qwen3.6-35B-A3B-FP8, a 35-billion-parameter Mixture-of-Experts model, and the API is currently free and fast, but its scalability and ability to handle larger models remain to be seen. The experiment aims to test Hetzner's infrastructure and learn whether users want such a service, but its success will depend on Hetzner's ability to invest in the necessary hardware to support larger models. AI summary

23 JUL 2026 · Swyx

Poolside's co-CEO on how his small team of top researchers built a model factory capable of training Laguna S - a 118B MOE beating Thinky's ~1T open weights model... and this is just the beginning.

23 JUL 2026 · Vercel

Ling 3.0 Flash from Ant Group is now available on AI Gateway. The model is free to use for the next three weeks, through August 3rd. Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It…

21 JUL 2026 · Paper

This paper investigates using hypernetworks for large-scale knowledge injection into language models, a technique that can improve their ability to answer factual questions. Practitioners may care because it could lead to more accurate and scalable language models for applications like customer service or question-answering systems.

20 JUL 2026 · Paper

This paper proposes a new way for large language models to learn from feedback, allowing them to retain more detailed information about the quality of their responses and learn from it in a more nuanced way. Practitioners might care because this approach could lead to better performance on tasks where the model doesn't have a clear way to evaluate its own output.

18 JUL 2026 · Sebastian Raschka

Researchers have developed a way to control the effort mode of large language models (LLMs), allowing them to switch between low-effort, medium-effort, and high-effort reasoning modes. This is achieved through training using reinforcement learning with verifiable rewards (RLVR), which provides a reward signal for correct and incorrect responses. AI summary

17 JUL 2026 · Swyx

Kimi K3, a 2.8T-parameter open model, has been released by Moonshot AI, rivaling the largest closed models and surpassing prior open competitors, with a comparable intelligence to Opus 4.8 and GPT-5.5, but behind Fable 5 and GPT-5.6 Sol. The model's pricing is at Sonnet 5 levels, with a cost per task of $0.94, and it uses Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) for improved performance. AI summary