Firehose

Filtered to tagged “inference optimization” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

24 JUL 2026 · Hacker News · 138 pts

Hetzner is launching an experimental Large Language Model (LLM) inference API, compatible with OpenAI models, allowing users to run inference on Hetzner's infrastructure without the need for custom hardware. The initial model is Qwen/Qwen3.6-35B-A3B-FP8, a 35-billion-parameter Mixture-of-Experts model, and the API is currently free and fast, but its scalability and ability to handle larger models remain to be seen. The experiment aims to test Hetzner's infrastructure and learn whether users want such a service, but its success will depend on Hetzner's ability to invest in the necessary hardware to support larger models. AI summary