Hetzner is launching an experimental Large Language Model (LLM) inference API, compatible with OpenAI models, allowing users to run inference on Hetzner's infrastructure without the need for custom hardware. The initial model is Qwen/Qwen3.6-35B-A3B-FP8, a 35-billion-parameter Mixture-of-Experts model, and the API is currently free and fast, but its scalability and ability to handle larger models remain to be seen. The experiment aims to test Hetzner's infrastructure and learn whether users want such a service, but its success will depend on Hetzner's ability to invest in the necessary hardware to support larger models. AI summary
Firehose
Filtered to tagged “inference optimization” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News
Browse by tag
artificial intelligence 43open-weight models 26agentic coding 20AI 14reinforcement learning 14AI safety 13continual learning 13AI agents 10cybersecurity 7large language models 7open-source 6Reinforcement learning 6benchmarking 5finance 5tech 5web development 5Agentic AI 4AI ethics 4autoregressive models 4Diffusion models 4language models 4Recursive self-improvement 4security 4vision-language models 4AI infrastructure 3AI security 3Code generation 3diffusion models 3diffusion transformers 3general 3