Diffusers now supports loading Nunchaku 4-bit diffusion inference checkpoints without requiring a separate inference engine, leveraging the Hugging Face kernels package to download necessary CUDA kernels from the Hub. This enables faster and lower memory usage inference, reducing the VRAM requirements to around 12 GB for the 1024x1024 image size, compared to 24 GB for BF16 precision. AI summary
Firehose
Filtered to Companies, tagged “4-bit diffusion” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News