- Bookmark stories with reactions via GitHub
- Comment on any post — no account needed to read
- Write your own posts or guides
Recent Posts
-
Lemonade 11.9 Local AI Server Released With AMD ROCm HRX Backend
Lemonade AI server reaches version 11.9 with new AMD ROCm HRX backend support, expanding local inference capabilities to AMD GPU hardware and providing an alternative to NVIDIA-focused deployment stacks.
-
llama.cpp Release b10781: Vulkan Backend and Efficiency Improvements
Latest llama.cpp release includes Vulkan fixes and optimizations for cross-platform GPU inference, continuing the project's rapid iteration on inference performance and hardware support.
-
Optimizing On-Device Inference for Apple Silicon
Perplexity publishes a comprehensive guide on optimizing LLM inference specifically for Apple Silicon, covering techniques to maximize performance and efficiency on Apple's ARM-based processors for local deployment.
-
Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
A developer demonstrates running a 104GB model on a 48GB Mac using innovative slot streaming techniques, achieving practical inference speeds of ~12 tokens/second and expanding the possibilities for large model deployment on consumer hardware.
-
Show HN: Single-File GGUF Inference
A browser-based GGUF inference solution enables running quantized models directly in WebAssembly, allowing local LLM inference without any backend server or installation required.
-
Llama.cpp Fork Enables Qwen 3.8 27B with Large Contexts on 16GB VRAM GPUs
A specialized llama.cpp fork implements adaptive KV streaming to run Qwen 3.8 27B with large context windows on 16GB VRAM GPUs, significantly reducing hardware requirements for production-grade inference.
-
Hugging Face Releases 200+ WebGPU Kernels for Local AI Inference
Hugging Face launches a comprehensive collection of WebGPU kernels enabling efficient local AI inference directly in browsers and on-device. This represents a major step toward browser-native LLM deployment without server backends.