Tagged "inference-speed"
-
Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
-
Running 104GB Qwen3.8-Flash-Next on 48GB Mac with Slotstream at ~12 tok/s
-
DSpark Speculative Decoding: Speeding Up LLM Inference
-
Gemma 4 vs Phi-4 Mini vs Qwen3.5: On-Device AI Comparison 2026
-
Gemma 4 MoE for Agentic Coding: Testing Open-Weight Models on AMD APU Hardware
-
llama.cpp Optimizes DFlash Encoder with KV Cache Injection
-
AMD ROCm 10 Arrives With ROCm.AI GA: Hyperloom Agents and 3.3x Inference Lift
-
VRAM Optimization Breakthrough: Single Setting Change Doubles Local Model Speed
-
vLLM-iOS Achieves 88% Faster Multi-Agent Inference Through Continuous Batching on iPhone
-
Benchmarking Qwen3.8 27B Quantizations: 4-bit Shows Strong Performance, 1-bit Collapses
-
vLLM v0.28.0 Features Major Kimi-K3 Optimization and Decode Context Parallel Support
-
Llama.cpp Build 10620: Continued Optimization for Local Inference
-
llama.cpp Build 10605: Mamba2 GEMM Optimization Improves State-Space Model Performance
-
Liquid AI Releases DSpark Version of Compact LFM2.5 Models with Up to 2.67x Speedup
-
vLLM's Disaggregated Serving Cuts GPU Interference, Delivering 2.5x Higher Goodput
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Ollama v0.33.0 Release Candidate Adds Claude Desktop Integration and Performance Improvements
-
Liquid AI Releases LFM2.5-DSpark Draft Models with 3.18x Faster Decoding
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
-
AMD EPYC ZenDNN Accelerates llama.cpp Prompt Processing 4.5x
-
AMD Adds Day 0 Qwen3.8 Support, Radeon AI PRO R9700 Hits 51.8 Tokens per Second
-
Meta's Muse Glimmer Achieves Fast On-Device Agentic AI with ExecuTorch
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
Local Model Performance Benchmarks on MacBook Pro M5 Max: Real-World Inference Metrics
-
Ollama v0.32.10: Faster Prefill Performance on NVFP4 Models with System Config Support
-
vLLM v0.27.0 Brings Major Performance Improvements and New Model Support
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Benchmarking Local LLMs on Consumer Hardware: Real-World Performance Data
-
vLLM v0.27.0 – Kimi K3 Support and 561 Commits from 242 Contributors
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
Ollama v0.32.6: Faster Apple GPU Inference with Speculative Decoding
-
Gainz.fast – Local Inference, Faster
-
28.9M-Parameter LLM Runs Locally on ESP32-S3 at 9 Tokens/s
-
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
-
Your Smartwatch Now Detects a Heart Irregularity in Milliseconds – Without Ever Touching the Cloud
-
Kioxia's UFS 5.0 Embedded Flash Enables Practical On-Device AI
-
K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
I Thought My Local AI Would Replace My Claude Subscription — Then I Tried Automating My PC
-
llama.cpp's 4.26× Intel Gain Has a Narrow Catch
-
Nvidia Boosts Token Throughput 5x With Software Optimizations, Reshaping AI Inference Economics
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Show HN: OpenVole 4.5 Is Out
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
Ollama's New MLX Engine Delivers Significant Performance Gains on Mac
-
Theoretical Bottlenecks for Scaling LLM Inference to Achieve Higher Token per Second
-
How to Choose Between Small and Frontier Models
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
ORA: Smaller Models. Same Intelligence
-
Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
-
Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
-
FlashRT: Execution State for Latency-First AI
-
Ray Serve LLM Achieves 24x Performance Improvement in Distributed Inference
-
Google's DiffusionGemma Brings Novel Text Generation to Local LLMs
-
AMD Brings Data Center-Level AI Performance to PCs
-
CacheWise Optimizes KVCache Reuse for LLM Coding Agents
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
Google's DiffusionGemma Achieves 4x Faster Text Generation for Local Deployment
-
vLLM vs Ollama 2026: 793 vs 41 TPS Performance Benchmark
-
Hermes with Ollama Emerges as Top Choice for Desktop AI Tools
-
Apple Enhances Siri With On-Device AI for Faster, Private Voice Responses
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Best Local LLM Setup for RTX 5090: llama.cpp Fork with TurboQuant
-
NVIDIA Dynamo Snapshot Accelerates AI Inference Startup on Kubernetes
-
Show HN: Lowfat – Pluggable CLI Filter Saving 91.8% of LLM Tokens
-
Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
-
Perplexity Unveils Hybrid Local-Cloud Inference System for Intelligent Task Distribution
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
Qualcomm Reveals Snapdragon C with Advanced On-Device AI Engine
-
NVIDIA Launches N1X/N1 CPU-GPU SoC for PC Market, Targeting Heavy On-Device AI Users
-
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
Dell Launches 14 Plus Laptop with Intel Core Ultra 9 and 32GB RAM at $1,499.99, Enabling Local Model Inference
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
Google Makes Gemini 3.5 Flash the Default AI Model for Billions of Users
-
llama.cpp Checkpoint Fix Accelerates Local Coding Agents
-
110 Tokens/Second on RTX 4070 Super with Qwen 3.6 35B
-
Intel llm-scaler-vllm 1.4 Released With Updated Components and Arc Pro B70 Support
-
Hardware LLM Taalas Reaches >14,000 TPS on Llama 3.1 8B
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
Google Tensor SDK Beta with LiteRT Enables Efficient On-Device AI
-
llama.cpp Adds Multi-Token Prediction, Doubles Qwen 3.6B Throughput for Local Inference
-
Orthrus Reshapes Economics of Local AI Inference with New Optimization Approach
-
llama.cpp Delivers Sharp Performance Gains for AMD RDNA3 Users
-
DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
-
One LM Studio Setting Makes Local LLMs Competitive With Cloud Models
-
Bun's Experimental Rust Rewrite Achieves 99.8% Test Compatibility on Linux
-
Lemonade Gives AMD Startups a Wider Path to Local Inference
-
Google Accelerates Gemma 4 Inference Speed 3x With Multi-Token Prediction Drafters
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Anker's Thus Chip Puts AI On-Device, Promising Faster Responses And Better Privacy
-
PFlash Claims 10x Prefill Speedup Over llama.cpp
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Linux Crushes Windows on llama.cpp Inference by Double Digits
-
Show HN: We built an OCR server that can process 270 dense images/s on a 5090
-
llama.cpp Merges Speculative Checkpointing for Major Inference Speed Boost
-
LlaMa.cpp Robot Wars
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
-
Fine-Tuned Qwen3.5-0.8B for OCR Outperforms Previous 2B Release
-
oMLX Framework Implements DFlash Attention for Optimized Inference
-
Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
-
The Best Local AI Model for Home Assistant Isn't Always the Biggest One
-
Google Gemma 4 Delivers Exceptional Speed and Accuracy for Local Inference
-
DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
-
On-Device AI: Achieving Powerful AI Capabilities Without Internet Connectivity
-
DMax: New Parallel Decoding Paradigm for Diffusion Language Models
-
Google's Gemini Nano 4 Offers Faster, Smarter Local Inference Capabilities
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Qwen 3.5 122B Achieves 198 Tokens/sec on Dual RTX PRO 6000 Blackwell GPUs
-
Speculative Decoding Made My Local LLM Actually Usable
-
Gemma 4 26B Achieves Impressive Local Performance With Proper Configuration
-
TurboQuant-Optimized llama.cpp Fork Delivers GFX906 GPU Acceleration
-
TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
-
HunyuanOCR 1B: High-Quality OCR Now Viable on Budget Consumer Hardware
-
Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
-
GMKtec NucBox K17 Launches with 97 TOPS AI Performance for Local Inference
-
Kokoro TTS Achieves 20× Realtime Speed on CPU-Only On-Device Inference
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
NVIDIA Accelerates Gemma 4 for Local Agentic AI on RTX GPUs
-
Google Gemma 4 Released with GGUF Quantizations
-
Gemma 4 26B A4B Outperforms Qwen 3.5 35B on Apple Silicon
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
TinyGPU Adds Mac Support for External Nvidia GPU Acceleration
-
Ollama Adopts Apple's MLX Framework for Faster Local AI on Mac
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
Linux Significantly Outperforms Windows for Local LLM Inference
-
TurboQuant: Understanding the Quantization Breakthrough
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
M5 Max Delivers 1.7x Faster Inference Than M3 Max on Qwen 3.5 Models
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Liquid AI's LFM2-24B Achieves 50 Tokens/Second in Web Browser via WebGPU
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
Critical: LiteLLM Supply Chain Attack Detected, Bifrost Alternative Released
-
Llama.cpp Benchmark: RTX 5090 vs Enterprise Systems Compared
-
Arm SME2 Technology Expands CPU Capabilities for On-Device AI
-
Carmack Proposes Using Long Fiber Lines as L2 Cache for Streaming AI Data