Tagged "inference-efficiency"
- IBM Releases Granite 4.2 Models Optimized for Local LLM Deployment
- MSI Crosshair A16 HX: Professional Gaming Laptop Built for AI and Gaming
- GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
- Nvidia Accelerates Chip Engineering with AI Agents
- Sol-5.6 and Opus 5 Models Demonstrate Strong One-Shot Game Performance
- Show HN: TS Compiler Knowledge Graph Reducing AI Tokens About 90%
- Arm China Unveils "Tianxuan" CPU and Xingchen 300 Platform, Targeting Ubiquitous AIoT with On-Device AI Portfolio
- Don't Sleep on BitNet (2025)
- Tencent Open-Sources Hy3 295B MoE Model Built for STEM Reasoning
- Compressor V2: Three Compression Layers for 50% LLM Agent Cost Cut
- Mistral AI Launches Mistral Vibe
- The Brain vs. Deep Learning Part I: Computational Complexity Analysis
- Gemma 4 Replaces Entire Local LLM Stack for Many Practitioners
- Running a Local LLM on a 12-Year-Old Raspberry Pi: Practical Edge Inference
- DistillFast: AI Cost Optimization Tool for Model Efficiency
- Bun's Experimental Rust Rewrite Achieves 99.8% Test Compatibility on Linux
- Google Releases Gemma 4 Multi-Token Prediction Drafters To Accelerate AI Inference
- A 49-Line Physics Classifier That Beats kNN on 76% of Benchmarks
- Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
- Economic Implications of AI Adoption: Why Local Deployment Matters for Cost Control
- Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
- CricketBrain: Neuromorphic Signal Processor in Rust (0.175us/step, 944 bytes)
- TurboQuant: Understanding the Quantization Breakthrough
- RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
- Energy-Based Models Compared Against Frontier AI for Sudoku Solving