Tagged "moe-architecture"
- Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
- Ollama Releases NVIDIA Nemotron 3.5 Lightning for Agent Execution
- Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
- Tencent Open-Sources Hy3 295B MoE Model Built for STEM Reasoning
- Dynamic Expert Cache in llama.cpp Achieves 27% Faster Inference on Large MoE Models