Perplexity Open-Sources Lily: 1.35x Faster Inference Than MLX on Apple Silicon

1 min read

Perplexity has open-sourced Lily, a new inference framework purpose-built for Apple Silicon (M1, M2, M3 series) that achieves 1.35x faster inference compared to the widely-used MLX framework. The project represents a meaningful advancement in the maturity of the Apple Silicon LLM inference ecosystem, which has historically lagged behind CUDA and ROCm optimisations.

Lily's performance gains likely stem from deeper integration with Metal Performance Shaders and ANE (Apple Neural Engine) capabilities, allowing better hardware utilisation than more generic frameworks. For the growing base of developers working on MacBook Pros and Mac Studios, this creates a viable path to faster local inference without external accelerators.

The release strengthens Apple Silicon as a serious option for local LLM development and deployment. As M-series chips continue to gain memory capacity and computational power, frameworks like Lily make them competitive alternatives to cloud inference for latency-sensitive applications, model experimentation, and privacy-first deployments.

Read the full article on Google News.


Source: Google News · Relevance: 8/10