Magnitude (YC S25) Launches Self-Compiling Inference Engine, Claims 92% Faster Metal Decode Than llama.cpp
Summary
Y Combinator S25 startup Magnitude launched its open-source inference engine on Sept 30, compiling and tuning kernels on the target device before a model runs rather than shipping precompiled kernels. The project claims up to 2x throughput over llama.cpp — 92% faster decode on Metal and 19% faster on CUDA in their benchmarks — with support for Apple Silicon, Nvidia, AMD GPUs, and CPU-only systems across macOS, Linux and Windows.
Originally reported by github.com
Read the original article →Original headline: Magnitude (YC S25) Launches Self-Compiling Inference Engine, Claims 92% Faster Metal Decode Than llama.cpp