Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.
We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case. Inference engines today all make a performance tradeoff.
Magnitude is built for maximum performance on your hardware and running local agents: on-device compilation and tuning, focus on best architectures, dynamic memory allocation, and hybrid paged attention. Magnitude is fully open source (Apache 2.0). We built it in Rust, including a custom GPU kernel runtime and autotuner.
Benchmarked against llama.cpp with Qwen 3.6 35B A3B (4 bit), 64k context, no speculative decoding: Metal (Mac M4 Pro 48 GB) - 92% faster decode (30 tok/s → 57 tok/s), CUDA (DGX Spark) - 19% faster decode (49 tok/s → 58 tok/s). Magnitude ships as a desktop app that you can easily connect with whatever agents you already use.
Source: Hacker News · Summarized by HeadlinesBriefing