Magnitude · 2026-09-30 · major
Magnitude — a local inference engine for agents, up to 2x faster than llama.cpp
Magnitude is an open-source Rust inference engine that tunes its kernels on your own device to run open models for local coding agents. v0.2.0 shipped its own engine, and the Launch HN reached the front page.
An open inference engine that tunes itself to your machine so local agents run open models faster.
Key specs
| GitHub stars | 5,831 |
|---|---|
| Decode vs llama.cpp (metal) | +92% |
Quick facts
| Maker | Magnitude (YC S25) |
|---|---|
| License | Apache-2.0 |
| Version | CLI v0.2.0 (Sep 30), now 0.2.3 |
| Platforms | macOS, Linux, Windows |
| Hardware | Apple Silicon, NVIDIA, AMD or CPU |
| Local API | OpenAI- and Anthropic-compatible, port 10100 |
| Price | Free |
What is it?
Magnitude v0.2.0 replaces the app's old llama.cpp backend with its own inference engine, written in Rust with a custom GPU kernel runtime and autotuner. It ships as a desktop app with a CLI that downloads open-weight models and starts them only when a connected agent such as Pi, OpenCode, Codex or Claude Code needs one.
How does it work?
Kernels have flexible parameters that are tuned on your actual hardware before a model runs, and they are hand-written for the most popular open-weight families. Memory is reserved only for the weights; the heap grows with each agent session and is freed when the agent stops. A hybrid paged attention design lets parallel sessions share prefix caches without slowing a single session.
Why does it matter?
Local agents run long sessions, often several at once, and general engines trade that speed away. In the makers' test on Qwen 3.6 35B A3B, decode went from 30 to 57 tokens/s on an M4 Pro and from 49 to 58 on a DGX Spark, with 27% less memory per agent.
Who is it for?
developers running coding agents on local models
Frequently asked questions
- How much faster is Magnitude than llama.cpp?
- Magnitude's makers benchmarked it against llama.cpp on Qwen 3.6 35B A3B at 4-bit with 64k context and no speculative decoding. On a Mac M4 Pro, decode was 92% faster (30 to 57 tokens/s) and prefill 9% faster. On a DGX Spark with CUDA, decode was 19% faster and prefill 23% faster. These are the makers' own numbers.
- Which coding agents work with Magnitude?
- Magnitude has one-click connections for Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline. Any other tool can use Magnitude's local API, which accepts both OpenAI-compatible and Anthropic-compatible requests at http://127.0.0.1:10100, so most agent harnesses can point at it without code changes.
- Does Magnitude cost anything or send data to the cloud?
- Magnitude is free and open source under the Apache-2.0 license. The project says there are no token costs and nothing leaves your machine: after a model is downloaded, inference runs fully on-device. Remote access over a network is optional and needs the server's configured API key.
- Can Magnitude run on a server without a desktop?
- Yes. Magnitude v0.2.0 added the magnitude serve command, which runs the inference service in the foreground without opening the desktop window on macOS, Windows and Linux. The Linux packages support Ubuntu 22.04, Debian 12 and Fedora or RHEL 9+ with glibc 2.35 or later.
- What is on Magnitude's roadmap?
- Magnitude's founders listed three planned features in their Launch HN post: expert streaming that keeps mixture-of-experts weights in RAM or on disk and loads them just in time, a full kernel compiler that picks fusions for your hardware, and multi-device use of CPU, GPUs, RAM and disk together.
Try it
https://magnitude.dev/download