Magnitude is an Apache-2.0 open-source on-device inference engine for agents that optimizes itself for the user’s exact hardware. It ships as a desktop app that runs open models and connects the agent you already use; the app includes the magnitude CLI with no separate install (magnitude.dev, GitHub). Launch HN discussion lands around Sep 30, 2026, aligned with the CLI 0.2.0 own-engine ship (Launch HN).
This is local inference on your machine—not a hosted model router or cloud multi-node cluster product, and not a rebrand of llama.cpp, Ollama, or MLX.
Kernels tuned on your device
Where generalist engines often ship precompiled kernels for broad hardware classes, Magnitude says it compiles and tunes its kernels on your actual device before a model runs so they fit the chip in front of you (magnitude.dev). Concurrent sessions can share prefix caches to limit slowdown; memory flexes and is freed when agents stop, as the project describes.
Platforms, privacy, and agents
Supported operating systems: macOS, Linux, and Windows. Hardware backends per the FAQ: any Apple Silicon, NVIDIA, or AMD GPU, or CPU-only—no fixed minimum; smaller machines run smaller models (magnitude.dev).
Privacy framing from the primary: prompts, files, and models stay on the machine; no internet is needed once a model is downloaded; no token costs, nothing leaves your machine, under Apache 2.0.
One-click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works through an OpenAI-compatible API (magnitude.dev).
CLI 0.2.x — Magnitude’s own engine
GitHub releases mark @magnitudedev/[email protected] (2026-09-30) as replacing the llama.cpp-based inference path with Magnitude’s own engine, which automatically optimizes itself for the hardware. Patches run through the 0.2.x family (tip 0.2.3 on 2026-10-01 includes a Windows installer fix). Prefer the 0.2.x / own-engine frame over pinning a single patch unless you are citing a release note (releases).
Vendor speed claims (as Magnitude states)
All tok/s and percentage figures below are Magnitude’s site measurements—not independently reproduced here. Fixture label: Qwen 3.6 35B A3B, 4-bit, 64k context, no speculative decoding (magnitude.dev):
| Surface | Decode | Prefill |
|---|---|---|
| Metal Mac M4 Pro 48 GB | 30 → 57 tok/s (92% faster) | 466 → 507 tok/s (9% faster) |
| CUDA DGX Spark | 49 → 58 tok/s (19% faster) | 2,033 → 2,507 tok/s (23% faster) |
Headline vendor line: up to 2× faster than llama.cpp; memory 27% less per agent on the primary. Treat those as attributed marketing benches. Community threads also debate Metal baselines and other engines—useful color, not a newsroom verdict.
Rust crossover (one note)
The inference engine workspace is Rust (Cargo tree under inference-v4); founders on Launch HN describe a custom GPU kernel runtime and autotuner (Launch HN, repo). That is engine-language color for a tools / cli local-agents story—not a Rust-framework desk lead.
Who should care
Builders who want private, on-device open-model inference wired into existing coding agents can start at magnitude.dev and the GitHub repo. Attribute every tok/s claim to Magnitude; do not invent cloud pricing from this launch—the primary publishes free / OSS / no token costs for the on-device path.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.