Magnitude has launched an open-source inference engine designed to optimize model performance for AI agents based on the user's exact hardware. By compiling and tuning kernels directly on the device, the engine aims to run open models up to 2x faster than llama.cpp.
Magnitude Launches as Self-Optimizing Open Source Inference Engine for AI Agents
The engine is compatible with various hardware, including Apple Silicon, NVIDIA, AMD GPUs, and CPUs. It is distributed as a desktop application that includes the Magnitude CLI, requiring no separate installation.
Magnitude features one-click connectivity for several existing agents, such as Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. It also supports any agent through an OpenAI-compatible API. For security and privacy, the engine is designed to keep prompts, files, and models on the local machine, requiring no internet connection once a model is downloaded.
Sources
- Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents (Hacker News Frontpage, 2026-09-30)
- Magnitude