A new open-source inference engine named Magnitude has been released, specifically designed to optimize the execution of open models for AI agents. Unlike generalist engines, Magnitude compiles and tunes its kernels on the actual device being used, allowing the software to fit the specific chip and memory configuration of the hardware.
Magnitude: Open-Source Inference Engine Optimizes AI Agent Performance via On-Device Kernel Tuning
In benchmarks comparing the engine to llama.cpp using a 4-bit quantized "Qwen 3.6 35B A3B" model with a 64K token context length, Magnitude demonstrated significant performance gains. On an Apple M4 Pro with 48GB of memory, Magnitude recorded a generation speed of approximately 57 tokens per second, a 92% increase over llama.cpp's 30 tokens per second. In a CUDA environment using an NVIDIA DGX Spark, Magnitude achieved approximately 58 tokens per second for generation and 2507 tokens per second for prefill, representing increases of 19% and 23% respectively.
The engine also optimizes memory usage during long agentic sessions. It allocates only the memory required to hold model weights initially and dynamically increases the allocation as the session progresses. Once an agent stops, the allocated memory is released to allow the user to continue other tasks on the same PC. Additionally, Magnitude implements a "prefix caching" mechanism that allows multiple sessions to reuse computation results for common system prompts or repetitive text, improving efficiency for AI agents that frequently pass similar instructions.
Magnitude supports macOS, Windows, and Linux, and is compatible with Apple Silicon, NVIDIA GPUs, AMD GPUs, and CPUs. It offers one-click connectivity to various agent tools, including Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, as well as an OpenAI-compatible API. Once models are downloaded, the system functions entirely offline, ensuring that prompts and files remain on the local machine.
Sources
- ハードウェアに合わせて自動最適化してオープンモデルを最大2倍高速化するAIエージェント向け推論エンジン「Magnitude」、Appleシリコン・NVIDIA・AMD・CPUに対応 (GIGAZINE, 2026-10-01)
- GitHub