English

Releases

llama.cpp v0.5.0 Public Release — Backend performance optimizations and expanded model support

llama.cpp has released version 0.5.0, focusing on backend performance, correctness, and broader model coverage. The update includes optimizations for CUDA and Metal, more robust server operations, and support for several new model architectures.

New Features and Improvements

  • Performance Optimizations:
    • Accelerated CUDA conv2d using implicit GEMM.
    • Added Metal MoE and SSM_CONV fusion optimizations.
    • Enabled CUDA graphs for MTP drafting.
  • New Model Support:
    • Added support for HRM-Text (DFM Mimir 1B).
    • Added conversion support for MiMo-V2.6 and support for HunyuanOCR via DFlash.
    • Extended support for Nemotron MTP and Nemotron-H models.
    • Added Qwen4Exp hyper-connection operations and sparse flash attention.
  • Server and API Enhancements:
    • Allowed the server to bind to multiple addresses via --host.
    • Added input_image support to server function-call outputs.
    • Improved JSON Schema and PEG handling.
    • Added environment variables for temperature, top-p, min-p, and penalties.

Bug Fixes

  • Fixed token counting API crashes during sleep in the server.
  • Fixed tensor-parallel split state and granularity issues for fused QKV models.
  • Fixed SigLIP buffer overruns for tall/wide images.
  • Resolved various UI issues, including mobile breakpoint and content overflow problems.
  • Fixed several parser errors for Ling 3.0, DeepSeek V3.2/V4, qwen3-coder, Muse Glimmer, and Gemma 4.

Sources

  1. v0.5.0 (2026-09-23)