English

Product LaunchesHugging FaceTransformers

Transformers Adds Support for Efficient GGUF Model Execution

Hugging Face has added support for running GGUF models efficiently within the Transformers library. This update allows users to run large language models optimized for local memory directly through the familiar Transformers APIs by loading checkpoints from the Hugging Face Hub.

GGUF, developed by the llama.cpp team, is a widely used format for local inference that packages model weights and metadata into a single file. It supports various quantization levels, allowing users to trade a certain amount of precision for a smaller memory footprint.

To achieve performance close to llama.cpp, Transformers reuses underlying ggml kernels through the kernels library and reduces overhead in the generate function. The initial focus of this integration is local inference on Apple Silicon, with the Qwen3.5 architecture as the first implementation. For users running on Metal, the loader automatically uses compatible ggml/Metal layer kernels.

The integration aims to bridge the gap between llama.cpp, which serves as a foundation for local inference, and Transformers, which serves as a foundation for model definition. By bringing ggml's performance to models that llama.cpp might not natively support, this approach expands the opportunity to accelerate a wider range of architectures, including research models and new multimodal modalities.

Sources

  1. Transformers now runs llama.cpp quants (Hugging Face Blog, 2026-09-22)
  2. Update on GitHub