🔍 Read the full analysis: A New Way To Run Llama.cpp Quants With Transformers on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Hugging Face has added support for loading GGUF-quantized checkpoints through Transformers’ familiar from_pretrained API. The current implementation targets Apple Silicon and Qwen3.5, is available on the main branch, and uses llama.cpp’s ggml kernels when compatible kernels are available.
Hugging Face has added GGUF model loading to its Transformers library, as detailed in the original analysis, allowing users to run compatible quantized checkpoints through the familiar from_pretrained interface. The initial implementation targets Apple Silicon Macs and the Qwen3.5 architecture, and is currently on the library’s main branch ahead of a stable release.
Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the gguf_file argument to from_pretrained. The library can then generate text from the loaded model. Hugging Face says the implementation reuses llama.cpp’s ggml kernels, the low-level code used to run models in the llama.cpp ecosystem.
On supported Apple Silicon systems, when model weights remain packed on Metal, Transformers automatically loads compatible ggml/Metal layer kernels and uses ggml-org/ggml-attn for attention. If that attention kernel is unavailable, the library falls back to PyTorch’s standard sdpa implementation and issues a warning. Users can also select sdpa directly. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory.
The stated requirements include an Apple Silicon Mac, a supported PyTorch version, and recent Transformers and compatible kernel library versions. Hugging Face says its published kernel builds generally support the two latest PyTorch releases. The same checkpoints can also be served using transformers serve, which provides an OpenAI-compatible API on the user’s machine for clients configured to connect to that endpoint.
GGUF Comes to Transformers Workflows
The addition lets developers who already use Transformers access GGUF checkpoints without relying on a separate llama.cpp-based application for model loading. GGUF is widely used for local inference and is available through model publishers and communities including Unsloth, LM Studio Community, bartowski and ggml-org. The new loader may make those checkpoints easier to use in existing Transformers workflows, though current device and architecture support is limited.
Quantization reduces model file size, which can make local use possible on hardware with less memory. Hugging Face’s example for Unsloth’s Qwen3.5-4B lists a BF16 file at 8.42 GB, compared with 2.74 GB for Q4_K_M. Smaller files can lower memory demands, but the resulting quality depends on the model, quantization level and task. Hugging Face advises users to evaluate the options on their own workloads.
Hugging Face says it compares performance with llama.cpp across three GGUF checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. The announcement identifies llama.cpp as its reference for local inference. Performance depends on the hardware and model, so the comparison does not establish a single result for every user’s system.
Top picks for "llama quant transformer"
As an affiliate, we earn on qualifying purchases.
A Familiar Format, New Loader
GGUF is a file format developed for the llama.cpp ecosystem. It can package model weights and related information, such as tokenizer data and, optionally, a chat template. Checkpoints are available at different quantization levels. A Q4_K_M file uses mostly four-bit weights while retaining higher precision for some tensors; higher-bit variants generally take more storage.
Hugging Face’s size examples for Qwen3.5-4B include Q6_K at 3.53 GB and Q5_K_M at 3.14 GB, alongside the BF16 and Q4_K_M files. The company recommends beginning with Q4_K_M and trying Q5_K_M or Q6_K when more memory is available. Those are recommendations, not guarantees of a particular balance between output quality and speed.
The announcement follows public demonstrations of local models running in coding tools. Hugging Face co-founder Julien Chaumond recently described Qwen3.6 27B running through llama.cpp in the Pi coding agent on a MacBook Pro. That demonstration was a separate report about a different model and setup; it does not establish the performance of the new Transformers implementation.
“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”
— Hugging Face announcement
Support Still Has Narrow Limits
The announcement’s initial support is limited to Apple Silicon and Qwen3.5. It does not give a timeline for CUDA, Linux or Windows support, or specify when additional model architectures will be supported. The feature is on the Transformers main branch; Hugging Face has not announced when it will reach a stable release.
The announcement refers to benchmark comparisons against llama.cpp but the results depend on the selected models and hardware. Readers should consult the published benchmark details before applying them to a different setup. It also remains unclear how quickly the project will add architectures or hardware backends. The expected quality impact of a quantization level varies by model and task, and requires evaluation against the intended workload.
Stable Release and Broader Support
The next milestone is a stable Transformers release that includes the feature; no release date has been provided. Until then, users can access it from the main branch, subject to the stated software and hardware requirements. Users considering a particular checkpoint should check its format and compatibility, confirm that the required kernels are available, and account for the higher memory use that can result from dequantization.
Further developments to watch include support for additional model architectures and devices. Hugging Face has not announced a roadmap or dates for those additions. Its GGUF documentation and kernel library are the places to check for changes to supported quantization types and runtime requirements. Benchmark results and user testing across more configurations will help clarify how closely performance matches llama.cpp in practice.
Key Questions
How do I load a GGUF checkpoint?
On a supported setup, pass the checkpoint file using the gguf_file argument to Transformers’ from_pretrained method. The current feature is on the main branch and has specific hardware and software requirements.
Which hardware and model architecture are supported initially?
The announcement identifies Apple Silicon Macs and Qwen3.5 as the initial targets. It gives no timeline for other devices or architectures.
Does Transformers use llama.cpp code?
Hugging Face says the implementation reuses llama.cpp’s ggml kernels for compatible operations. If a compatible attention kernel is unavailable, it can fall back to sdpa.
Will a smaller quantized file always give the same results?
No. Hugging Face says the quality impact depends on the model and task. Smaller files can reduce memory requirements, but users should assess output quality on their own workloads.
When will this be included in a stable release?
Hugging Face has not announced a stable release date. The feature is currently available on the Transformers main branch.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
