Next upAI x Bio Pitch Contest
News

Hugging Face adds llama.cpp GGUF support to Transformers

Transformers can now load and run llama.cpp-style GGUF quantized checkpoints through its standard Python API, with initial acceleration focused on Apple Silicon.

D
Sep 25, 2026 · 1 min read

Hugging Face has added support to Transformers for loading and running GGUF quantized checkpoints used by llama.cpp through the library’s standard from_pretrained API. The change lets developers on Apple Silicon Macs use those checkpoints in Transformers without first converting them to another format.

The initial implementation uses Metal GPU kernels delivered through Hugging Face’s kernels library. This brings the llama.cpp quantization ecosystem into Transformers while preserving its existing Python model-loading interface.

Developers select GGUF model files by passing a filename through the gguf_file argument to AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename). Hugging Face also shows the format working with the Transformers command-line server, where a model identifier is paired with a specific .gguf file.

Hugging Face says a test of Qwen3.5-4B using the Q4_K_M quantization on a MacBook Pro with an M2 Max processor and 32GB of memory delivered throughput close to llama.cpp. Because the result has not been independently verified, it remains a vendor-reported comparison rather than a broader performance finding.

More news