Hugging Face adds llama.cpp GGUF support to Transformers
Transformers can now load and run llama.cpp-style GGUF quantized checkpoints through its standard Python API, with initial acceleration focused on Apple Silicon.
Hugging Face has added support to Transformers for loading and running GGUF quantized checkpoints used by llama.cpp through the library’s standard from_pretrained API. The change lets developers on Apple Silicon Macs use those checkpoints in Transformers without first converting them to another format.
The initial implementation uses Metal GPU kernels delivered through Hugging Face’s kernels library. This brings the llama.cpp quantization ecosystem into Transformers while preserving its existing Python model-loading interface.
Developers select GGUF model files by passing a filename through the gguf_file argument to AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename). Hugging Face also shows the format working with the Transformers command-line server, where a model identifier is paired with a specific .gguf file.
Hugging Face says a test of Qwen3.5-4B using the Q4_K_M quantization on a MacBook Pro with an M2 Max processor and 32GB of memory delivered throughput close to llama.cpp. Because the result has not been independently verified, it remains a vendor-reported comparison rather than a broader performance finding.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
