GGUFs in transformers natively!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey there folks! Aritra here from Hugging Face. I wanted to update you all about the latest changes in `transformers`. We now natively support GGUFs (llama cpp quants). You can use it like so: After loading, you're using the normal Transformers APIs. Why did we want to do this?
On supported Apple Silicon setups, we're also reusing ggml kernels so the model can run directly from its packed quantized weights. On the Qwen checkpoints we tested on an M2 Max, Transformers reached:
This isn't meant to replace llama.cpp. If you only care about maximum local inference performance, llama.cpp is still probably the better choice. The point is more that you can now use the same GGUF models in a more flexible environment. Read more: https://huggingface.co/blog/transformers-llama-cpp-quants [link] [comments] |
More from r/LocalLLaMA
-
Pirate Face - pirate bay for LLMs
Sep 23
-
DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic
Sep 23
-
Nathan Lambert's written Congressional testimony on the state of open models - Chinese open-weight downloads now 2x America's, >80% of OpenRouter open-model usage
Sep 23
-
The attack on open weights continues to grow.
Sep 23
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.