Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
Abstract
A framework quantizes vision-language models for mobile deployment using self-generated training data and a 2.7-bit format, compressing Llama 3.2 11B Vision Instruct to 3.7 GB with preserved visual question answering performance.
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.
Get this paper in your agent:
hf papers read 2608.21134 curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper
No model linking this paper
Datasets citing this paper
No dataset linking this paper
Spaces citing this paper
No Space linking this paper
Collections including this paper
More from Hugging Face Daily Papers
-
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Aug 24
-
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
Aug 24
-
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
Aug 24
-
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
Aug 24
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.