Published at TMLR 09/2026: <a href=\"https://openreview.net/forum?id=3nZb43fvAQ\" rel=\"nofollow\">https://openreview.net/forum?id=3nZb43fvAQ</a></p>\n","updatedAt":"2026-09-09T07:13:07.449Z","author":{"_id":"668aeec93d34648deb2aad43","avatarUrl":"/avatars/684fc4b7fdcc9424940ca84865cc3fa4.svg","fullname":"Nan Jiang","name":"jiangnanhugo","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7261409759521484},"editors":["jiangnanhugo"],"editorAvatarUrls":["/avatars/684fc4b7fdcc9424940ca84865cc3fa4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2510.08999","authors":[{"_id":"6aa106a6d0174964227beeab","name":"Ziyi Wang","hidden":false},{"_id":"6aa106a6d0174964227beeac","user":{"_id":"668aeec93d34648deb2aad43","avatarUrl":"/avatars/684fc4b7fdcc9424940ca84865cc3fa4.svg","isPro":false,"fullname":"Nan Jiang","user":"jiangnanhugo","type":"user","name":"jiangnanhugo"},"name":"Nan Jiang","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.809Z","hidden":false},{"_id":"6aa106a6d0174964227beead","name":"Guang Lin","hidden":false},{"_id":"6aa106a6d0174964227beeae","name":"Qifan Song","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/668aeec93d34648deb2aad43/YusXl4e3OC5kN6I3iE_EL.jpeg"],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions","submittedOnDailyBy":{"_id":"668aeec93d34648deb2aad43","avatarUrl":"/avatars/684fc4b7fdcc9424940ca84865cc3fa4.svg","isPro":false,"fullname":"Nan Jiang","user":"jiangnanhugo","type":"user","name":"jiangnanhugo"},"summary":"Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization individually, often resulting in suboptimal compression rates to preserve acceptable performance drops. We introduce a unified framework for simultaneous pruning and low-bit quantization via Bayesian variational learning (\\method), which achieves higher compression rates than prior baselines while maintaining comparable performance. The key idea is to employ a spike-and-slab prior to induce sparsity and model quantized weights using Gaussian Mixture Models (GMMs) to enable low-bit precision. Due to the intractability of the objective involving spike-and-slab priors with GMMs, we derive an efficient approximation that facilitates effective compression with minimal accuracy loss. In theory, we provide a consistent result for our proposed variational approach to a sparse and quantized deep neural network. Extensive experiments on compressing ResNet, BERT-base, Llama3.2, and Qwen2.5 models show that our method achieves higher compression rates than a line of existing methods with comparable performance drops. Project page: https://comeusr.github.io/SQS_Webpage.","upvotes":14,"discussionId":"6aa106a6d0174964227beeaf","projectPage":"https://comeusr.github.io/SQS_Webpage/","githubRepo":"https://github.com/comeusr/SQS_TMLR","githubRepoAddedBy":"user","ai_summary":"A unified Bayesian variational framework combining spike-and-slab sparsity and Gaussian mixture quantization achieves high compression rates for large neural networks with minimal accuracy loss.","ai_keywords":["Bayesian variational learning","spike-and-slab prior","Gaussian Mixture Models","pruning","low-bit quantization","sparse and quantized deep neural network"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6400300fe7767a89533fa2c0","name":"Purdue","fullname":"Purdue University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677733893266-64002f32cafc9d54986439cb.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a6a81c9ad5a6f2f636078bf","avatarUrl":"/avatars/c6156e9fbbac708d46c6c24a665fa1ac.svg","isPro":false,"fullname":"Steven Martinez","user":"steven-martinez","type":"user"},{"_id":"6a6c8a48e582cd54da2aad70","avatarUrl":"/avatars/113d731d611604d4c648c6c90f8eb524.svg","isPro":false,"fullname":"Thomas Thomas","user":"velvetThomas","type":"user"},{"_id":"6a6d58512a48f063703b8d3d","avatarUrl":"/avatars/25aa979d4bd497cee90a5429fa13f354.svg","isPro":false,"fullname":"Richard Martinez","user":"Quiet-Richard","type":"user"},{"_id":"6a7d427a0b4fb44febcd1f5e","avatarUrl":"/avatars/9f4e0e073d07d911a61a2d48c45e7e9c.svg","isPro":false,"fullname":"Matthew Lewis","user":"Matthew-Lewis","type":"user"},{"_id":"6a9ae02eaf1530c75fdaada2","avatarUrl":"/avatars/d525cbd62bec3e03a54c5bb43e0234a7.svg","isPro":false,"fullname":"渡辺 美加子","user":"kanakobayashi","type":"user"},{"_id":"6aa0da3acd38c4d253cfe8dc","avatarUrl":"/avatars/f5aca34f6f702b30bd4a11112d6ebce9.svg","isPro":false,"fullname":"杨楠","user":"NovaFanghou","type":"user"},{"_id":"6aa0f0b05efccda06029cfa7","avatarUrl":"/avatars/e13f0437aefc37a5bea0dac58311ffad.svg","isPro":false,"fullname":"高橋 亮介","user":"CedarEvan","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a6aa002977fbfce4bad937b","avatarUrl":"/avatars/04ced1974d585c0bd7507cda07ae61c7.svg","isPro":false,"fullname":"Charles White","user":"zenithLens","type":"user"},{"_id":"6a6c8c10c9a43ea10742deee","avatarUrl":"/avatars/9b3742550832094a81800eb88787eee8.svg","isPro":false,"fullname":"Joshua Anderson","user":"graniteGlade","type":"user"},{"_id":"6a6dea1d067f0e2726f81e06","avatarUrl":"/avatars/02afbceec3ca6047937af545cccef437.svg","isPro":false,"fullname":"Edward Lee","user":"Azure-Leo","type":"user"},{"_id":"6a9ae3eae09eec58fe776b9d","avatarUrl":"/avatars/1427b9015368d6b47a6110f28507217f.svg","isPro":false,"fullname":"中村 太郎","user":"Granite-Rika49","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6400300fe7767a89533fa2c0","name":"Purdue","fullname":"Purdue University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677733893266-64002f32cafc9d54986439cb.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2510/2510.08999.md","query":{}}">
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
Abstract
A unified Bayesian variational framework combining spike-and-slab sparsity and Gaussian mixture quantization achieves high compression rates for large neural networks with minimal accuracy loss.
Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization individually, often resulting in suboptimal compression rates to preserve acceptable performance drops. We introduce a unified framework for simultaneous pruning and low-bit quantization via Bayesian variational learning (\method), which achieves higher compression rates than prior baselines while maintaining comparable performance. The key idea is to employ a spike-and-slab prior to induce sparsity and model quantized weights using Gaussian Mixture Models (GMMs) to enable low-bit precision. Due to the intractability of the objective involving spike-and-slab priors with GMMs, we derive an efficient approximation that facilitates effective compression with minimal accuracy loss. In theory, we provide a consistent result for our proposed variational approach to a sparse and quantized deep neural network. Extensive experiments on compressing ResNet, BERT-base, Llama3.2, and Qwen2.5 models show that our method achieves higher compression rates than a line of existing methods with comparable performance drops. Project page: https://comeusr.github.io/SQS_Webpage.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2510.08999 in a model README.md to link it from this page.
Cite arxiv.org/abs/2510.08999 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2510.08999 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.