bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Don't get too hyped + take with a grain of salt as there have been an endless amount of quantization schemes with big promises that never really became a thing. Tim Dettmers is a pretty well known researcher though, so maybe something will come of this. Time shall tell. Another tweet about the method, DS4 Pro on a single B300 (288 GB VRAM): https://xcancel.com/Tim_Dettmers/status/2087624491362820364 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.