llama.cpp releases · · 1 min read

b10481: CUDA: MMVQ nwarps=8 for bs=1 for dense models on DGX Spark (#26843)

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

  • CUDA: MMVQ nwarps=8 for bs=1 for dense models on DGX Spark

Signed-off-by: ynankani [email protected]

  • skip moe experts and allow others based on k geometry (allow only small idle tail)

Signed-off-by: ynankani [email protected]

  • rename MMVQ DGX Spark params to GB10 and fix MSVC constexpr lambda capture

Signed-off-by: ynankani [email protected]


Signed-off-by: ynankani [email protected]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases