llama.cpp releases · June 19, 2026 · 1 min read

b9717

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

ggml-cpu: support K tails in power10 Q8/Q4 MMA matmul (#24753)

ggml-cpu: support K tails in Power10 MMA Q8/Q4 matmul

This patch removes the requirement that K be divisible by kc in the tinyBlas_Q0_PPC tiled matmul path. Process the final K panel using its actual depth and pass the reduced panel size through packing and kernel execution. This allows more workloads to use the MMA kernel and reduces fallback to mnpack.

Apply suggestion from @taronaeo

Co-authored-by: Aaron Teo [email protected]

macOS/iOS:

macOS Apple Silicon (arm64)
macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
macOS Intel (x64)
iOS XCFramework

Linux:

Android:

Android arm64 (CPU)

Windows:

openEuler:

DISABLED
openEuler x86 (310p)
openEuler x86 (910b, ACL Graph)
openEuler aarch64 (310p)
openEuler aarch64 (910b, ACL Graph)

UI:

Discussion (0)

No comments yet. Sign in and be the first to say something.

Discussion (0)

More from llama.cpp releases