llama.cpp releases
500 articles archived · Visit source ↗ · RSS
-
llama.cpp releases dev-tools 1h ago
b11227
context : do not re-reserve the scheduler when toggling causal_attn ( #28751 ) context : do not re-reserve the scheduler when toggling causal_attn llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this…
27 -
llama.cpp releases dev-tools 2h ago
b11226
Enables Windows ARM64 build with MSVC cl.exe ( #28362 ) can reproduce the issue vlad sees fix fma issue drop volatile fix volatile runtime task add arm flag if needed fix hsum compile error fix syntax in quants strengthen sve probing make the syntax fixes one liners remove debug…
13 -
llama.cpp releases dev-tools 3h ago
b11225
tests : fix ggml init ( #29554 ) tests : init ggml for test-recurrent-state-rollback cont : same for test-save-load-state cont : add to test-state-restore-fragmented + add TODOs Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50700507…
14 -
llama.cpp releases dev-tools 4h ago
b11224
vulkan: fix wrong results when a mul_mat reads a slice of a larger cache ( #28956 ) vulkan: read the batch stride of an in place src0 from nb[2] A dim01 contiguous tensor can still be a view whose batches are strided by more than ne[1] rows, the first rows of a KV cache for…
25 -
llama.cpp releases dev-tools 13h ago
b11223
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) ( #28876 ) server : allow splitting RANK pooling for causal LLM rerankers Rerank models fall into two categories: bidirectional cross-encoders (BERT, etc.) that require all tokens in a…
26 -
llama.cpp releases dev-tools 17h ago
b11222
common : avoid side effects around params parsing ( #29537 ) register --rpc unconditionally and call llama_supports_rpc() only from its handler print server "initialization ..." log after args are parsed Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL Website: https://llama.app…
32 -
llama.cpp releases dev-tools 18h ago
b11221
common : make string_split throw on invalid values ( #29518 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50573055 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon…
32 -
llama.cpp releases dev-tools 20h ago
b11218
jinja : add support for dict builtin ( #29477 ) add support for dict builtin add tests Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50563209 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)…
15 -
llama.cpp releases dev-tools 20h ago
b11217
opencl: refine bin kernel loading condition ( #29503 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50560500 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
19 -
llama.cpp releases dev-tools 21h ago
b11216
sycl: FWHT kernels for block widths above 512 ( #29243 ) The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus 384/640/768/1280 via the Kronecker/Paley construction added separately in Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through…
17 -
llama.cpp releases dev-tools 21h ago
b11215
CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 ( #26289 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50554486 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS…
28 -
llama.cpp releases dev-tools 22h ago
b11214
HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes ( #28907 ) HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes CI: hip-quality-check: ignore spills for very large mfma mma kernels Website: https://llama.app Attestations:…
34 -
llama.cpp releases dev-tools 23h ago
b11213
vulkan: fix argsort kernel selection for Adreno ( #29469 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50544436 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
21 -
llama.cpp releases dev-tools 23h ago
b11212
common : throw instead of abort on grammar without llguidance ( #29516 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50541746 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
14 -
llama.cpp releases dev-tools 1d ago
b11211
RPC: use RDMA completion channel to not spin ( #29440 ) RPC: use RDMA completion queue to not spin add TODO for apple RDMA Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50534798 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
25 -
llama.cpp releases dev-tools 1d ago
b11209
llama-bench : fix OOB access of hf_file ( #29515 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50528155 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…
8 -
llama.cpp releases dev-tools 1d ago
b11208
hrm : fix layer placement of z_l_init weight ( #29512 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50522371 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
16 -
llama.cpp releases dev-tools 1d ago
b11207
hexagon: support tiled Q4_0 and Q8_0 GET_ROWS ( #29511 ) hexagon: support tiled Q4_0 and Q8_0 GET_ROWS hex-get-rows: fix macros hex-get-rows: use tiled HVX dequantization Assisted-by: OpenCode hex-get-rows: fix register spills and clean up checks for unsupported ops…
24 -
llama.cpp releases dev-tools 1d ago
b11206
hexagon: support for backend sampler ( #29502 ) hex-topk: trying to improve/cleanup the pipeline hex-sampling: add STEP op hex-sampler: add SUM op hex-sampler: update CPY to support sampling cases hex-binary: add support for chunking to handle large logits hex-argmax: super…
37 -
llama.cpp releases dev-tools 1d ago
b11205
cuda: support Nemotron 3 Puzzle state size 96 for ssm scan ( #28717 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50461180 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
13 -
llama.cpp releases dev-tools 1d ago
b11203
cuda: add F16 input to the FWHT ( #29096 ) cuda: add F16 input to the FWHT The CUDA FWHT accepts F32 input only. This makes the source type a template parameter, so the kernel reads an F16 source directly instead of requiring a converted copy. The F32 path is unchanged.…
33 -
llama.cpp releases dev-tools 1d ago
b11202
server : fix wake_fd warning on Windows ( #29479 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50455005 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…
19 -
llama.cpp releases dev-tools 1d ago
b11201
Revert "Change max context length for auto-fitting with unified KV ( #28849 )" ( #29437 ) This reverts commit b04d4e5 . Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50433369 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon…
17 -
llama.cpp releases dev-tools 2d ago
b11200
jinja : implement sameas test ( #29448 ) implement sameas test add tests Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50397675 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
17 -
llama.cpp releases dev-tools 2d ago
b11199
jinja : fix compile error ( #29468 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50395469 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…
36 -
llama.cpp releases dev-tools 2d ago
b11195
ggml-cpu: tiled mul_mat for k-quants ( #27851 ) Added tiled mul_mat. For each mul_mat_one_chunk, quants are unpacked into (max) 256x256 tiles of int8, one routine per quent. Then microkernel computes 16x16 tiles before writing out 256x256 float reults to main memory.…
31 -
llama.cpp releases dev-tools 2d ago
b11194
opencl: add A8 Q8_0 non-MoE dp4a binary kernel ( #29439 )
17 -
llama.cpp releases dev-tools 2d ago
b11193
hexagon: find software divide calls using binary inspection tool ( #29449 ) hex-scripts: fix table alignment hex-scripts: find sw div calls using binary inspection tool Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50359675…
28 -
llama.cpp releases dev-tools 2d ago
b11192
vendor : update cpp-httplib to 0.58.0 ( #29407 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50337762 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
22 -
llama.cpp releases dev-tools 2d ago
b11191
common,rpc : simplify fs_create_directory_with_parents() ( #29432 ) The original function was broken on Windows for some unicode paths Paths without a trailing separator now create the last directory too, matching the function name. All current callers already include a trailing…
26 -
llama.cpp releases dev-tools 2d ago
b11190
mtmd: fix mel preprocessor in LFM2 audio ( #29403 ) which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: uses log(x + 2^-24) instead of clamping to the log floor…
27 -
llama.cpp releases dev-tools 2d ago
b11189
opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin , kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin ( #29401 ) opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel opencl: fix s transpose - s only transposed for bin kernels Co-authored-by: Li He…
18 -
llama.cpp releases dev-tools 2d ago
b11188
Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support ( #29373 ) ( #29409 ) vulkan : fix build issue of legacy glslc version by adding GGML_VULKAN_COOPMAT_GLSLC_SUPPORT macro check for Intel FA shader compiling vulkan : add preprocess…
21 -
llama.cpp releases dev-tools 2d ago
b11185
common : extract shared unicode path/string helpers ( #29415 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50264913 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon…
22 -
llama.cpp releases dev-tools 2d ago
b11184
metal: FWHT kernels for block widths above 512 ( #29095 ) metal: FWHT kernels for block widths above 512 The Metal FWHT covers widths 64 to 512, one row per simdgroup with N/32 values per lane. Wider blocks need more registers per lane than that layout allows. kernel_fwht_tg…
28 -
llama.cpp releases dev-tools 2d ago
b11182
llama : add llama_prec_policy + model-driven W4A4 path ( #24364 ) Rebase and update based on #26675 Signed-off-by: ynankani [email protected] CI failure fix(launh_bounds overload on HIP) and cleanup Signed-off-by: ynankani [email protected] Address review comments…
10 -
llama.cpp releases dev-tools 2d ago
b11181
HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 ( #29231 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50216799 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…
25 -
llama.cpp releases dev-tools 2d ago
b11180
rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes ( #29283 ) rpc : include nb in the get_alloc_size cache key and floor the result at ggml_nbytes cont : remove redundant comment cont : add TODO Co-authored-by: Georgi Gerganov [email protected]…
17 -
llama.cpp releases dev-tools 2d ago
b11179
[SYCL] support sparse FA ( #28796 ) fix conflict fix format issue rm unused code Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50188435 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED…
26 -
llama.cpp releases dev-tools 2d ago
b11178
musa: fix PH1 (MTT S5000) operator failures and build issues ( #29193 ) musa: use 16-byte copies for MUSA like sm_70+ ggml_cuda_get_max_cpy_bytes() derives the copy width from CUDA_ARCH . mcc never defines it, so MUSA fell into the generic branch and returned 8 bytes instead of…
35 -
llama.cpp releases dev-tools 3d ago
b11177
CUDA: fuse RMS_NORM + SCALE into one kernel ( #29393 ) #28068 builds the GDN q/k l2norm as ggml_scale(ggml_rms_norm(x, eps/n), 1/sqrt(n)). This adds 2 SCALE nodes per GDN layer, 96 extra kernel launches per ubatch on Qwen3.8-27B (48 GDN layers). The extra kernels take no…
38 -
llama.cpp releases dev-tools 3d ago
b11176
llama : fix tensor split for fused qkv with uneven K/V head sizes ( #2 …
6 -
llama.cpp releases dev-tools 3d ago
b11175
hexagon: add q5_k quant type support ( #29123 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50064037 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
8 -
llama.cpp releases dev-tools 3d ago
b11173
metal : fix graph capture and handle empty graphs ( #29390 ) return early when the graph has no nodes drop the redundant reset of capture_compute: the decrement at the top of the function already transitions the counter from 0 to -1, so a capture happens exactly once hint at…
27 -
llama.cpp releases dev-tools 3d ago
b11172
metal : optimize sparse FA + clean-up ( #29377 ) metal : cache sparse FA indices in shared memory Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp metal : simplify shared memory size calculation Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp pi : update general…
7 -
llama.cpp releases dev-tools 3d ago
b11171
sync : ggml ( #29396 ) ggml : bump version to 0.25.2 (ggml/1642) ggml : fix ubsan error in ggml_graph_nbytes (ggml/1644) ggml : bump version to 0.25.3 (ggml/1645) sync : ggml Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50020953…
21 -
llama.cpp releases dev-tools 3d ago
b11170
hexagon: handle multi-sequence in concat_2d ( #29344 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50014220 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
36 -
llama.cpp releases dev-tools 3d ago
b11169
llama-grammar: fix numeric truncation for token_id parsing ( #29382 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50007794 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
21