llama.cpp releases
500 articles archived · Visit source ↗ · RSS
-
llama.cpp releases dev-tools 3d ago
b11168
hexagon: dynamic quantizer improvements ( #29395 ) hexagon: fix accuracy issue in Q8_0 N=1 MUL_MAT hex-quant: fix register spills hex-mm: use dma for all dyn.quant paths Co-authored-by: Aparna M P [email protected] hex-mm: remove obsolete run_quant_task hex-mm: update…
35 -
llama.cpp releases dev-tools 3d ago
b11167
hexagon: support I32 CPY and CONT ( #29379 ) Assisted-by: OpenCode Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49994729 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)…
24 -
llama.cpp releases dev-tools 3d ago
b11166
cuda : add F16 kernel support for CONV_2D_DW ( #29064 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49985945 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
12 -
llama.cpp releases dev-tools 3d ago
b11165
test: flush status ( #28352 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49978694 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64…
12 -
llama.cpp releases dev-tools 3d ago
b11163
llama: add llama_batch_ext ( #24669 ) (wip) add llama_batch_ext wip updated design updated impl change signature unused var demo common_prompt_batch_decode fix pos tmp disable test-batch-alloc fix compat nits: add const no more pos_max add comment about…
18 -
llama.cpp releases dev-tools 3d ago
b11160
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 ( #27952 ) vulkan: add int8 coopmat quantized matmul shader apply scales inline use scalar sums probe and directly access coopmat values instead of going through shmem add q8_0 support add BK_STEP to shader,…
11 -
llama.cpp releases dev-tools 3d ago
b11159
vulkan: handle misalignment in conv_2d and conv_3d ( #29365 ) vulkan: handle misalignment in conv_2d and conv_3d fix test-backend-ops print Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49846182 macOS/iOS: macOS Apple Silicon (arm64)…
6 -
llama.cpp releases dev-tools 3d ago
b11158
vulkan: tune KHR cooperative matrix support for Adreno GPUs ( #29328 ) Enable coopmat support for Vulkan backend Fixed the mul_mat_s Removed the debug statement Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49831613 macOS/iOS: macOS…
35 -
llama.cpp releases dev-tools 4d ago
b11157
cuda : add conv3d with implicit GEMM ( #29137 ) cuda : add conv3d with implicit GEMM cuda : refine conv3d implicit GEMM and handle empty kernels Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49789544 macOS/iOS: macOS Apple Silicon…
21 -
llama.cpp releases dev-tools 4d ago
b11156
model : add Ling 3.0 VL support ( #29151 ) model : fold Ling 3.0 VL into the BailingMoeV3 architecture Assisted-by: Scout model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections Co-authored-by: aetherbird [email protected] Website:…
36 -
llama.cpp releases dev-tools 4d ago
b11155
server,common : fix the GCC 12 stringop-overread false positive (again) ( #29325 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49775037 macOS/iOS: macOS Apple Silicon (arm64) macOS…
34 -
llama.cpp releases dev-tools 4d ago
b11154
test-save-load-state : print a per-model results table in --models mode ( #29316 ) test-save-load-state : print a per-model results table in --models mode in --models mode the output was very heavy: every model printed its token dumps, per-test headers and PASS lines. instead,…
35 -
llama.cpp releases dev-tools 4d ago
b11153
hexagon: reject MUL_MAT_ID when src1 precision is F32 ( #29348 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49757881 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)…
11 -
llama.cpp releases dev-tools 4d ago
b11151: convert : allow vision target for DFlash/Dspark (#29339)
Resolve the target arch with get_model_architecture so vision targets (e.g. Lfm2VlForConditionalGeneration) map to their text model for the vocab. Fix double rope reorder for LFM2/LFM2.5 DSpark drafters
17 -
llama.cpp releases dev-tools 4d ago
b11149
tests: add -b/--backend option to test-llama-archs for testing a specific backend ( #27372 ) tests: add backend option to test-llama-archs Update tests/test-llama-archs.cpp Co-authored-by: Johannes Gäßler [email protected] remove extra space Co-authored-by: Johannes Gäßler…
24 -
llama.cpp releases dev-tools 4d ago
v0.5.0
Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP…
26 -
llama.cpp releases dev-tools 4d ago
b11147
opencl: add A8 Q6_K non-MoE dp4a binary kernel ( #29057 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49630180 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
26 -
llama.cpp releases dev-tools 4d ago
b11146
llama.cpp : bump version to 0.5.0 ( #29333 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49623059 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
33 -
llama.cpp releases dev-tools 4d ago
b11140
CUDA: enable sparse-fa for dsv4 prefill (again) ( #29298 ) CUDA: enable sparse-fa for dsv4 prefill (again) CUDA: unroll the query loop of the sparse mask scan The query loop of flash_attn_mask_to_sparse_indices has a runtime trip count, which keeps the unrolled scan over the…
35 -
llama.cpp releases dev-tools 4d ago
b11139
server: fix token counting API crash on sleep ( #29309 ) server: wake up sleeping server correctly server: wake up sleeping server correctly (local aliases removed) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49582694 macOS/iOS:…
26 -
llama.cpp releases dev-tools 4d ago
b11138
jinja : parse unary +/- before variables ( #29244 ) jinja : parse unary +/- before variables Lexer already emits unary_operator for -n / +n, and runtime executes unary -. Parse them at multiplicative precedence so slices like items[:-n] and GigaChat indent[:-indent_factor] work.…
8 -
llama.cpp releases dev-tools 4d ago
b11136
server: accept OpenAI video_url content type and data: video URIs ( #27921 ) The OpenAI chat completions API specifies content part type "video_url" with a {"url": ...} object, and clients typically send data: URIs (e.g. data:video/mp4;base64,...). The llama-server only accepted…
11 -
llama.cpp releases dev-tools 4d ago
b11132
model : support Gemma4 DSpark draft backbone ( #29226 ) dspark: add Gemma 4 draft support Add GGUF conversion and runtime support for full-attention and SWA Gemma 4 DSpark drafts, including tied output weights and boolean backbone metadata. Assisted-by: Codex dflash: infer Gemma…
8 -
llama.cpp releases dev-tools 4d ago
b11130
make-release : update summary prompt Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49525233 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…
30 -
llama.cpp releases dev-tools 4d ago
b11126
vulkan: add IQ4_XS MMQ/MMV matmul kernels ( #28415 ) vulkan: optimize IQ4_XS matmul kernels Assisted-by: OpenAI Codex vulkan: address IQ4_XS review nits drop the dead LOAD_VEC_A != 8 branch in the IQ4_XS shmem load; iq4_xs is in lut_load_vec_a()'s "8" list, so that path is never…
15 -
llama.cpp releases dev-tools 5d ago
b11125
ggml-meta: resolve multi buffer views ( #29266 ) ggml-meta: resolve multi buffer views add TODO to revisit if graph allocator gets refactored Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49502431 macOS/iOS: macOS Apple Silicon…
20 -
llama.cpp releases dev-tools 5d ago
b11124
cuda: top-k MoE should always fire ( #28432 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49495074 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
23 -
llama.cpp releases dev-tools 5d ago
b11123
sycl : support new UT case for mul_mat_hadamard fp16 ( #29218 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49484593 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)…
4 -
llama.cpp releases dev-tools 5d ago
b11122
sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions ( #28931 ) sycl : extend MMVQ GLU fusion to mixed quant types; add rms_norm+scale and ssm_conv+silu fusions fixing spacing issue and macro converted to template function Website: https://llama.app…
7 -
llama.cpp releases dev-tools 5d ago
b11121
sycl : support op get_rows_back, only support fp32/fp16 ( #25266 ) resovle confict support gedt_rows_back, update the ops.md Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49462636 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
15 -
llama.cpp releases dev-tools 5d ago
b11120
vulkan: hide internal symbols to prevent duplicate-dlopen state destruction ( #29139 ) Since #28732 our internal symbols are exported. A duplicate copy dlopened and dlclosed by ggml_backend_load_all() then interposes them, so its destructors destroy the live vk_instance and…
25 -
llama.cpp releases dev-tools 5d ago
b11119
sampler: reduce the size of the probe ( #29285 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49443624 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
33 -
llama.cpp releases dev-tools 5d ago
b11118
hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling ( #29282 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49404271 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…
6 -
llama.cpp releases dev-tools 5d ago
b11117
HIP : optimize IQ2/IQ3 ( __vsub4 __vcmpne4 ) using SWAR ( #27962 ) HIP : use bit manipulation for __vcmpne4 HIP : use bit manipulation for __vsub4 Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49401157 macOS/iOS: macOS Apple Silicon…
11 -
llama.cpp releases dev-tools 5d ago
b11115
opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin ( #29056 ) opencl: add A8 Q4_K non-MoE dp4a binary kernel opencl: rename binary kernel selection helpers Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49361383…
11 -
llama.cpp releases dev-tools 5d ago
b11114
server: fix router eviction races with the existing queue ( #29217 ) server: route every model load through the queue A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the…
20 -
llama.cpp releases dev-tools 5d ago
b11113
server: do not pass log file to children ( #29212 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49344105 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
34 -
llama.cpp releases dev-tools 5d ago
b11111
vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) ( #24406 ) vulkan : Intel FA kernel optimization for split k path vulkan : Host code update for Intel split k FA kernel path selection, fix A770 Linux op test failures vulkan : use symmetric…
18 -
llama.cpp releases dev-tools 5d ago
b11110
mtmd: add various sanity checks ( #29276 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49308522 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
10 -
llama.cpp releases dev-tools 5d ago
b11109
metal : gate mul_mm_id src1 rescale behind ggml_prec ( #29029 ) metal : gate mul_mm_id src1 rescale behind ggml_prec Assisted-by: Claude Fable 5.1 ggml-webgpu: reject MUL_MAT_ID when src1 precision is F32 cuda/vulkan: reject MUL_MAT_ID in supports_op when src1 prec is F32 fix…
10 -
llama.cpp releases dev-tools 5d ago
b11108
ggml : IQ1_M build prefix sums once per block ( #28706 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49294443 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
4 -
llama.cpp releases dev-tools 5d ago
b11105
jinja: use const for statement::execute and ::visit ( #29271 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49279983 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
29 -
llama.cpp releases dev-tools 5d ago
b11104
server: Add support for binding to multiple addresses ( #28690 ) Add support for binding llama-server to multiple addresses Assisted-by: Codex remove redundant thread handler make it clear about overlapping addr reject --port 0 with multiple tcp addr improve arg handler nits fix…
32 -
llama.cpp releases dev-tools 5d ago
b11103
spec : support DFlash for HunyuanOCR ( #28890 ) model : add DFlash layer-input taps for HunyuanVL DFlash speculative decoding needs the target graph to expose the residual stream entering each layer (res->t_layer_inp[il]) - the draft model reads those tensors to build its…
4 -
llama.cpp releases dev-tools 5d ago
b11102
convert: add MiMo-V2.6 support ( #29257 ) convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert Update conversion/mimo.py fix: use autoparser Co-authored-by: Sigbjørn Skjæret…
5 -
llama.cpp releases dev-tools 5d ago
b11101
cmake : allow repeated find_package calls for llama ( #29228 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49213556 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
38 -
llama.cpp releases dev-tools 6d ago
b11099
ci : publish snapdragon builds in release workflow ( #29007 ) The snapdragon CI builds packages only to feed the QDC device tests, so Hexagon NPU binaries never reached the releases page. Build both targets in release.yml and attach them as release assets. Website:…
15