llama.cpp releases
500 articles archived · Visit source ↗ · RSS
-
llama.cpp releases dev-tools 6d ago
b11097
opencl: add A8 Q4_0 non-MoE dp4a binary kernel ( #29055 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49141307 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
4 -
llama.cpp releases dev-tools 6d ago
b11096
ci : update Level Zero SDK to v1.33.1 and enable the L0/oneDNN CMake …
11 -
llama.cpp releases dev-tools 6d ago
b11095
hexagon: new HMX-optimized GATED_DELTA_NET ( #29199 ) hex-gdn: start putting together HMX support for GDN hex-gdn: working hmx but not-pipelined and slow for now hex-gdn: re-write vtcm layout handling and prep for pipelining hex-gdn: starting to pipeline hmx and dmas hex-gdn:…
11 -
llama.cpp releases dev-tools 6d ago
b11094
vendor : update cpp-httplib to 0.57.1 ( #29239 ) Signed-off-by: Adrien Gallouët [email protected] Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49092442 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…
17 -
llama.cpp releases dev-tools 6d ago
b11093
metal : fix mask bounds in flash attention block pre-pass ( #29220 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49087452 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
15 -
llama.cpp releases dev-tools 6d ago
b11090
cuda: fix sm_70 tile compilation error ( #29224 ) The 5-argument load_ldmatrix added in 1884824 only defines tile<16,8>, so the Volta tile<8,4> does not match. See #29222 for details. Building on 1884824 , generalize the tile shape of the 5-argument load_ldmatrix from <16,8> to…
14 -
llama.cpp releases dev-tools 6d ago
b11081
test-llama-archs : make tensor data stdev configurable and improve help ( #29133 ) test-llama-archs : make tensor data stdev configurable Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp test-llama-archs : expand usage and add examples Assisted-by:…
19 -
llama.cpp releases dev-tools 6d ago
b11080
tests/test-backend-ops : allow regex entries in the -o filter ( #29204 ) tests/test-backend-ops : allow regex entries in the -o filter so far -o only accepted a comma separated list of exact op names or full test case strings. entries that are not plain op names are now treated…
31 -
llama.cpp releases dev-tools 6d ago
b11078
args: add env vars for temperature, top-p, min-p and penalties ( #27380 ) Allow configuring --temp, --top-p, --min-p, --repeat-penalty, --presence-penalty and --frequency-penalty via LLAMA_ARG_* so llama-server can be fully controlled from an EnvironmentFile (e.g. systemd on…
18 -
llama.cpp releases dev-tools 6d ago
b11077
server : do not forward --api-key-file to router-spawned child instances ( #28938 ) In router mode, authentication belongs to the router. unset_reserved_args() already unset LLAMA_API_KEY, but did not unset LLAMA_ARG_API_KEY_FILE. When --api-key-file was passed, children…
4 -
llama.cpp releases dev-tools 6d ago
b11076
tests : remove stale comment ( #29140 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48998612 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
33 -
llama.cpp releases dev-tools 6d ago
b11074
json: Fixed json enum handling ( #28518 ) Fixed json enum handling Added common_json_value handling for enum values. Added tests/test-json.cpp to cover testing of some aspects of common_json. Removed tests as requested. Applied recommended style and simplification Simplified by…
17 -
llama.cpp releases dev-tools 6d ago
b11071
ci : Upgrade CUDA to 13.4 for Ubuntu CUDA Release Builds ( #29202 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48951751 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
35 -
llama.cpp releases dev-tools 7d ago
b11070
hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements ( #29197 ) hex-dma64: enable support extended buffer mappings and 64bit dma hex-dma64: expand binary ops to support more DMA scenarios hex-dma64: add binary-ops.h hex-dma64: add --hex-dma64 to…
8 -
llama.cpp releases dev-tools 7d ago
b11069
cuda : tune MMVQ to MMQ crossover for SM70 (Volta) ( #28912 ) tune MMVQ to MMQ crossover for SM70 (Volta) Signed-off-by: Yangyu Chen [email protected] Apply suggestion from @JohannesGaessler Apply suggestion from @JohannesGaessler Apply suggestion from @JohannesGaessler…
8 -
llama.cpp releases dev-tools 7d ago
b11068
metal : fix deprecation warnings from macOS 27 SDK ( #29136 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48890258 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
29 -
llama.cpp releases dev-tools 7d ago
b11067
webgpu : add fused gdn + cpy ( #28976 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48873164 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
20 -
llama.cpp releases dev-tools 7d ago
b11065
CUDA: tune FA for Gemma 4 on Ampere or newer ( #29152 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48802880 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
28 -
llama.cpp releases dev-tools 7d ago
b11064
metal : support arbitrary hc in dsv4_hc_pre ( #29169 ) the dsv4_hc_pre kernels hardcoded hc = 4 via a constexpr used with simd_shuffle, so the op was rejected by supports_op for any other hc and fell back to CPU. Kimi-K3 uses dsv4_hc_pre with hc equal to the number of banked…
17 -
llama.cpp releases dev-tools 8d ago
b11062
CUDA: enable sparse fa for qwen4 ( #28770 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48739244 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
35 -
llama.cpp releases dev-tools 8d ago
b11059
metal: add F16 input to the FWHT ( #29094 ) metal: add F16 input to the FWHT The Metal FWHT kernel accepts F32 input only. This change makes the source type a template parameter, so the kernel reads an F16 source directly instead of requiring a converted copy. The F32…
11 -
llama.cpp releases dev-tools 8d ago
b11057
chat : add dedicated Ling 3.0 (Bailing V3) parser ( #28682 ) chat: add dedicated Ling 3.0 (Bailing V3) parser Ling 3.0 Flash templates pre-open the think block in the generation prompt, so the model never emits an opening , and a tool call can arrive before any . The generated…
4 -
llama.cpp releases dev-tools 8d ago
b11056
hexagon: enable I32 GET_ROWS ( #29116 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48662388 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
29 -
llama.cpp releases dev-tools 8d ago
b11055
hexagon: add support for GEGLU_QUICK ( #29114 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48660160 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
14 -
llama.cpp releases dev-tools 8d ago
b11054
hexagon: enable support for TOP_K op ( #29113 ) hexagon: enable support for TOP_K op hex-topk: thread single-row TOP_K, raise VTCM-based size cap hex-topk: fix TOP_K mdev row partitioning hex-topk: optimize TOP_K large-row selection hexagon: clean up comment formatting hex-docs:…
24 -
llama.cpp releases dev-tools 8d ago
b11053
server : improve startup log messages ( #29125 ) server-models : show source per model in log Show [source] tag (preset/models_dir/cache) per model instead of cryptic * marker Show HF hub cache path in the 'Loaded cached model presets' log Add hf_cache::get_cache_dir() public…
14 -
llama.cpp releases dev-tools 8d ago
b11052
json-schema : accept escaped hyphen in regex patterns ( #29127 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48635641 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)…
14 -
llama.cpp releases dev-tools 9d ago
b11050
metal : fix FA support checks ( #29122 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48626162 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
26 -
llama.cpp releases dev-tools 9d ago
b11049
test-llama-archs : generate dummy test vocab ( #29084 ) Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48623939 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,…
37 -
llama.cpp releases dev-tools 9d ago
b11048
metal : support qwen4exp hc ops ( #29000 ) Add support for the new DSV4 HC op variants used by qwen4exp: hc_pre with per-element sigmoid gate (gated variant) hc_post with identity mixing (comb == nullptr) Assisted-by: pi:llama.cpp/Qwen3.8-27B Website: https://llama.app…
18 -
llama.cpp releases dev-tools 9d ago
b11047
cuda : fix CUB argsort corruption caused by in-place keys ( #28389 ) argsort_f32_i32_cuda_cub called the one-shot DeviceRadixSort::SortPairs API with d_keys_in == d_keys_out (temp_keys, temp_keys). CUB's internal double-buffer ping-pong requires distinct key buffers: with…
22 -
llama.cpp releases dev-tools 9d ago
b11046
opencl: add support for bin kernel flash_attn_f32_f16_bin ( #29046 ) opencl: add flash_attn_f32_f16_bin opencl: guarded prefill fa Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48577094 macOS/iOS: macOS Apple Silicon (arm64) macOS…
28 -
llama.cpp releases dev-tools 9d ago
b11045
hexagon: add ROLL op support ( #29105 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48565469 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux:…
18 -
llama.cpp releases dev-tools 9d ago
b11044
hexagon: im2col update ( #29103 ) ggml-hexagon: accept 1D and padded IM2COL ops ggml-hexagon: make pure-DDR IM2COL kernel is_2D-aware ggml-hexagon: extend IM2COL DMA patch-embed fast path to 1D ggml-hexagon: add blocked-staging general IM2COL DMA kernel Website:…
9 -
llama.cpp releases dev-tools 9d ago
b11043
hexagon: HMX flash-attention head_dim padding (support DK=DV=72) ( #26539 ) Allow HMX flash-attention to run with head_dim not a multiple of 64 (e.g. SigLIP head_dim=72), by operating on DK/DV rounded up to 64 with zero-filled tail lanes. Website: https://llama.app Attestations:…
29 -
llama.cpp releases dev-tools 9d ago
b11042
opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin ( #28678 ) opencl: add A8 Q6_K non-MoE binary kernel opencl: fix layout compatibility Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48517412 macOS/iOS: macOS…
14 -
llama.cpp releases dev-tools 9d ago
b11040
ggml : check for allocation failures to prevent crashes ( #28149 ) ggml : check for allocation failures to prevent crashes wording Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48473967 macOS/iOS: macOS Apple Silicon (arm64) macOS…
17 -
llama.cpp releases dev-tools 9d ago
b11039
Model-Saver: Write the SWA pattern, 15 more architectures roundtrip ( #29042 ) llama: read the SWA pattern as a period or a per-layer array Add llama_model_base::load_swa_pattern(), which reads sliding_window_pattern either as one flag per layer or as a period expanded by…
7 -
llama.cpp releases dev-tools 9d ago
b11037
ggml-webgpu: fix supports_op condition for GET_ROWS ( #28978 ) fix get_rows vec4 handling Add src strides checking to vec4_aligned of get_rows and the new test case. Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48438627 macOS/iOS:…
31 -
llama.cpp releases dev-tools 10d ago
b11036
ggml : handle graph buffer reservation failure ( #26070 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48431150 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
20 -
llama.cpp releases dev-tools 10d ago
b11034
vocab : add ufakzeka pre-tokenizer ( #29033 ) vocab : add ufakzeka pre-tokenizer vocab : move ufakzeka to the models list and regenerate the hash mapping Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48421591 macOS/iOS: macOS Apple…
6 -
llama.cpp releases dev-tools 10d ago
b11030
ci : bump android-actions/setup-android to 4.0.4 ( #29065 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/48395907 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
37 -
llama.cpp releases dev-tools 10d ago
b11029
vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 1024 experts ( #28501 ) vulkan: raise the hoisted row-id limit for mul_mat_id to 512 experts The expert-count shader (count_experts.comp) sizes its shared arrays with BLOCK_SIZE, which is 256. Because of that,…
16