llama.cpp releases · · 2 min read

b10614

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

metal: per-op source split + parallel compile (#26561)

  • metal : per-op source split + parallel compile (#24021)

  • preliminary extract common header

  • op source split

  • split metallib into 8 libs && load in parallel

  • derive kernel->library routing from functionNames

  • x-macro lib list + underscore filenames, dedup QK_NL, MRC fixes

  • op source split 8 to 20

  • improve robustness of source fallback

  • clean up

  • change bool -> atomic_bool

  • only prepend headers that source actually includes

  • no semaphore, use GCD global queue

  • dedup library compile path, fix NSError lifetime, rename gla

  • relocate upstream concat/rope_back/repeat kernel changes into split files

  • move ggml-common.h from common.h into dequantize.h to shrink binary size


Co-authored-by: lvyichen [email protected]

  • metal: add col2im_1d op (f32/f16/bf16) (#25176)

  • metal : add set_rows with src0 f16 (#25434)

  • metal : add CONV_2D_DW (depthwise convolution) support (#21565)

  • metal : add Q2_0 support (#25419)

  • metal: fuse snake activation (mul, sin, sqr, mul, add) (#25459)

  • ggml-metal: FWHT kernel for metal backend (#25924)

  • metal : port new kernels into the split sources

Move the kernels added on master after the split (lightning indexer,
DSv4 hyper-connections, silu_back, f16 bin ops, TQ2_0, the flash-attn KV
dequantization pass, rope offset/inplace, ssm_scan rollback, packed q8_0
dequantization and the tensor-API mat-mat K clamp) into the corresponding
kernels/*.metal sources. Copied verbatim, no functional change.


Co-authored-by: lvyichen [email protected]
Co-authored-by: Georgi Gerganov [email protected]

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases