llama.cpp releases · · 2 min read

b10615

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

metal : per-device tuned (Q, NE) for flash-attn vec (#26570)

  • metal : per-device tuned (Q, NE) for flash-attn vec (#25750)

  • rebase Q-generic FA vec body from 01dc936 (#23114)

  • add 53 f16 (Q,NE) flash-attn vec instantiations (vec 80 -> 133)

  • add FA vec (Q,NE) tuning table + dispatch wiring + SMEM cap fallback

  • add FA vec (Q,NE) perf sweep

  • fill tuning result

  • fold family table into a per-family representative SKU

  • refactor tuning result format

  • extend FA vec tuning to quantized KV caches

  • sync fa vec tuner bucketing with runtime, use pointwise tuning regret

  • update tuned table

  • format and cleanup

  • prefix fa_vec tuning procs with ggml_backend_metal_tuning_, drop unused fa_vec_override_active

  • add device id -> token lookup for the offline tuning tool

  • add ggml-metal-tuning skeleton

  • add op-agnostic perf cell + median timing for the tuner

  • add FA-vec graph build + tensor init to the tuner

  • tools : add FA-vec (Q,NE) sweep, compression and table emit

  • cool down and re-measure the dirty window on thermal drift

  • test-backend-ops : replace the FA vec tune mode with a bounded (Q,NE) slice

  • tools : document the Metal tuner, point the table comment at it

  • abort on unknown KV type, single-source fa_vec_legal_ne

  • cleanup

  • honor -o in the FA vec (Q,NE) slice

  • retune FA-vec (Q, NE) under a pointwise no-harm gate

  • cont : add fa-vec tunings for M1 Pro, M2 Ultra, M5 Max


Co-authored-by: Georgi Gerganov [email protected]

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases