llama.cpp releases · · 2 min read

b10413

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

common : auto-detect spec type from draft GGUF metadata (#26814)

  • common : auto-detect spec type from draft GGUF metadata

When -md loads a local draft model without --spec-type, the sidecar
inference in common_models_handler_apply only checks HF repo sidecars
and misses local files. The draft model loads into VRAM but speculative
decoding never activates (types stays NONE).

Read general.architecture from the draft GGUF header and map:
dflash + markov_w1.weight tensor -> draft-dspark
dflash without markov head -> draft-dflash

Assisted-by: opencode

  • common : address review feedback on spec-type auto-detect PR
  • Fix comment spacing to match surrounding style (/* .x = / not /.x =*/)
  • Add LOG_INF when auto-detection fires so users can see why spec decoding enabled
  • Document single-file assumption for split-GGUF edge case

Addresses bot review feedback on #26814.

  • common : move spec-type GGUF auto-detect into speculative module
  • add common_speculative_types_from_gguf() in speculative.cpp/.h
  • use gguf_context_ptr (RAII) from ggml-cpp.h
  • reduce comments to a single line per AGENTS.md style

Addresses review feedback on #26814

  • common : add doc note and join SPC_INF line in spec-type auto-detect

Assisted-by: opencode

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases