llama.cpp releases · · 3 min read

b11227

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

context : do not re-reserve the scheduler when toggling causal_attn (#28751)

  • context : do not re-reserve the scheduler when toggling causal_attn

llama_context::set_causal_attn() marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive sched_reserve() passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because causal_attn only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: causal_attn is a graph reuse key (llm_graph_params via cparams), so a new graph is built regardless of sched_need_reserve, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

images config H200 before -> after RTX 4090 before -> after
1 -c 8192 -ub 512 134 -> 105 ms (1.27×) 201 -> 119 ms (1.69×)
24 -c 8192 -ub 512 2278 -> 1562 ms (1.46×) 3559 -> 1748 ms (2.04×)
24 -c 32768 -ub 2048 5379 -> 1584 ms (3.40×) 13377 -> 1759 ms (7.61×)

Generated output remains identical before and after.

  • qwen4exp : make the indexer bias shape independent of causal_attn

The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.

  • context : state the causal_attn shape rule in the comment

  • cont : add TODOs


Co-authored-by: Georgi Gerganov [email protected]

Website:

Attestations:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases