llama.cpp releases · · 2 min read

b10431

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

ggml : recurrent state rollback for ggml_ssm_scan (#26623)

  • Initial changes for Recurrent state rollback for nemotron for cpu and cuda

  • Removing CPU RS rollback. Will enable it in subsequent PRs

  • addition of test case

  • Removing assert and calling runtime API to check if op is supported

  • removing extra API and updating the call sites for K

  • replace static cuda detection to runtime fused_op api

  • address review comments and fallback when SSM rollback not supprted

  • Adding changes for supporting RS-rollback in CPU. Also added test-backend-ops for cpu and cuda

  • removing memory manipulation as rs rollback is now supported in CPU

  • removing the static probe which is not needed now

  • correcting the format

  • address review comments

  • enabling test for all the backends, unsupported backends will fallback to CPU

  • Apply suggestions from code review

Co-authored-by: Georgi Gerganov [email protected]

  • choose different graph based on the result of fused_ssm_op is supported or not and also handled memory->n_rs_seq >1 case incase of op is not supported

  • Support K > 1 in ssm_scan for all backends

  • Fix CI Issues


Co-authored-by: Georgi Gerganov [email protected]
Co-authored-by: Gaurav Garg [email protected]

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases