r/MachineLearning · · 1 min read

Understanding and Enhancing Kimi Delta Attention [R]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Understanding and Enhancing Kimi Delta Attention [R]

TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1] and the delta rule learning rate is extended to [0, 2] which we call Complex KDA (CKDA). Our theory demonstrates that this form allows us to express any orthogonal diagonal-plus-rank-one matrix and track the S3, S4, and A5 groups, but not S5. Our experiments show that CKDA can learn S3 and S4, shows promising results on Audio continuation and it can train stably and be competitive with standard KDA on language modelling.
Paper title: Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

https://i.redd.it/v7oqopy3v1rh1.gif

submitted by /u/Yossarian_1234
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning