Understanding and Enhancing Kimi Delta Attention [R]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1] and the delta rule learning rate is extended to [0, 2] which we call Complex KDA (CKDA). Our theory demonstrates that this form allows us to express any orthogonal diagonal-plus-rank-one matrix and track the S3, S4, and A5 groups, but not S5. Our experiments show that CKDA can learn S3 and S4, shows promising results on Audio continuation and it can train stably and be competitive with standard KDA on language modelling. [link] [comments] |
More from r/MachineLearning
-
Xiaomi releases MiMo-V2.6: "Frontier intelligence, all the modalities, built in public." [N]
Sep 22
-
Paper on ArXiv for a year now, should I disclose about it in ICLR submission? [Discussion]
Sep 22
-
Jev's calibration was measured. The LLMs won [D]
Sep 21
-
I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R]
Sep 21
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.