Kimi Delta Attention

Unlock 2D Rotations: Exploring the Power of Complex Kimi Delta Attention

New research illuminates a significant advancement in AI attention mechanisms.

4 min readMachine Learning
Unlock 2D Rotations: Exploring the Power of Complex Kimi Delta Attention
Understanding and Enhancing Kimi Delta Attention [R]

TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1] and the delta rule learning rate is extended to [0, 2] which we call Complex KDA (CKDA). Our theory demonstrates that this form allows us to express any orthogonal diagonal-plus-rank-one matrix and track the S3, S4, and A5 groups, but not S5. Our experiments show that CKDA can learn S3 and S4, shows promising results on Audio continuation and it can train stably and be competitive with standard KDA on language modelling.
Paper title: Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

https://i.redd.it/v7oqopy3v1rh1.gif

submitted by /u/Yossarian_1234
[link] [comments]

Read the original at Machine Learning