Sliding-window beats linear attention
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention sinks and no post training. This could be big for memory constrained local LLM inference.
EDIT: Fixed the link to the paper
[link] [comments]
More from r/LocalLLaMA
-
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin
Sep 8
-
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
Sep 8
-
WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster
Sep 8
-
I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic
Sep 7
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.