Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic structure. We introduce Syntax-informed Positional Embeddings (SiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.</p>\n","updatedAt":"2026-08-11T23:54:28.079Z","author":{"_id":"6463739b51fa6e63060491b1","avatarUrl":"/avatars/f8ee64b3768ef46e847cfa15fcd689f0.svg","fullname":"Haris Riaz","name":"hriaz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8943895697593689},"editors":["hriaz"],"editorAvatarUrls":["/avatars/f8ee64b3768ef46e847cfa15fcd689f0.svg"],"reactions":[],"isReport":false}},{"id":"6a7bcebd6a516f9bf5f19cdb","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-12T01:39:09.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings](https://huggingface.co/papers/2608.03994) (2026)\n* [Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models](https://huggingface.co/papers/2607.08186) (2026)\n* [Maglev: Sliding Recurrent Memory](https://huggingface.co/papers/2608.02870) (2026)\n* [Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers](https://huggingface.co/papers/2607.03994) (2026)\n* [Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension](https://huggingface.co/papers/2608.03494) (2026)\n* [Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering](https://huggingface.co/papers/2606.18986) (2026)\n* [Convolution for Large Language Models](https://huggingface.co/papers/2607.18413) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2608.03994\">When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.08186\">Hidden Decoding at Scale: Latent Computation Scaling for Large Language Models</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.02870\">Maglev: Sliding Recurrent Memory</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03994\">Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.03494\">Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.18986\">Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.18413\">Convolution for Large Language Models</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-12T01:39:09.234Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7014821171760559},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06111","authors":[{"_id":"6a757adde1228e04b323831a","name":"Haris Riaz","hidden":false},{"_id":"6a757adde1228e04b323831b","name":"Hyungji Kim","hidden":false},{"_id":"6a757adde1228e04b323831c","name":"Mihai Surdeanu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6463739b51fa6e63060491b1/ADKJsyLq8h2riJ9_9Hsql.png","https://cdn-uploads.huggingface.co/production/uploads/6463739b51fa6e63060491b1/CY6KT8GJJWMlVCHAzdEvt.png","https://cdn-uploads.huggingface.co/production/uploads/6463739b51fa6e63060491b1/2JsIJr9mbbcOb2w6-VX_5.png","https://cdn-uploads.huggingface.co/production/uploads/6463739b51fa6e63060491b1/8M6LNco1GjORHgOaLmXx-.png","https://cdn-uploads.huggingface.co/production/uploads/6463739b51fa6e63060491b1/mVa3G20dPZfLy14l3KlLN.png"],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers","submittedOnDailyBy":{"_id":"6463739b51fa6e63060491b1","avatarUrl":"/avatars/f8ee64b3768ef46e847cfa15fcd689f0.svg","isPro":false,"fullname":"Haris Riaz","user":"hriaz","type":"user","name":"hriaz"},"summary":"Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic structure. We introduce Syntax-informed Positional Embeddings (SiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.","upvotes":1,"discussionId":"6a757adde1228e04b323831d","projectPage":"https://hriaz17.github.io/SiPE/","githubRepo":"https://github.com/hriaz17/SiPE","githubRepoAddedBy":"user","ai_summary":"SiPE integrates a lightweight syntactic prior from dependency parses into positional embeddings across transformer architectures, improving syntactic generalization and language understanding without altering self-attention or increasing inference cost.","ai_keywords":["positional embeddings","Transformers","syntactic structure","SiPE","dependency parses","absolute positional embeddings","relative positional embeddings","rotary positional embeddings","self-attention","SyntaxGym","perplexity","GLUE"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"64c1a137c1d1f89163c1c9da","name":"UArizona","fullname":"University of Arizona","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/SMavW1WgXf_tWtmZLEfEp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"5e6a3d4ea9afd5125d9ec064","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1584020801691-noauth.jpeg","isPro":true,"fullname":"Stefan Schweter","user":"stefan-it","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64c1a137c1d1f89163c1c9da","name":"UArizona","fullname":"University of Arizona","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/SMavW1WgXf_tWtmZLEfEp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06111.md","query":{}}">
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Abstract
SiPE integrates a lightweight syntactic prior from dependency parses into positional embeddings across transformer architectures, improving syntactic generalization and language understanding without altering self-attention or increasing inference cost.
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic structure. We introduce Syntax-informed Positional Embeddings (SiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
Community
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic structure. We introduce Syntax-informed Positional Embeddings (SiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate where and how the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly---while for encoders it is best added directly to the input embeddings, composing with each encoder's native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision---a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.06111 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.06111 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.06111 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.