In this IEEE OJSP paper, we introduce SylReg-LM 7B, an efficiently scalable interleaved syllable-text language model!<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/66815cb077ed01ba88279b1d/7eXmuwWDMlVyJQ0eGZIZR.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/66815cb077ed01ba88279b1d/7eXmuwWDMlVyJQ0eGZIZR.png\" alt=\"results\"></a></p>\n","updatedAt":"2026-07-07T02:57:24.279Z","author":{"_id":"66815cb077ed01ba88279b1d","avatarUrl":"/avatars/be1006b19e5fc892a2b19a7c60774882.svg","fullname":"Ryota Komatsu","name":"ryota-komatsu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.6992353200912476},"editors":["ryota-komatsu"],"editorAvatarUrls":["/avatars/be1006b19e5fc892a2b19a7c60774882.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04064","authors":[{"_id":"6a4c67ca25849b193a834088","user":{"_id":"66815cb077ed01ba88279b1d","avatarUrl":"/avatars/be1006b19e5fc892a2b19a7c60774882.svg","isPro":false,"fullname":"Ryota Komatsu","user":"ryota-komatsu","type":"user","name":"ryota-komatsu"},"name":"Ryota Komatsu","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:12:01.561Z","hidden":false},{"_id":"6a4c67ca25849b193a834089","name":"Kota Kawakita","hidden":false},{"_id":"6a4c67ca25849b193a83408a","name":"Takuma Okamoto","hidden":false},{"_id":"6a4c67ca25849b193a83408b","name":"Takahiro Shinozaki","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66815cb077ed01ba88279b1d/wkO8McdOksVxkOvwcZrzl.png"],"publishedAt":"2026-07-05T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization","submittedOnDailyBy":{"_id":"66815cb077ed01ba88279b1d","avatarUrl":"/avatars/be1006b19e5fc892a2b19a7c60774882.svg","isPro":false,"fullname":"Ryota Komatsu","user":"ryota-komatsu","type":"user","name":"ryota-komatsu"},"summary":"Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.","upvotes":1,"discussionId":"6a4c67ca25849b193a83408c","projectPage":"https://ryota-komatsu.github.io/speaker_disentangled_hubert/","githubRepo":"https://github.com/ryota-komatsu/speaker_disentangled_hubert","githubRepoAddedBy":"user","ai_summary":"A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeling performance.","ai_keywords":["syllabic tokenization","teacher-student distillation","HuBERT","latent speech frame representations","speaker-disentangled","cross-entropy objective","syllable boundary detection","syllabic segment clustering","speech language model","SpiRit-LM"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":46},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69a2c4f1816fc48aea554904","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/aURgGWS_ezoY9zI_g_Tcs.png","isPro":false,"fullname":"李嘉豪","user":"evelyn-jones4","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04064.md","query":{}}">
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Abstract
A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeling performance.
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.
Community
In this IEEE OJSP paper, we introduce SylReg-LM 7B, an efficiently scalable interleaved syllable-text language model!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.04064 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.