Hugging Face Daily Papers · · 3 min read

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Blog &amp; demo: <a href=\"https://huggingface.co/spaces/wayu-ai/wayu-paxa-tts-blog\">https://huggingface.co/spaces/wayu-ai/wayu-paxa-tts-blog</a><br>Paper: <a href=\"https://arxiv.org/abs/2609.03502\" rel=\"nofollow\">https://arxiv.org/abs/2609.03502</a></p>\n","updatedAt":"2026-09-14T11:41:23.258Z","author":{"_id":"62d192c2d50433c35eb1b48e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/VjmDu8GOIuLuQNBQdQLLS.png","fullname":"Kunat Pipatanakul","name":"kunato","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":20,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3284059464931488},"editors":["kunato"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/VjmDu8GOIuLuQNBQdQLLS.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03502","authors":[{"_id":"6a9a259c8f7c3b7557239413","name":"Kunat Pipatanakul","hidden":false},{"_id":"6a9a259c8f7c3b7557239414","name":"Potsawee Manakul","hidden":false},{"_id":"6a9a259c8f7c3b7557239415","name":"Warit Sirichotedumrong","hidden":false},{"_id":"6a9a259c8f7c3b7557239416","name":"Sittipong Sripaisarnmongkol","hidden":false},{"_id":"6a9a259c8f7c3b7557239417","name":"Pakorn Nathong","hidden":false},{"_id":"6a9a259c8f7c3b7557239418","name":"Phatrasek Jirabovonvisut","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech","submittedOnDailyBy":{"_id":"62d192c2d50433c35eb1b48e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/VjmDu8GOIuLuQNBQdQLLS.png","isPro":true,"fullname":"Kunat Pipatanakul","user":"kunato","type":"user","name":"kunato"},"summary":"In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.","upvotes":1,"discussionId":"6a9a259c8f7c3b7557239419","ai_summary":"A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors.","ai_keywords":["voice-cloning","fixed-voice student","synthetic speech","text preparation","quality filtering","rejection sampling","Thai TTS","lexical tone","code-switching","CER","speaker similarity","prosody pause accuracy"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6a61f2d436983d28935dc359","name":"wayu-ai","fullname":"Wayu Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/zDMksIhbUZd9b8SG1CaGt.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62d192c2d50433c35eb1b48e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/VjmDu8GOIuLuQNBQdQLLS.png","isPro":true,"fullname":"Kunat Pipatanakul","user":"kunato","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a61f2d436983d28935dc359","name":"wayu-ai","fullname":"Wayu Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d192c2d50433c35eb1b48e/zDMksIhbUZd9b8SG1CaGt.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03502.md","query":{}}">
Papers
arxiv:2609.03502

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Published on Sep 3
· Submitted by
Kunat Pipatanakul
on Sep 14
Authors:
,

Abstract

A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors.

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03502
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers