Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native script and Romanized forms. We present INDICTALK, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multiturn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that INDICTALK produces fluent, coherent, and naturally codemixed conversations across both script variants. We will release INDICTALK to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.</p>\n","updatedAt":"2026-07-28T04:05:29.659Z","author":{"_id":"66e1425c919f283fbd7dfb5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png","fullname":"Rajvee Sheth","name":"RajveeSheth","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8730567097663879},"editors":["RajveeSheth"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.23242","authors":[{"_id":"6a68170d73f69d5af2bec5f4","user":{"_id":"692ad60788d89514d82ef5cd","avatarUrl":"/avatars/5029f6b0e77e1135593311e4365c564b.svg","isPro":false,"fullname":"Sahil Gawande","user":"Prolexsahil","type":"user","name":"Prolexsahil"},"name":"Sahil Deepak Gawande","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.574Z","hidden":false},{"_id":"6a68170d73f69d5af2bec5f5","name":"Mayank Singh","hidden":false}],"publishedAt":"2026-07-25T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages","submittedOnDailyBy":{"_id":"66e1425c919f283fbd7dfb5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png","isPro":false,"fullname":"Rajvee Sheth","user":"RajveeSheth","type":"user","name":"RajveeSheth"},"summary":"Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .","upvotes":2,"discussionId":"6a68170d73f69d5af2bec5f6","organization":{"_id":"667eb54cc9fc5e32c079544d","name":"LingoIITGN","fullname":"Lingo Research Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/667b8f8ba271fc5a8e6929de/8xB-4Az0x50XC4PS3ZIbL.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"692ad60788d89514d82ef5cd","avatarUrl":"/avatars/5029f6b0e77e1135593311e4365c564b.svg","isPro":false,"fullname":"Sahil Gawande","user":"Prolexsahil","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"667eb54cc9fc5e32c079544d","name":"LingoIITGN","fullname":"Lingo Research Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/667b8f8ba271fc5a8e6929de/8xB-4Az0x50XC4PS3ZIbL.jpeg"},"query":{}}">
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
Abstract
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .
Community
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native script and Romanized forms. We present INDICTALK, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multiturn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that INDICTALK produces fluent, coherent, and naturally codemixed conversations across both script variants. We will release INDICTALK to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.23242 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.23242 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.23242 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.