Hugging Face Daily Papers · · 4 min read

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native script and Romanized forms. We present INDICTALK, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multiturn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that INDICTALK produces fluent, coherent, and naturally codemixed conversations across both script variants. We will release INDICTALK to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.</p>\n","updatedAt":"2026-07-28T04:05:29.659Z","author":{"_id":"66e1425c919f283fbd7dfb5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png","fullname":"Rajvee Sheth","name":"RajveeSheth","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8730567097663879},"editors":["RajveeSheth"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.23242","authors":[{"_id":"6a68170d73f69d5af2bec5f4","user":{"_id":"692ad60788d89514d82ef5cd","avatarUrl":"/avatars/5029f6b0e77e1135593311e4365c564b.svg","isPro":false,"fullname":"Sahil Gawande","user":"Prolexsahil","type":"user","name":"Prolexsahil"},"name":"Sahil Deepak Gawande","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.574Z","hidden":false},{"_id":"6a68170d73f69d5af2bec5f5","name":"Mayank Singh","hidden":false}],"publishedAt":"2026-07-25T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages","submittedOnDailyBy":{"_id":"66e1425c919f283fbd7dfb5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66e1425c919f283fbd7dfb5e/lfX3gDTLKvwX5QgAivq4S.png","isPro":false,"fullname":"Rajvee Sheth","user":"RajveeSheth","type":"user","name":"RajveeSheth"},"summary":"Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .","upvotes":2,"discussionId":"6a68170d73f69d5af2bec5f6","organization":{"_id":"667eb54cc9fc5e32c079544d","name":"LingoIITGN","fullname":"Lingo Research Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/667b8f8ba271fc5a8e6929de/8xB-4Az0x50XC4PS3ZIbL.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"692ad60788d89514d82ef5cd","avatarUrl":"/avatars/5029f6b0e77e1135593311e4365c564b.svg","isPro":false,"fullname":"Sahil Gawande","user":"Prolexsahil","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"667eb54cc9fc5e32c079544d","name":"LingoIITGN","fullname":"Lingo Research Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/667b8f8ba271fc5a8e6929de/8xB-4Az0x50XC4PS3ZIbL.jpeg"},"query":{}}">
Papers
arxiv:2607.23242

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Published on Jul 25
· Submitted by
Rajvee Sheth
on Jul 28

Abstract

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .

Community

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native script and Romanized forms. We present INDICTALK, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multiturn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that INDICTALK produces fluent, coherent, and naturally codemixed conversations across both script variants. We will release INDICTALK to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.23242 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.23242 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.23242 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers