News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion arXiv:2607.02862v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID… 24 arXiv — NLP / Computation & Language research 1mo ago S-DiverSe: Spanish Diverse Speech arXiv:2607.03207v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech, yet speech affected by neurological conditions remains a challenge. We present S-DiverSe (Spanish Diverse Speech), a corpus of 3.2 hours of in-the-wild… 20 arXiv — NLP / Computation & Language research 1mo ago Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization arXiv:2607.04064v1 Announce Type: new Abstract: Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the… 19 arXiv — NLP / Computation & Language research 1mo ago Towards Digital Preservation of Efik: TTS for a Low-Resource African Language arXiv:2607.04515v1 Announce Type: new Abstract: Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end… 24 arXiv — NLP / Computation & Language research 1mo ago Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition arXiv:2607.04814v1 Announce Type: new Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage linguistic relatedness to enhance… 28 arXiv — NLP / Computation & Language research 1mo ago DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling arXiv:2607.04941v1 Announce Type: new Abstract: Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited for SDLM… 33 arXiv — NLP / Computation & Language research 1mo ago RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain arXiv:2607.05171v1 Announce Type: new Abstract: Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This calls for a foundation model of… 33 arXiv — NLP / Computation & Language research 1mo ago Unified Audio Intelligence Without Regressing on Text Intelligence arXiv:2607.05196v1 Announce Type: new Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a… 7 Hugging Face Daily Papers research 1mo ago Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization Abstract A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeling performance. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Unsupervised syllabic… 8 r/MachineLearning community 1mo ago CPU TTS benchmark with UTMOS MOS scoring: Kokoro, Supertonic, Inflect-Nano, and Kyutai's new Pocket TTS [P] Sharing a CPU TTS benchmark with objective MOS scores in case it's useful for anyone evaluating small TTS models. Adding this because Kyutai's Pocket TTS is architecturally different from the others in the field and I hadn't seen a head-to-head with it yet. Models: Kokoro 82M… 33 r/LocalLLaMA community 1mo ago Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS Kyutai dropped Pocket TTS a bit ago and I've been sitting on it for a benchmark. Finally ran it head to head against the three CPU TTS models that have been getting attention (Kokoro 82M, Supertonic 3, Inflect-Nano-v1). 180 timed runs, 36 audio samples, objective MOS scores via… 8 r/LocalLLaMA community 1mo ago As promised, here is the GitHub link for my 100% local voice-to-voice assistant I've posted up earlier versions of this project before, promising a GitHub link, but never got around to pushing the code from my local system. Sorry all, I have a very busy life :P Anyway, without further ado: GitHub: https://github.com/igorbarshteyn/athena Athena is a fully… 38 r/LocalLLaMA community 1mo ago Gemma Avatar: Talk to Gemma 4-31B face to face This is a voice chat with Gemma 4 31B where you talk to a 3D avatar. It listens while you speak, answers with a voice and a face (the avatar is exposed to the LLM as function tools: set_mood, make_hand_gesture, make_facial_expression) and Gemma decides the expressions on its… 8 arXiv — Machine Learning research 1mo ago Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling arXiv:2607.01830v1 Announce Type: new Abstract: Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria,… 34 arXiv — NLP / Computation & Language research 1mo ago SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings arXiv:2607.01238v1 Announce Type: new Abstract: Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems… 13 arXiv — NLP / Computation & Language research 1mo ago From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages arXiv:2607.01502v1 Announce Type: new Abstract: Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architectures in… 37 arXiv — NLP / Computation & Language research 1mo ago Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data… 20 arXiv — NLP / Computation & Language research 1mo ago NAVER LABS Europe Submission to the Instruction-following 2026 Short Track arXiv:2607.01960v1 Announce Type: new Abstract: In this paper, we describe NAVER LABS Europe's submission to the instruction-following speech processing short track at IWSLT 2026. We participate again in the constrained setting, developing systems capable of jointly performing… 21 arXiv — NLP / Computation & Language research 1mo ago Towards a Phonology-Informed Evaluation of Multilingual TTS arXiv:2607.01965v1 Announce Type: new Abstract: Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We… 13 arXiv — NLP / Computation & Language research 1mo ago Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words arXiv:2607.02002v1 Announce Type: new Abstract: Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (CEs). The present study investigates whether CEs also predict spoken word… 29 arXiv — NLP / Computation & Language research 1mo ago Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning arXiv:2607.02214v1 Announce Type: new Abstract: Instruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large language models (LLMs), as it requires learning a new modality and a wide range of speech-specific instructions in… 12 r/LocalLLaMA community 1mo ago Talking with Gemma 4 31B! Hi! I'm Andi from Hugging Face. This is a fully open-source and free to test/pull/modify demo I'm bringing today. It's a voice demo creating a pipeline of: - Nvidia's parakeet - Gemma 4 31B (served by cerebras!) - My custom inference for Qwen3TTS It sees and searches the web… 13 r/LocalLLaMA community 1mo ago I built a local LLM NPC backend focused on NPC-to-NPC conversations I just released a research project I did last year as open source. It is a fully local speech-to-speech backend for LLM NPCs. So speech-to-text, local LLM, text-to-speech, no cloud needed. The main focus was NPCs talking to each other, not just answering the player, and my study… 17 Smol AI News news-outlet 1mo ago not much happened today **OpenAI** announced **GPT-5.6 Sol**, **Terra**, and **Luna** with strong improvements in coding, math, persistence, and computer use, receiving positive early tester feedback. The launch includes **GPT-Live**, a full-duplex voice architecture enabling simultaneous listening and… 5 arXiv — Machine Learning research 1mo ago Automatic Detection of Stress from Speech in the Trier Social Stress Test arXiv:2607.00986v1 Announce Type: new Abstract: Automatically detecting stress in speech provides an unobtrusive way to gain insights relevant to behavioral research or clinical assessment. This study investigates the automatic differentiation between a stressful and… 12 arXiv — NLP / Computation & Language research 1mo ago Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study arXiv:2607.00143v1 Announce Type: new Abstract: Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech… 6 arXiv — NLP / Computation & Language research 1mo ago Speech Playground: An Interactive Tool for Speech Analysis and Comparison arXiv:2607.00418v1 Announce Type: new Abstract: This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can be cumbersome to integrate them with modern deep learning representations and… 26 arXiv — NLP / Computation & Language research 1mo ago Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents arXiv:2511.07397v3 Announce Type: replace Abstract: Voice agents face a fundamental tension: the reasoning, retrieval, and tool use that make foundation models capable are iterative and slow, while conversational interaction demands responses on a millisecond timescale. Smaller,… 22 r/LocalLLaMA community 1mo ago My reasons to run local models I can finetune any model on any dataset I want. I can use techniques like speculative decoding and other sota approaches to get the max tps The llm provides like anthropic and openai are not getting access to my data The hardware is reusable for vision text speech, and I can run… 10 r/LocalLLaMA community 1mo ago gemma-4-31B on Cerebras is better than ChatGPT voice mode open models will win on inference too 🚀   submitted by   /u/paf1138 [link]   [comments] 34 Hugging Face Daily Papers research 1mo ago FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Abstract Flexible Spoken Language Model (FlexiSLM) introduces dynamic frame rate capabilities for speech input and output, achieving superior performance over fixed-frame-rate models while enabling controllable inference speed. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Spoken… 15 Hugging Face Daily Papers research 1mo ago RedVox: Safety and Fairness Gaps in Speech Models Across Languages Abstract Multilingual safety and fairness benchmark for speech models reveals persistent vulnerabilities across languages and naturalistic conditions. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech-capable models are increasingly deployed in real-world applications across… 36 arXiv — Machine Learning research 1mo ago Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection arXiv:2606.30675v1 Announce Type: cross Abstract: Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic biomarkers remains challenging. We propose a multimodal framework leveraging Whisper for… 28 arXiv — NLP / Computation & Language research 1mo ago Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems arXiv:2606.31055v1 Announce Type: new Abstract: Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because $F_0$, speaking rate, articulation rate, and pausing shift with… 7 arXiv — NLP / Computation & Language research 1mo ago What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR arXiv:2606.31112v1 Announce Type: new Abstract: ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including… 31 arXiv — NLP / Computation & Language research 1mo ago Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection arXiv:2606.31186v1 Announce Type: new Abstract: Spontaneous speech is a vital non-invasive biomarker for Alzheimer's Disease (AD), yet many systems overlook non-linear structural disruptions and clinical heterogeneity in pathological language. We propose a Multi-View Gated Graph… 31 arXiv — NLP / Computation & Language research 1mo ago Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck arXiv:2606.31411v1 Announce Type: new Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is… 4 arXiv — NLP / Computation & Language research 1mo ago Building an ASR Solution for Training and Assessing Children's Reading arXiv:2606.31508v1 Announce Type: new Abstract: Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value for reproducible literacy assessment. We present an open-source system for… 30 arXiv — NLP / Computation & Language research 1mo ago Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition arXiv:2606.31642v1 Announce Type: new Abstract: Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone… 18 arXiv — NLP / Computation & Language research 1mo ago Adapting Foundation ASR Models to Dysarthric Speech: A Case Study arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker,… 11 arXiv — NLP / Computation & Language research 1mo ago LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, we… 25 arXiv — NLP / Computation & Language research 1mo ago ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection arXiv:2606.30646v1 Announce Type: cross Abstract: Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs, providing a non-invasive proxy for cognitive assessment. Yet most speech-based dementia… 18 arXiv — NLP / Computation & Language research 1mo ago UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling arXiv:2606.31128v1 Announce Type: cross Abstract: Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion… 30 r/LocalLLaMA community 1mo ago [audio.cpp] VibeVoice 1.5B released — 90-min podcast in 22.95 min, 4.08x real-time, 2.86x faster than Python without quantization. Native C++/ggml I’m the author of audio.cpp, a C++/ggml runtime for local audio models. I just added VibeVoice 1.5B support and wanted to share the benchmark because long-form multi-speaker TTS is a good stress test for local inference runtimes. Result on RTX 5090: VibeVoice 1.5B Audio length:… 26 Hugging Face official-blog 1mo ago Hugging Face and Cerebras bring Gemma 4 to real-time voice AI Back to Articles a]:hidden"> Hugging Face and Cerebras bring Gemma 4 to real-time voice AI Published July 1, 2026 Update on GitHub Upvote - Amir Mahla A-Mahla Andres Marafioti andito Leandro von Werra lvwerra Saurabh Vyas vyassaurabh cerebras For voice AI, latency is a critical… 37 Hugging Face Daily Papers research 1mo ago One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications Abstract A universal speech enhancement model with configurable algorithmic and computational latency controls using parallel convolutions and early-exit mechanisms. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Different real-time speech applications impose distinct latency… 9 Hugging Face Daily Papers research 1mo ago Interleaved Speech Language Models Latently Work In Text Abstract Interleaved speech-text language models exhibit an implicit transcription phase where text tokens become decodable in intermediate layers, followed by text-based prediction before speech domain transformation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech language… 16 arXiv — NLP / Computation & Language research 1mo ago Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain arXiv:2606.28772v1 Announce Type: new Abstract: Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is not neutral: 42.6% of all annotator disagreement in HateXplain concentrates… 28 arXiv — NLP / Computation & Language research 1mo ago How to Leverage Synthetic Speech for LLM-Based ASR Systems? arXiv:2606.29031v1 Announce Type: new Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic… 15 arXiv — NLP / Computation & Language research 1mo ago Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs arXiv:2606.29534v1 Announce Type: new Abstract: Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a… 23 Page 5 of 10 · 500 articles ← Newer Older →