News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow Google DeepMind official-blog 4d ago Gemini 3.8 text-to-speech says hello Gemini 3.8 text-to-speech says hello Sep 23, 2026 | x.com Facebook LinkedIn Mail Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini… 26 r/LocalLLaMA community 4d ago Streaming Nemotron 3 Diarization I’ve been playing with Nemotron 3 Diarization , and it fills a gap I’ve had with local voice agents: keeping track of who is speaking. It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to… 4 arXiv — NLP / Computation & Language research 5d ago "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It arXiv:2609.25021v1 Announce Type: cross Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the… 21 arXiv — Machine Learning research 5d ago A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization arXiv:2609.25471v1 Announce Type: new Abstract: Semi-supervised federated learning (SSFL) trains models on clients' unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly… 36 arXiv — NLP / Computation & Language research 5d ago Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models arXiv:2609.25890v1 Announce Type: new Abstract: Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention,… 6 arXiv — NLP / Computation & Language research 5d ago Enriching Speech Emotion Representations with Conversational Context arXiv:2609.26422v1 Announce Type: new Abstract: Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces.… 15 arXiv — NLP / Computation & Language research 5d ago Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation arXiv:2609.26536v1 Announce Type: new Abstract: In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this,… 24 arXiv — NLP / Computation & Language research 5d ago Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction arXiv:2609.25176v1 Announce Type: cross Abstract: Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think… 9 arXiv — NLP / Computation & Language research 5d ago Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement arXiv:2609.25948v1 Announce Type: cross Abstract: Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained… 25 Vercel — AI dev-tools 5d ago Gemini 3.8 text-to-speech models now available on AI Gateway Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS from Google are now available on AI Gateway . Both models take text and generate speech in more than 100 languages. They support long-form narration, control over delivery, and two-speaker dialogue.… 20 arXiv — NLP / Computation & Language research 6d ago The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker-sensitive cues. We present the Bairong system for the MLC-SLM 2026 Challenge,… 9 arXiv — NLP / Computation & Language research 6d ago Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection arXiv:2609.22261v1 Announce Type: new Abstract: Lexicon-driven misogyny detectors cannot, by construction, distinguish a slur used against a woman from the same slur mentioned in counter-speech ("don't call her that") -- yet exactly this distinction governs whether moderation… 30 arXiv — NLP / Computation & Language research 6d ago Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models arXiv:2609.22452v1 Announce Type: new Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio,… 33 arXiv — NLP / Computation & Language research 6d ago COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural… 7 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B we're so back?!?   submitted by   /u/VoiceApprehensive893 [link]   [comments] 24 MIT Technology Review — AI news-outlet 6d ago She died at the San Diego border. A surveillance camera was in plain sight She had only walked for a couple of hours, and already she was lost.  It was early afternoon on Sept. 14, 2025, when 30-year-old Graciela Gómez Hernández crossed the border from the eastern edge of Tijuana into Southern California, sending voice messages to her mother and… 38 arXiv — Machine Learning research 7d ago Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars arXiv:2609.21109v1 Announce Type: new Abstract: Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models… 13 arXiv — Machine Learning research 7d ago Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding arXiv:2609.21288v1 Announce Type: new Abstract: Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5… 5 arXiv — NLP / Computation & Language research 7d ago Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR arXiv:2609.20828v1 Announce Type: new Abstract: ASR systems optimised for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both critical for language-learning feedback. We present a three-stage pipeline for speakers from… 13 arXiv — NLP / Computation & Language research 7d ago Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge arXiv:2609.20833v1 Announce Type: new Abstract: This paper presents the Transsion Speech Team submission to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational speech. We propose a cascaded framework consisting… 25 arXiv — NLP / Computation & Language research 7d ago Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition arXiv:2609.20839v1 Announce Type: new Abstract: Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing… 23 arXiv — NLP / Computation & Language research 7d ago Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction arXiv:2609.21392v1 Announce Type: new Abstract: Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of… 28 arXiv — NLP / Computation & Language research 7d ago Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER arXiv:2609.21663v1 Announce Type: new Abstract: Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question:… 11 arXiv — NLP / Computation & Language research 7d ago Per-Aetiology Contrastive Severity Embeddings with Phonological Pseudo-Labelling for Multilingual Dysarthric Speech arXiv:2609.21789v1 Announce Type: new Abstract: Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space. We test that pooling assumption with four matched HuBERT-base contrastive… 22 arXiv — NLP / Computation & Language research 7d ago Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts arXiv:2609.21844v1 Announce Type: new Abstract: Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title… 23 arXiv — NLP / Computation & Language research 7d ago NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities arXiv:2609.21967v1 Announce Type: new Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel… 33 arXiv — NLP / Computation & Language research 7d ago Cross-Lingual Parkinson's Disease Severity Assessment Using Pre-trained Speech Embeddings: A Multi-Class Evaluation arXiv:2609.20875v1 Announce Type: cross Abstract: Parkinson's disease (PD) often manifests through speech impairments, facilitating accessible, non-invasive, and cost-effective severity assessment for early diagnosis and progression tracking. Despite advances in speech… 31 arXiv — NLP / Computation & Language research 7d ago Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation arXiv:2609.20995v1 Announce Type: cross Abstract: Natural spoken interaction requires more than streaming ASR, language generation, and speech synthesis: a system must react to overlap without canceling on every acknowledgment, prepare a response before a turn is certain, and… 21 arXiv — NLP / Computation & Language research 7d ago The Hidden Cost of Digits: Number Normalization and WER in ASR Systems arXiv:2609.21084v1 Announce Type: cross Abstract: Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts… 9 arXiv — NLP / Computation & Language research 7d ago Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation arXiv:2609.21683v1 Announce Type: cross Abstract: Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next… 12 arXiv — NLP / Computation & Language research 7d ago Hierarchical attention interpretation: an interpretable speech-level transformer for bi-modal depression detection arXiv:2309.13476v3 Announce Type: replace Abstract: Depression is a common mental disorder. Automatic depression detection tools using speech, enabled by machine learning, help early screening of depression. This paper addresses two limitations that may hinder the clinical… 18 r/MachineLearning community 8d ago Inside sanoTTS — a 294,279-parameter TTS system [P] How sanoTTS works? I have vibe coded this site to show what's inside sanoTTS? Every tensor shown on the page is a real intermediate value captured from the shipped int8 model while it synthesized an actual sentence; no mock-ups, no stand-in data. Just check this out:… 35 r/LocalLLaMA community 8d ago What is the best tts to create audio books? One that support emotions? I tried google and on almost each mode.people complains it’s not good enough? is there any good tts right now that support English and can express emotions for audio and not read it in monotone voice?   submitted by   /u/Alarmed_Wind_4035 [link]   [comments] 21 r/LocalLLaMA community 9d ago Ternary Bonsai 2 27B (1.75bpw) vs. Gemma 26B-A4B MoE Introduction My audiobook pipeline that has to decide who speaks each line of dialogue in a novel, so the TTS can cast voices per character. It's been running on Gemma 4 26B-A4B (QAT Q4). Bonsai 2 27B looked like it should win: a 27B-class model in 5.9 GB means a stronger base… 13 arXiv — NLP / Computation & Language research 10d ago A frontend-backend architecture for tool calls in full-duplex speech models arXiv:2609.19334v1 Announce Type: new Abstract: Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture… 38 arXiv — NLP / Computation & Language research 10d ago Full-Duplex Speech Models Take the Floor When Asked, Not When Needed arXiv:2609.19596v1 Announce Type: new Abstract: Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a… 14 arXiv — NLP / Computation & Language research 10d ago Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data arXiv:2609.19805v1 Announce Type: new Abstract: Grapheme-to-phoneme (G2P) conversion turns raw text into its phonemic form and is an essential part of both text-to-speech (TTS) and automatic speech recognition (ASR) systems. It is required to be fast, stable and context-aware.… 14 arXiv — NLP / Computation & Language research 10d ago Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition arXiv:2609.20081v1 Announce Type: new Abstract: SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We… 29 arXiv — NLP / Computation & Language research 10d ago Design of the IBM Granite 5.0 TurboCTC ASR Model arXiv:2609.20104v1 Announce Type: new Abstract: We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal… 24 arXiv — NLP / Computation & Language research 10d ago Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech arXiv:2609.20223v1 Announce Type: new Abstract: We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which… 34 arXiv — NLP / Computation & Language research 10d ago The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation arXiv:2609.20232v1 Announce Type: new Abstract: We introduce the \textbf{Public Discourse Corpus (PDC)}, the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across… 34 arXiv — NLP / Computation & Language research 10d ago Unifying Models of Intergroup Hostility in Online Discourse arXiv:2609.20808v1 Announce Type: new Abstract: Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational… 33 arXiv — NLP / Computation & Language research 10d ago A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech arXiv:2609.19398v1 Announce Type: cross Abstract: Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by… 6 arXiv — NLP / Computation & Language research 10d ago From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization arXiv:2609.19630v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the… 25 arXiv — NLP / Computation & Language research 10d ago Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain arXiv:2609.20504v1 Announce Type: cross Abstract: FarmerChat is Digital Green's AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet… 26 The Information — AI news-outlet 10d ago Microsoft Execs Voiced Concern About OpenAI’s Data Use, New York Times Claims Microsoft executives expressed concerns in recent years that OpenAI’s use of paywalled articles from the New York Times and other publishers to train its models could pose legal issues, and OpenAI executives discussed how doing so could threaten those publishers’ businesses—but… 25 TechCrunch — AI news-outlet 11d ago Iceland-based Treble raises $18 million for its voice simulation platform Treble's voice simulation platform is used by voice AI model developers, AI wearable, and robotics companies 34 arXiv — NLP / Computation & Language research 11d ago Myovox: Reading Speech from the Muscles of the Face arXiv:2609.17548v1 Announce Type: new Abstract: Myovox, from myo (muscle) and vox (voice), decodes open-vocabulary English text from 31-channel surface electromyography (sEMG) recorded from the muscles of the face during vocalized speech. It takes the single-subject emg2speech… 24 arXiv — NLP / Computation & Language research 11d ago T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition arXiv:2609.18194v1 Announce Type: new Abstract: In Taiwanese Hokkien automatic speech recognition (ASR), prior studies often treat tone sandhi as a major challenge under the assumption that models fail to process implicit phonological variations. However, our experiments on… 36 arXiv — NLP / Computation & Language research 11d ago Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs arXiv:2609.18516v1 Announce Type: new Abstract: While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally… 11 Page 2 of 10 · 500 articles ← Newer Older →