News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation arXiv:2609.30416v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally… 30 arXiv — NLP / Computation & Language research 7h ago Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition arXiv:2609.30439v1 Announce Type: new Abstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech… 18 arXiv — NLP / Computation & Language research 7h ago Inquesto Score: A reliability Protocol For Voice Agents arXiv:2609.30514v1 Announce Type: new Abstract: Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto… 22 arXiv — NLP / Computation & Language research 7h ago Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance arXiv:2609.30773v1 Announce Type: new Abstract: Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech… 26 arXiv — NLP / Computation & Language research 7h ago I-Parakeet: Integer-Only Conformer ASR on Mobile NPU arXiv:2609.30846v1 Announce Type: new Abstract: In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard… 13 arXiv — NLP / Computation & Language research 7h ago Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring arXiv:2609.30924v1 Announce Type: new Abstract: Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P)… 26 arXiv — NLP / Computation & Language research 7h ago THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer arXiv:2609.30984v1 Announce Type: new Abstract: Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces,… 8 arXiv — NLP / Computation & Language research 7h ago Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge arXiv:2609.31511v1 Announce Type: new Abstract: We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a… 9 arXiv — NLP / Computation & Language research 7h ago MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos arXiv:2609.31553v1 Announce Type: new Abstract: Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of… 24 arXiv — NLP / Computation & Language research 7h ago Asymmetric Classifier-Free Guidance for Target-Speaker ASR arXiv:2609.30476v1 Announce Type: cross Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech… 7 arXiv — NLP / Computation & Language research 7h ago Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring arXiv:2609.31293v1 Announce Type: cross Abstract: Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or… 29 arXiv — NLP / Computation & Language research 7h ago Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR arXiv:2604.06487v3 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection… 34 r/LocalLLaMA community 2d ago I built a tiny (332MB) CPU-friendly model for document sorting that actually knows when to say "none fits" (BeeNara) Hey r/LocalLLaMA ! I wanted to share a small project I’ve been working on called BeeNara Why I built this: I was looking for a way to automatically sort my local documents (invoices, letters, contracts) into my personal folders. While local LLMs are amazing, I noticed that… 35 r/LocalLLaMA community 2d ago I compared diarization models on 15 clinical conversations: Nemotron 3, Pyannote, Sortformer and VibeVoice I've been working on clinical speaker attribution at Omi and wanted to compare the current diarization models on the same audio. I used 15 mock doctor–patient consultations from PriMock57, about 2.4 hours. Full recordings, automatic speaker counts, without telling the models… 35 ThursdAI news-outlet 3d ago Opus 5.5 is your new workhorse! OpenAI ships GPT 6 Sol and Luna before DevDay and Meta goes all in on Muse! Your friday read is here From CoreWeave: Opus 5.5 is back and 40% cheaper, GPT-6 halves prices, Grok orders Starbucks from your Tesla, and Google clones your voice in 30 seconds 11 arXiv — NLP / Computation & Language research 3d ago An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection arXiv:2609.28703v1 Announce Type: new Abstract: Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary… 37 arXiv — NLP / Computation & Language research 3d ago PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs arXiv:2609.28727v1 Announce Type: new Abstract: Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level… 35 arXiv — NLP / Computation & Language research 3d ago Temporal Taxation Compounds Under Post-Training Compression of Whisper Models arXiv:2609.28739v1 Announce Type: new Abstract: Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which… 31 arXiv — NLP / Computation & Language research 3d ago BanglaKontho: Closing the Long-Form Gap in Bangla Text-to-Speech arXiv:2609.29146v1 Announce Type: new Abstract: Bangla, the seventh most spoken language in the world, remains under-resourced for neural text-to-speech. Public Bangla speech corpora are dominated by short read-prompt utterances collected for speech recognition, leaving… 13 arXiv — NLP / Computation & Language research 3d ago Parts-of-Speech as Emergent Categories in SAE Latent Space arXiv:2609.29362v1 Announce Type: new Abstract: Sparse AutoEncoders (SAEs) offer a promising way to inspect language model representations, but it is still unclear what kind of linguistic structure their latents expose. We use part-of-speech (PoS) categories as a controlled test… 4 arXiv — NLP / Computation & Language research 3d ago BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech arXiv:2609.29371v1 Announce Type: new Abstract: This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining… 8 arXiv — NLP / Computation & Language research 3d ago agentic-ger: terminology recovery in long-form speech using global context arXiv:2609.29428v1 Announce Type: new Abstract: Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the… 25 arXiv — NLP / Computation & Language research 3d ago YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech arXiv:2609.29448v1 Announce Type: new Abstract: We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date,… 30 arXiv — NLP / Computation & Language research 3d ago How To Do Things With Prompts arXiv:2609.29657v1 Announce Type: new Abstract: When users address large language models, they produce directive speech acts whose pragmatic features differ from those of both everyday conversation and traditional human-computer interaction, and these features change as users… 24 arXiv — NLP / Computation & Language research 3d ago LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity arXiv:2609.29672v1 Announce Type: new Abstract: Artificial intelligence helps education most where an essential provision has been rationed by cost. For language learners that provision is a teacher's voice, which binds listening, reading, speaking, and writing into one act.… 8 arXiv — NLP / Computation & Language research 3d ago Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages arXiv:2609.29798v1 Announce Type: new Abstract: This paper presents an end-to-end study of automatic speech recognition (ASR) for adolescent health communication in three Ghanaian languages (Twi, Dagbani, and Ewe). The work proceeds in three connected stages; First, we benchmark… 33 arXiv — NLP / Computation & Language research 3d ago Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition arXiv:2609.29800v1 Announce Type: new Abstract: Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of… 29 arXiv — NLP / Computation & Language research 3d ago VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching arXiv:2609.30005v1 Announce Type: new Abstract: Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these… 23 arXiv — NLP / Computation & Language research 3d ago A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition arXiv:2609.30160v1 Announce Type: new Abstract: Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling,… 19 arXiv — NLP / Computation & Language research 3d ago Do Audio Language Models Hear and Read Distinctive Features Alike? arXiv:2609.30167v1 Announce Type: new Abstract: Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes… 10 arXiv — NLP / Computation & Language research 3d ago Spooftral: Can Voxtral Audio-Language Model Detect Speech Spoofing? arXiv:2609.28713v1 Announce Type: cross Abstract: Self-supervised learning (SSL) countermeasures (CMs) have shown strong performance in recent years. However, they often show degraded performance while facing unseen spoofing attacks and mismatched conditions. This study examines… 10 arXiv — NLP / Computation & Language research 3d ago BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge arXiv:2609.28758v1 Announce Type: cross Abstract: We describe our submission to the Unsupervised Speech in the Wild (UPS) Challenge at Interspeech 2026, a bidirectional Mamba-2 (BiMamba2) encoder trained with masked discrete-unit prediction following the HuBERT-style paradigm.… 14 arXiv — NLP / Computation & Language research 3d ago Learning New Words from Unlabeled Test Data in Automatic Speech Recognition arXiv:2609.28877v1 Announce Type: cross Abstract: New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual… 21 arXiv — NLP / Computation & Language research 3d ago Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS arXiv:2609.28988v1 Announce Type: cross Abstract: We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer… 27 TechCrunch — AI news-outlet 3d ago 20 minutes with the CEO of ElevenLabs, now reportedly valued at $22B ElevenLabs powers the AI voice on the other end of a lot of customer service calls, and its CEO told me this week that businesses should probably tell you that — at least until getting a machine is what everyone expects anyway. 11 OpenAI official-blog 3d ago Ringg’s AI agents resolve up to 65% of customer calls with OpenAI Using GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1. 35 Latent.Space news-outlet 4d ago [AINews] Meta Connect 2026: Muse glasses, voice, video, and Charm Team Zuck is absolutely on fire. 11 arXiv — NLP / Computation & Language research 4d ago NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task arXiv:2609.27086v1 Announce Type: new Abstract: NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks… 23 arXiv — NLP / Computation & Language research 4d ago Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach arXiv:2609.27205v1 Announce Type: new Abstract: Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P)… 17 arXiv — NLP / Computation & Language research 4d ago Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition arXiv:2609.27289v1 Announce Type: new Abstract: Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an… 38 arXiv — NLP / Computation & Language research 4d ago Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond arXiv:2609.27650v1 Announce Type: new Abstract: Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a… 25 arXiv — NLP / Computation & Language research 4d ago Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery arXiv:2609.27980v1 Announce Type: new Abstract: Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo}… 32 arXiv — NLP / Computation & Language research 4d ago A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models arXiv:2609.28060v1 Announce Type: new Abstract: Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose… 34 arXiv — NLP / Computation & Language research 4d ago Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study arXiv:2609.26823v1 Announce Type: cross Abstract: Post-training quantization of speech language models is often summarized with text-output scores and nominal bit widths. Those numbers alone do not establish behavior that depends on information missing from a transcript, or… 38 arXiv — NLP / Computation & Language research 4d ago Quieter Than the Room: Representation Drift and Task Robustness in Speech Encoders arXiv:2609.27195v1 Announce Type: cross Abstract: Non-speech interference can change a speech representation without causing comparable task loss. We test eight frozen encoders on four tasks, adding non-speech sounds throughout recordings, during speech, or in pauses. Under… 16 arXiv — NLP / Computation & Language research 4d ago Psychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models arXiv:2609.27378v1 Announce Type: cross Abstract: End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization… 22 arXiv — NLP / Computation & Language research 4d ago When Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models arXiv:2609.27382v1 Announce Type: cross Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24… 8 Simon Willison community 4d ago Gemini 3.8 TTS Playground Tool: Gemini 3.8 TTS Playground Google released two new Gemini text-to-speech models today - gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts . They come with a library of over 2,000 voices, plus the ability to create a custom voice with "just a 30-second audio sample of your… 16 TechCrunch — AI news-outlet 4d ago ChatGPT mobile app gets voice-based agentic features Pro and Plus users will be able to use the Work tab on their phones to complete agentic tasks. 14 Hacker News — AI on Front Page community 4d ago Gemini 3.8 text-to-speech Article URL: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ Comments URL: https://news.ycombinator.com/item?id=49817615 Points: 214 # Comments: 111 35 Page 1 of 10 · 500 articles Older →