News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — Machine Learning research 18d ago Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate… 31 arXiv — NLP / Computation & Language research 18d ago MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond arXiv:2607.22100v1 Announce Type: new Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering.… 24 arXiv — NLP / Computation & Language research 18d ago Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine-tuning, requiring large, task-specific speech… 17 r/LocalLLaMA community 18d ago ~20s that you'll never get back There is no quality or value to this post, however, I hope you might find humor in this broken output from my local Qwen TTS setup. Evidently I messed something up.   submitted by   /u/Full_Dimension_3495 [link]   [comments] 34 r/LocalLLaMA community 19d ago ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding… 36 r/LocalLLaMA community 20d ago I released Inflect v2: two ultra-tiny complete TTS models under 4M and 10M parameters I’ve spent the past month trying to find the point where an extremely small TTS model stops feeling like a size experiment and starts feeling genuinely useful. Today I’m releasing Inflect v2 , with two complete local text-to-speech models: Inflect-Nano-v2: 3.96M parameters,… 9 Ars Technica — AI news-outlet 20d ago Canadian legislator reads out apparent LLM response in floor speech "Here’s a more natural, flowing version of that section..." 17 TechCrunch — AI news-outlet 20d ago OpenAI’s new voice mode makes it to the ChatGPT desktop app ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents. 27 arXiv — NLP / Computation & Language research 21d ago Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions? arXiv:2607.20460v1 Announce Type: new Abstract: Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behavior when explicitly instructed. This is critical for real-world deployment, where… 17 arXiv — NLP / Computation & Language research 21d ago DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models… 16 r/LocalLLaMA community 21d ago [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR and two community models OuteTTS TTS and… 7 TechCrunch — AI news-outlet 21d ago Anthropic updates Claude voice mode with more capable models Claude's new voice model will let you reschedule your meeting or draft an email. 30 arXiv — NLP / Computation & Language research 22d ago Abstraction Induces the Brain Alignment of Language and Speech Models arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very little is known about the… 13 arXiv — NLP / Computation & Language research 22d ago Simultaneous Speech-to-Speech Translation Without Aligned Data arXiv:2602.11072v2 Announce Type: replace Abstract: Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencies. Traditional approaches rely on supervised training with word-level aligned… 25 arXiv — NLP / Computation & Language research 22d ago The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation arXiv:2604.26347v2 Announce Type: replace-cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on… 20 Interconnects research 22d ago Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next A podcast with Florian Brand. 21 r/LocalLLaMA community 22d ago We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions We’re open sourcing an alpha release of NeuTTS-2E : an on-device TTS model with 125M active parameters and 7 controllable emotions. The goal was simple: when you select “angry,” “fearful,” or “happy,” the delivery should follow that instruction rather than whatever emotion the… 14 OpenAI official-blog 23d ago Introducing OpenAI Presence Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows. 10 arXiv — NLP / Computation & Language research 23d ago From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin arXiv:2607.18912v1 Announce Type: new Abstract: Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We… 24 arXiv — NLP / Computation & Language research 23d ago Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing arXiv:2607.18934v1 Announce Type: new Abstract: Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of… 22 arXiv — NLP / Computation & Language research 23d ago Constrained CTC Decoding for Efficient Diacritic Restoration arXiv:2607.18946v1 Announce Type: new Abstract: In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently been… 34 arXiv — NLP / Computation & Language research 23d ago Content is What Remains: Invariant Speech Tokenization from Parallel Utterances arXiv:2607.19033v1 Announce Type: new Abstract: Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions… 26 arXiv — NLP / Computation & Language research 23d ago Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results arXiv:2607.19049v1 Announce Type: new Abstract: Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and… 8 arXiv — NLP / Computation & Language research 23d ago A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour arXiv:2607.18317v1 Announce Type: cross Abstract: We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text as input… 32 arXiv — NLP / Computation & Language research 23d ago Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer arXiv:2607.18662v1 Announce Type: cross Abstract: We present a practical recipe for building a compact Hindi text-to-speech (TTS) model by distilling a large flow-matching teacher (IndicF5, 337M-parameter DiT) under a severe data budget (~17.6 hours). Training a small model from… 12 arXiv — Machine Learning research 24d ago Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models arXiv:2607.17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese… 16 arXiv — NLP / Computation & Language research 24d ago AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures arXiv:2607.17237v1 Announce Type: new Abstract: AI_LectureNote is a historical, readability-oriented post-ASR workflow for Korean-English medical lectures. It rewrites speech-to-text output into study transcripts while restoring Latin-script medical terms rather than Korean… 5 arXiv — NLP / Computation & Language research 24d ago When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation arXiv:2607.17766v1 Announce Type: new Abstract: Extra context is valuable for simultaneous speech translation of technical talks, but injecting the entire document context into every streaming segment is often too coarse. Through diagnostic experiments, we find that context… 37 arXiv — NLP / Computation & Language research 24d ago ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions arXiv:2607.17812v1 Announce Type: new Abstract: As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We… 6 Hugging Face Daily Papers research 24d ago FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Abstract Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism. Existing… 10 r/LocalLLaMA community 24d ago Running a 13M ASR conformer on a microcontroller Hello everyone, I wanted to share a recent project of mine, which brings a 13.1 million parameter convolution transformer model to a < $10 microcontroller (more specifically, the ESP32-S3). It's a distilled and quantized version of nvidias small conformer model from huggingface.… 38 r/LocalLLaMA community 25d ago Introducing Scylla's Band, a new TTS model + inference framework with Android sample! Hey all! https://github.com/lowkeytea/scyllasband -> inference code https://huggingface.co/spybyscript/scyllasband -> model, LiteRT, ONNX, and voices https://lowkeytea.github.io/scyllasband/ -> sample audio for the voices, emotions, and languages. The tldr: 10 voices, 7… 13 Smol AI News news-outlet 25d ago not much happened today **US policy debates** are moving toward restricting Chinese open models like **Kimi**, with potential **procurement restrictions** and **Entity List designations**. Technical voices including **@APompliano**, **@ClementDelangue**, and **@mmitchell_ai** warn this could harm… 18 r/LocalLLaMA community 25d ago Good ASR and TTS models? Hey everyone, Something I don't see discussed often here are ASR and TTS models. I've been using Whisper and Kokoro (old models, I know!) with koboldcpp for a while now but wondered if there are now solid replacements available. Know of Qwen3-ASR and Qwen3-TTS, but haven't found… 12 arXiv — Machine Learning research 25d ago SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models arXiv:2607.15697v1 Announce Type: cross Abstract: Backdoor attacks pose a critical threat to neural network models, allowing attackers to implant a backdoor during the training phase by manipulating a small portion of the training data. In security-sensitive applications such as… 34 arXiv — NLP / Computation & Language research 25d ago Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension arXiv:2607.15856v1 Announce Type: new Abstract: Naturalistic language comprehension requires listeners to process both local probabilistic expectations and contextual semantic relations. Surprisal has been widely used to quantify local word unexpectedness, but evidence that it… 4 arXiv — NLP / Computation & Language research 25d ago Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers arXiv:2607.16085v1 Announce Type: new Abstract: Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example… 17 Hacker News — AI on Front Page community 28d ago EEG shows brain can simultaneous encode two speech streams Article URL: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3003876 Comments URL: https://news.ycombinator.com/item?id=48943745 Points: 209 # Comments: 131 5 arXiv — NLP / Computation & Language research 28d ago PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction arXiv:2412.03230v3 Announce Type: replace Abstract: Chinese ASR correction is challenging because errors are often \emph{phonetic} (many characters share similar Pinyin) while the correction model must also obey a \emph{length constraint} under noisy N-best hypotheses. Existing… 18 OpenAI official-blog 29d ago How Cars24 scales conversations and builds faster with OpenAI Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company. 13 r/LocalLLaMA community 29d ago Hermes on Android (Graphene OS) https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gateway running on my laptop as the backend, with Llama.cpp and Qwen 3.6 35b. Paired… 14 arXiv — NLP / Computation & Language research 1mo ago Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification arXiv:2607.11946v1 Announce Type: new Abstract: Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting language revitalisation and digital inclusion. However, extreme data scarcity… 9 arXiv — NLP / Computation & Language research 1mo ago Toward Metaphor-Fluid Conversation Design for Voice User Interfaces arXiv:2502.11554v3 Announce Type: replace-cross Abstract: Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, human-centric metaphors that fail to adapt to diverse contexts and user needs. This… 7 Vercel — AI dev-tools 1mo ago How Speechify serves 500,000 dynamic pages to 60 million users on Vercel Speechify on Vercel 500,000+ pages served across 40+ languages Cut costs 50% by auto-scaling with Fluid compute Zero user impact on bad deploys with Instant Rollbacks Speechify started as a tool for people with dyslexia. Cliff Weitzman, Founder & CEO, built it because reading… 31 r/LocalLLaMA community 1mo ago [audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released audio.cpp again. Hopefully you are not sick of it yet :) Release 0.3 adds five new models: Supertonic 3, MOSS-TTS-Local, MOSS-TTS-Nano, IndexTTS2, and Irodori-TTS. The highlight is Supertonic 3. It can hit 200 ×+ real time on CUDA (RTX 5090), 6×+ on CPU, and around 47 ms TTFT in… 37 Hugging Face official-blog 1mo ago Introducing Real World VoiceEQ: Measuring the human quality of voice AI Back to Articles a]:hidden"> Introducing Real World VoiceEQ: Measuring the human quality of voice AI Published July 15, 2026 Update on GitHub Upvote 6 David Ayllon dayllon HumeAI Alice aliceebaird HumeAI Jeff Brooks jeffbrooks HumeAI Franc Camps Febrer francamps HumeAI Jakub… 25 TechCrunch — AI news-outlet 1mo ago The founder of Hinge raised $18M to build a new AI dating service, Overtone Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions." 26 Hacker News — AI on Front Page community 1mo ago Speech Recognition and TTS in less than 500kb Article URL: https://github.com/moonshine-ai/moonshine/tree/main/micro Comments URL: https://news.ycombinator.com/item?id=48911793 Points: 299 # Comments: 33 21 arXiv — NLP / Computation & Language research 1mo ago Efficiently Adapting Spoken Language Models for the Singaporean Context arXiv:2607.10092v1 Announce Type: new Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual,… 19 arXiv — NLP / Computation & Language research 1mo ago Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR arXiv:2607.10256v1 Announce Type: new Abstract: This paper investigates how language similarity can improve cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource settings. Warlpiri, an Australian Aboriginal language, has very limited transcribed… 5 Page 3 of 10 · 500 articles ← Newer Older →