News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse arXiv:2607.10846v1 Announce Type: new Abstract: Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce… 38 arXiv — NLP / Computation & Language research 1mo ago Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR arXiv:2607.11163v1 Announce Type: new Abstract: Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue,… 36 arXiv — NLP / Computation & Language research 1mo ago FAD-SA-GRU: Enhancing Hate Speech Detection in Algerian Dialect Through Feature-Augmented Self-Attention GRU Networks arXiv:2607.11279v1 Announce Type: new Abstract: The widespread adoption of social media platforms has transformed online communication by enabling users to exchange information and opinions instantly. However, these platforms have also facilitated the dissemination of abusive… 34 arXiv — NLP / Computation & Language research 1mo ago Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection arXiv:2607.11597v1 Announce Type: new Abstract: The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced… 10 r/LocalLLaMA community 1mo ago Self-hosted voice for any agent/harness of your choice (open-source) For a while now I've been maintaining tts-bench ( https://github.com/5uck1ess/tts-bench ) and a blind voting arena ( https://5uck1ess-tts-arena.hf.space ) where people A/B test open text-to-speech (TTS) models without knowing which is which. One problem I've always wanted to… 31 Hacker News — AI on Front Page community 1mo ago Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor Article URL: https://get-inscribe.com/blog/apple-speech-api-benchmark.html Comments URL: https://news.ycombinator.com/item?id=48894752 Points: 206 # Comments: 104 10 arXiv — NLP / Computation & Language research 1mo ago Phone Segmentation and Recognition through Phonological Activation Mapping arXiv:2607.09020v1 Announce Type: cross Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models… 32 arXiv — NLP / Computation & Language research 1mo ago FreyaTTS Technical Report arXiv:2607.09530v1 Announce Type: new Abstract: We introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synthesis. Freya-TTS is a 183.2M-parameter non-autoregressive conditional flow-matching… 33 arXiv — NLP / Computation & Language research 1mo ago Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR arXiv:2607.09598v1 Announce Type: new Abstract: Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause… 26 arXiv — NLP / Computation & Language research 1mo ago Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation arXiv:2511.17813v3 Announce Type: replace Abstract: LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anonymous… 5 Hugging Face Daily Papers research 1mo ago Phone Segmentation and Recognition through Phonological Activation Mapping Abstract Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to… 16 r/LocalLLaMA community 1mo ago Current state of Voice-To-Voice models Hi, Has there been any improvement to Voice models (like RVC) in the last two years, or has nothing changed? Thanks!   submitted by   /u/Iwishlife [link]   [comments] 36 r/LocalLLaMA community 1mo ago How fast can I get a voice assistant to respond without a GPU? Qwen3-ASR and Kokoro-TTS ONNX on CPU. Been testing out the ONNX models to see how far I can push the CPU to take on ASR and TTS, so the GPU is completely free for running the LLM. The video attached shows me testing latency on a 2022 Macbook M2 and an AMD Ryzen 9 7900. This is just running the regex fast commands,… 17 OpenAI official-blog 1mo ago How Deutsche Telekom is rewiring telecommunications with AI How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice. 37 arXiv — NLP / Computation & Language research 1mo ago A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents arXiv:2607.07985v1 Announce Type: new Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1… 5 arXiv — NLP / Computation & Language research 1mo ago COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation arXiv:2607.08117v1 Announce Type: new Abstract: Contextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA (Contextualized ASR Leveraging Biasing… 7 arXiv — NLP / Computation & Language research 1mo ago Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech arXiv:2607.08208v1 Announce Type: new Abstract: This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted… 12 arXiv — NLP / Computation & Language research 1mo ago Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment arXiv:2607.08256v1 Announce Type: new Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a… 18 arXiv — NLP / Computation & Language research 1mo ago When Synthetic Speech Is All You Have: Better Call GRPO arXiv:2607.08409v1 Announce Type: new Abstract: LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute. Yet synthetic speech… 6 arXiv — NLP / Computation & Language research 1mo ago UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech arXiv:2508.09767v3 Announce Type: replace-cross Abstract: We propose UtterTune, a lightweight method for adapting a multilingual text-to-speech (TTS) system built on a large language model (LLM). It improves control of pronunciation in the target language while preserving… 25 Hugging Face Daily Papers research 1mo ago Vidu S1: A Real-Time Interactive Video Generation Model Abstract Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce Vidu S1, a real-time… 15 TechCrunch — AI news-outlet 1mo ago Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia The Paris-based ElevenLabs competitor, just announced a hefty seed extension round. 11 arXiv — NLP / Computation & Language research 1mo ago Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts arXiv:2607.06611v1 Announce Type: new Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words. Recent solutions rely on audio foundation… 10 arXiv — NLP / Computation & Language research 1mo ago Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs arXiv:2607.06831v1 Announce Type: new Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an… 14 arXiv — NLP / Computation & Language research 1mo ago Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese arXiv:2607.07408v1 Announce Type: new Abstract: Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for… 34 arXiv — NLP / Computation & Language research 1mo ago Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders arXiv:2607.07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.… 23 arXiv — NLP / Computation & Language research 1mo ago Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems arXiv:2512.17648v2 Announce Type: replace Abstract: Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality.… 7 r/LocalLLaMA community 1mo ago [audio.cpp] What Does the Fox Say: 4 ASR models (Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR) in native C++/GGML, init streaming support, and 327s of audio transcribed in 2.17s. I just pushed a new audio.cpp update with streaming support and 4 ASR models: Nemotron 3.5 ASR, Higgs Audio STT, VibeVoice ASR, and Hviske ASR (da only). Overall 1.07x to 2.41x faster than Python. I decided to drop Parakeet-TDT since good implementations already exist, and I… 25 Simon Willison community 1mo ago Introducing GPT‑Live Introducing GPT‑Live OpenAI finally upgraded the model used by ChatGPT voice mode! I've had preview access for a few weeks in the iPhone app, and the new model is very impressive. It also has the ability to spin off harder tasks to GPT-5.5: For questions that require web search,… 13 Hugging Face Daily Papers research 1mo ago VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech Abstract Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and accent cues. Generated by… 21 TechCrunch — AI news-outlet 1mo ago OpenAI releases new voice models for more natural live conversations OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation. 23 arXiv — NLP / Computation & Language research 1mo ago Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition arXiv:2607.05612v1 Announce Type: new Abstract: Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately linear relation in log-log space. Modern end-to-end… 7 arXiv — NLP / Computation & Language research 1mo ago NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task arXiv:2607.05623v1 Announce Type: new Abstract: We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder… 27 arXiv — NLP / Computation & Language research 1mo ago Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments arXiv:2607.05964v1 Announce Type: new Abstract: Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across… 27 arXiv — NLP / Computation & Language research 1mo ago From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition arXiv:2607.06289v1 Announce Type: new Abstract: Dhivehi, the national language of the Maldives, is currently under-resourced for automatic speech recognition (ASR) and other NLP tasks. This study investigates whether cross-lingual transfer learning from Sinhala, a linguistically… 29 arXiv — NLP / Computation & Language research 1mo ago Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs arXiv:2607.06540v1 Announce Type: new Abstract: Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors… 13 arXiv — NLP / Computation & Language research 1mo ago BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech arXiv:2607.06054v1 Announce Type: cross Abstract: Off-the-shelf TTS systems are poorly adapted to Taiwanese Mandarin. Their accent defaults to other Mandarin variants, their tokenizers over-segment common Taiwanese text, and their pronunciation degrades at code-switching… 33 arXiv — NLP / Computation & Language research 1mo ago WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS arXiv:2607.06461v1 Announce Type: cross Abstract: While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In… 18 arXiv — NLP / Computation & Language research 1mo ago Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs arXiv:2601.12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a… 24 Hugging Face Daily Papers research 1mo ago Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Abstract A supervised contrastive alignment framework maps WavLM embeddings from English and Mandarin into a shared clinical space for depression detection, addressing cross-lingual generalization challenges and revealing performance artifacts caused by speaker identity leakage.… 38 Vercel — AI dev-tools 1mo ago Chat SDK adds Dial support Chat SDK now supports Dial with the new vendor-official adapter . Build bots that send and receive SMS, MMS, and iMessage on a real phone number, with bidirectional media and inbound voice-call transcripts. Replies use the standard Chat SDK thread and message APIs, with… 19 OpenAI official-blog 1mo ago Introducing GPT-Live A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice. 30 Hacker News — AI on Front Page community 1mo ago Local, CPU-Friendly, High-Quality TTS (Text-to-Speech) with Kokoro Article URL: https://ariya.io/2026/03/local-cpu-friendly-high-quality-tts-text-to-speech-with-kokoro/ Comments URL: https://news.ycombinator.com/item?id=48821576 Points: 372 # Comments: 75 7 r/LocalLLaMA community 1mo ago Gepard : 0.6B streaming TTS built for real-time dialogue - 20× realtime factor, ~50ms time-to-first-audio, vLLM-native, Apache 2.0 We just open-sourced Gepard 1.0 , a TTS model built for real-time conversation. It’s streaming-first: audio starts the moment text arrives, generated frame by frame instead of waiting for a full sentence. - ~555M params : Qwen3.5 0.8B backbone (14 layers) + Nemo NanoCodec (FSQ,… 36 Hugging Face Daily Papers research 1mo ago Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study Abstract Temporal aggregation methods for speech-based depression detection show inconsistent performance across different backbones and training runs, highlighting the need for robust benchmarking criteria. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Speech-based depression… 6 Hugging Face Daily Papers research 1mo ago Unified Audio Intelligence Without Regressing on Text Intelligence Abstract A unified audio-text large language model is presented that integrates audio and text processing through a shared transformer decoder, achieving superior performance across multiple audio and speech tasks while maintaining strong text reasoning capabilities. Generated… 33 r/LocalLLaMA community 1mo ago Are there any local ASR models that surpass Whisper right now? Hey everyone, I'm currently using faster-whisper(medium/large turbo) for local speech recognition, running it on an 8GB VRAM GPU. It works great, but I was wondering if there are any new open-source/local models that outright beat Whisper at this point? Here is exactly what I'm… 18 arXiv — Machine Learning research 1mo ago GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech arXiv:2607.02633v1 Announce Type: new Abstract: We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling. Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispronounce… 19 arXiv — NLP / Computation & Language research 1mo ago Reinforcement Learning for Data-Efficient Code-Switched ASR arXiv:2607.02757v1 Announce Type: new Abstract: Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable rewards… 6 arXiv — NLP / Computation & Language research 1mo ago LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering arXiv:2607.02763v1 Announce Type: new Abstract: Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM methods in low-resource settings. This paper investigates whether text-to-speech… 20 Page 4 of 10 · 500 articles ← Newer Older →