News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow Hugging Face Daily Papers research 1d ago A Common Measure of Communication for Speech Brain-Computer Interfaces Abstract Open-vocabulary mutual information provides a unified metric to compare speech brain-computer interfaces across different vocabularies and conditions, revealing trade-offs between vocabulary coverage and decoding accuracy. Generated by thinkingmachines/Inkling-Small… 18 r/LocalLLaMA community 2d ago How do you guys handle your personal RAG setup I am getting into developing a RAG setup, for getting information out of existing documents, new document ingestion, web searches, and good visuals. I am planning to use it for, alongside the regular "chat to my data", ingesting personal docs, invoices, creating tables views and… 7 arXiv — Machine Learning research 2d ago The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100 arXiv:2609.03231v1 Announce Type: new Abstract: The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a… 13 arXiv — NLP / Computation & Language research 2d ago Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition arXiv:2609.02901v1 Announce Type: new Abstract: Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically… 24 arXiv — NLP / Computation & Language research 2d ago Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions arXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However,… 11 arXiv — NLP / Computation & Language research 2d ago Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech arXiv:2609.03502v1 Announce Type: new Abstract: In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a… 33 arXiv — NLP / Computation & Language research 2d ago Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis arXiv:2609.03992v1 Announce Type: new Abstract: We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs… 7 arXiv — NLP / Computation & Language research 2d ago SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training arXiv:2609.02941v1 Announce Type: cross Abstract: Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches… 32 arXiv — NLP / Computation & Language research 2d ago VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis arXiv:2609.03203v1 Announce Type: cross Abstract: Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause,… 16 r/LocalLLaMA community 2d ago I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size I have been trying to squeeze TTS stack down far enough to run in a $3 chip which has 512kb of SRAM without NPU. While trying to get to that milestone i built sanoTTS which has - 11 voices, 6 languages - params size ranging from 294k - 2.2m. For comparison we are 1000x smaller… 35 The Information — AI news-outlet 3d ago Anthropic Splits From Google, OpenAI Over State AI Safety Bill Anthropic is at odds with other big tech firms over a Massachusetts Senate proposal requiring big AI developers to hire independent evaluators to assess catastrophic risks posed by their models every four months. The Massachusetts proposal could set a new standard for AI… 19 Hugging Face Daily Papers research 3d ago VibeVoice-ASR-Streaming Technical Report Abstract A streaming, LLM-based end-to-end model unifies speaker-attributed speech recognition and diarization for low-latency real-time applications. Generated by thinkingmachines/Inkling-Small Traditional speaker-attributed ASR systems treated ASR and speaker diarization as… 34 arXiv — Machine Learning research 3d ago A Common Measure of Communication for Speech Brain-Computer Interfaces arXiv:2609.02887v1 Announce Type: new Abstract: Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction.… 16 arXiv — Machine Learning research 3d ago Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models arXiv:2609.01723v1 Announce Type: cross Abstract: Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content.… 36 arXiv — NLP / Computation & Language research 3d ago SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition arXiv:2609.01737v1 Announce Type: new Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a… 26 arXiv — NLP / Computation & Language research 3d ago AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking arXiv:2609.01828v1 Announce Type: new Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem. A strong per-turn text editor… 29 arXiv — NLP / Computation & Language research 3d ago Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language arXiv:2609.02606v1 Announce Type: new Abstract: Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural… 23 arXiv — NLP / Computation & Language research 3d ago Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases arXiv:2609.02735v1 Announce Type: new Abstract: Per-patient adapters are the preferred production architecture for dysarthric automatic speech recognition (ASR), yet parameter-efficient fine-tuning (PEFT) variants have not been compared in the speaker-dependent, per-patient… 7 arXiv — NLP / Computation & Language research 3d ago Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction arXiv:2609.02623v1 Announce Type: cross Abstract: Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given… 6 r/LocalLLaMA community 3d ago Microsoft VibeVoice-ASR-Streaming Released   submitted by   /u/Acceptable-Cycle4645 [link]   [comments] 23 r/LocalLLaMA community 3d ago VoxGen, an AMD-optimized TTS inference engine for VoxCPM 2 models Hi, everyone, I’ve just released VoxGen, a lightweight native inference engine for VoxCPM2, written in Rust and using Vulkan compute instead of Python/PyTorch/CUDA. Why VoxGen? The main reason I started the project was because I needed a decent local text-to-speech solution. I… 10 llama.cpp releases dev-tools 4d ago b10760 mtmd: Fix Qwen3-tts-0.6b ( #28231 ) mtmd: load the qwen3-tts code predictor proj_in as optional The talker and the code predictor share the hidden size on the 0.6B checkpoints, so the reference builds no small_to_mtp_projection and the conversion emits no tensor for it. The… 27 arXiv — NLP / Computation & Language research 4d ago Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts arXiv:2609.00330v1 Announce Type: new Abstract: In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging:… 33 arXiv — NLP / Computation & Language research 4d ago Staged Linguistic Seeding: Grounded Query Expansion for Verified-Unit QA in AI Contact Centers arXiv:2609.00844v1 Announce Type: new Abstract: Customer-service QA in an AI contact center (AICC) runs under deployment constraints that benchmark QA misses: tight voice-hotline latency and a high cost for unsupported or wrong automatic answers. We deploy a system that answers… 27 arXiv — NLP / Computation & Language research 4d ago Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech arXiv:2609.01016v1 Announce Type: new Abstract: Current speech synthesis struggles with code-switching, which mixes a foreign language phrase into a primary language utterance, causing the phrase to be spoken with the primary language's accent rather than its native one. We… 29 arXiv — NLP / Computation & Language research 4d ago Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation arXiv:2609.01246v1 Announce Type: new Abstract: Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we… 20 r/LocalLLaMA community 4d ago Really stunned by the Singularity comment section These are screenshots from the r/Singularity comment section. I'm speechless. This doesn't even have downvotes. How can someone cheer for a monopoly run by a few elites?   submitted by   /u/Howard_banister [link]   [comments] 13 The Information — AI news-outlet 5d ago Meta Unveils New Audio AI Model Meta Platforms on Tuesday unveiled a new audio transcription model, Muse Voice Transcribe, that CEO Mark Zuckerberg said can transcribe speech to text and segment audio based on who is speaking. According to Zuckerberg’s announcement in a Threads post , the model was trained on… 34 r/MachineLearning community 5d ago We released TontaubeV1, a character-level TTS model for long-form generation [P] Hey everyone, My brother and I just released TontaubeV1, a 2.9B-parameter open-weight TTS model focused on expressive speech, long-form generation/narration, and low-latency local inference. It is primarily aimed at English and German and supports zero-shot voice cloning from up… 14 r/LocalLLaMA community 5d ago Vellium v1.1.0 — Live voice, local STT/TTS and easier llama.cpp setup Vellium is an open-source, local-first desktop app for AI chat, character roleplay and long-form writing. Recent updates have focused mostly on making local voice and model setups easier to use. Live mode now supports microphone input, local or Whisper-compatible speech… 15 arXiv — Machine Learning research 5d ago V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational Awareness arXiv:2608.28981v1 Announce Type: new Abstract: As air traffic volumes in the National Airspace System continue to expand, in particular in the low altitude airspaces, the need for scalable decision support tools used by air traffic controllers will also require more… 33 arXiv — NLP / Computation & Language research 5d ago Test-Time Scaling for Scientific Equation Discovery arXiv:2608.28660v1 Announce Type: new Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an… 31 arXiv — NLP / Computation & Language research 5d ago No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus arXiv:2608.28875v1 Announce Type: new Abstract: Supplying context at inference time to a large multimodal model is an inexpensive lever for adapting speech transcription to a domain, and earlier results on smaller models reported large gains. This work tested that mechanism… 25 arXiv — NLP / Computation & Language research 5d ago VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition arXiv:2608.28916v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems are commonly evaluated with word error rate (WER), yet many voice workflows depend on exact written values for identifiers, paths, and measured quantities. A transcript can appear fluent… 31 arXiv — NLP / Computation & Language research 5d ago VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models arXiv:2608.28932v1 Announce Type: new Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed… 16 arXiv — NLP / Computation & Language research 5d ago HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding arXiv:2608.29120v1 Announce Type: new Abstract: Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a… 30 arXiv — NLP / Computation & Language research 5d ago Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages arXiv:2608.29239v1 Announce Type: new Abstract: Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with… 36 arXiv — NLP / Computation & Language research 5d ago When Patients Cut In: Extending Clinical Conversational AI Safety to Interruptions arXiv:2608.29241v1 Announce Type: new Abstract: Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a cascaded architecture (speech-to-text -> LLM -> text-to-speech), so when a patient… 17 TechCrunch — AI news-outlet 6d ago Harvard Law dropout raises $6M for Blue Voice to build a ‘Harvey for police officers’ The seed round for the app that provides real-time legal and policy guidance to officers was led by SignalFire and Las Olas VC. 5 r/LocalLLaMA community 6d ago pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost   submitted by   /u/paf1138 [link]   [comments] 4 arXiv — NLP / Computation & Language research 6d ago Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech arXiv:2608.27462v1 Announce Type: new Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While… 31 arXiv — NLP / Computation & Language research 6d ago Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness arXiv:2608.27988v1 Announce Type: new Abstract: Smooth speaker transitions are fundamental to effective conversation and rely on an interlocutor's ability to predict when to enter the conversation. This ability depends on accurately interpreting and expressing the verbal and… 11 arXiv — NLP / Computation & Language research 6d ago A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls arXiv:2608.28040v1 Announce Type: new Abstract: Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what… 17 arXiv — NLP / Computation & Language research 6d ago Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation arXiv:2608.28508v1 Announce Type: new Abstract: Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale and multilingual analysis. We introduce two corpus-level metrics based on self-supervised (SSL) speech representations for… 6 arXiv — NLP / Computation & Language research 6d ago SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation arXiv:2608.27783v1 Announce Type: cross Abstract: Speech LLMs are usually graded after they answer, although an operating system first has to decide whether a waveform should be sent to the model. We define the Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge)… 33 arXiv — NLP / Computation & Language research 6d ago Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation arXiv:2608.27817v1 Announce Type: cross Abstract: Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence… 33 arXiv — NLP / Computation & Language research 6d ago Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages arXiv:2608.27848v1 Announce Type: cross Abstract: Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST), little… 4 arXiv — NLP / Computation & Language research 6d ago When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI arXiv:2608.28518v1 Announce Type: cross Abstract: We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI… 38 arXiv — NLP / Computation & Language research 6d ago OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion arXiv:2512.00234v3 Announce Type: replace Abstract: There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech… 5 Vercel — AI dev-tools 6d ago Vercel Sandbox now calculates snapshot storage costs daily Vercel now calculates charges for Sandbox snapshot storage from each day’s average usage and adds them to your monthly invoice. Previously, billing used one average across the entire monthly billing period. You can track these daily charges on the team’s Usage page and see how… 22 Page 1 of 10 · 500 articles Older →