News / #voice Tag Voice 500 articles archived under #voice · RSS Sign in to follow arXiv — NLP / Computation & Language research 5h ago Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition arXiv:2608.12327v1 Announce Type: new Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol. We fine-tune six pretrained models (XLSR-53, IndicWav2Vec, MMS-1B, Whisper-Medium,… 34 arXiv — NLP / Computation & Language research 5h ago FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines arXiv:2608.12814v1 Announce Type: new Abstract: FastThaiG2P provides sub-millisecond Thai grapheme-to-phoneme conversion for text-to-speech pipelines (International Phonetic Alphabet and Kokoro-TTS conventions) using a PyThaiNLP-tokenized, extensible dictionary and normalization… 31 arXiv — NLP / Computation & Language research 5h ago CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model arXiv:2608.13101v1 Announce Type: new Abstract: Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and… 36 arXiv — NLP / Computation & Language research 5h ago Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection arXiv:2608.13425v1 Announce Type: new Abstract: Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related… 28 arXiv — NLP / Computation & Language research 1d ago DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition arXiv:2608.11441v1 Announce Type: new Abstract: Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from… 18 arXiv — NLP / Computation & Language research 1d ago Easper: An Accessible ASR Pipeline for Language Documentation arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present… 4 arXiv — NLP / Computation & Language research 1d ago Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low… 18 arXiv — NLP / Computation & Language research 1d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 arXiv — NLP / Computation & Language research 1d ago RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation arXiv:2608.12099v1 Announce Type: cross Abstract: We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a growing key-value cache, Mamba propagates a fixed-size… 21 arXiv — NLP / Computation & Language research 1d ago Marco-Voice Technical Report arXiv:2508.02038v5 Announce Type: replace Abstract: This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in… 13 TechCrunch — AI news-outlet 1d ago Why Stream ring-maker Sandbar says the future of AI wearables is voice AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. Now, a whole wave of… 4 TechCrunch — AI news-outlet 1d ago Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing earbuds all promising to capture your meetings and turn them into summaries and action items. Now, a whole wave of… 19 llama.cpp releases dev-tools 2d ago b10369 mtmd: support pocket-tts ( #26871 ) adapt the api text model ok working impl, need verify and clean up mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample was built as one convolution and one… 16 arXiv — NLP / Computation & Language research 2d ago ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct… 5 arXiv — NLP / Computation & Language research 2d ago Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first… 28 arXiv — NLP / Computation & Language research 2d ago X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction arXiv:2608.10878v1 Announce Type: new Abstract: Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular… 27 arXiv — NLP / Computation & Language research 2d ago myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality… 14 arXiv — NLP / Computation & Language research 2d ago Edge Phoneme Recognition for Children's Speech through Age-Aware Training arXiv:2608.10206v1 Announce Type: cross Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a… 4 arXiv — NLP / Computation & Language research 2d ago DuplexWorld: Can voice agents help you get through the day? arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing… 38 Hugging Face Daily Papers research 2d ago Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Abstract Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator. Generated by thinkingmachines/Inkling-Small Omni-modal dialogue models can understand… 20 r/LocalLLaMA community 2d ago [llama.cpp PR #26608] Ling-3.0 support (unmerged) aetherbird has done some great work getting Ling-3.0 to work in llama.cpp. The architecture is generally identical to deepseekv2. I recently added a microscopic 40 line PR to his that adds support for the Tiny model, works great. Using it for home assistant voice with decent… 23 arXiv — NLP / Computation & Language research 3d ago DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects arXiv:2608.08067v1 Announce Type: new Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the… 14 arXiv — NLP / Computation & Language research 3d ago From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios arXiv:2608.08510v1 Announce Type: new Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still… 33 arXiv — NLP / Computation & Language research 3d ago Multilingual Emotion Neurons in Large Audio-Language Models arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion… 23 arXiv — NLP / Computation & Language research 3d ago Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much… 35 arXiv — NLP / Computation & Language research 3d ago Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based… 35 Hugging Face official-blog 3d ago Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS Back to Articles a]:hidden"> Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS Enterprise + Article Published August 10, 2026 Upvote 4 Maryam Motamedi maryameee nvidia Mikyas Desta mdestanv nvidia Jason Li blisc nvidia… 20 arXiv — Machine Learning research 4d ago The Sparsity Whisperer arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly… 13 arXiv — NLP / Computation & Language research 4d ago Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing… 29 arXiv — NLP / Computation & Language research 4d ago Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a… 4 arXiv — NLP / Computation & Language research 4d ago Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response… 13 r/LocalLLaMA community 5d ago I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS   submitted by   /u/InternationalGap3698 [link]   [comments] 6 r/MachineLearning community 6d ago Real-Time Conversational Agents (RTCA) Workshop @ NeurIPS 2026 — submissions now open, deadline Aug 29 AoE [N] Real-Time Conversational Agents (RTCA) workshop at NeurIPS 2026 (Sydney, Dec 11–12). Submissions are now open on OpenReview. What the workshop is about Conversational AI has crossed into real-time deployment — voice modes, embodied avatars, full-duplex speech agents — but the… 28 r/LocalLLaMA community 6d ago parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM High-performance inference of NVIDIA's Parakeet TDT 0.6B V2 English transcription model, in the browser. Check out the live demo: https://parakeet.narcotic.sh/ A fully custom, dependancy-free implementation with raw WebGPU compute shaders and SIMD WebAssembly audio frontend. 1… 19 llama.cpp releases dev-tools 6d ago b10311 mtmd: stop feeding the text stream again during Qwen3-TTS generation ( #26706 ) The reference implementation has two mutually exclusive prompt layouts. In non streaming mode the prefill carries the whole utterance text plus tts_eos summed with codec_pad, and the trailing text… 29 Hugging Face Daily Papers research 6d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Abstract Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities,… 20 arXiv — NLP / Computation & Language research 7d ago A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper arXiv:2608.05165v1 Announce Type: new Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation… 10 arXiv — NLP / Computation & Language research 7d ago How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs arXiv:2608.05759v1 Announce Type: new Abstract: Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this:… 19 arXiv — NLP / Computation & Language research 7d ago FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India arXiv:2608.06027v1 Announce Type: new Abstract: In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health… 12 arXiv — NLP / Computation & Language research 7d ago Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI arXiv:2608.06141v1 Announce Type: new Abstract: This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and… 31 arXiv — NLP / Computation & Language research 7d ago ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment arXiv:2608.06110v1 Announce Type: cross Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared… 28 arXiv — NLP / Computation & Language research 7d ago Integrating Human Linguistic Insights into AI: Theory-Driven Representation for Multilingual Text-to-Speech arXiv:2204.07228v2 Announce Type: replace Abstract: This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike… 8 r/LocalLLaMA community 7d ago Echo Dot 2 can run 28M LLM at decent speed Code and instructions available here: https://github.com/albertoZurini/echo-dot-2-playground Hello there! After a few days of experimenting I was able to get a completely local voice pipeline running on an Amazon Echo Dot 2. The interesting part for this community is that the… 26 r/LocalLLaMA community 7d ago 🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp 🐦⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_multilingual_357m#run-magpietts-locally-with-nemo-speechcpp I run open… 9 r/MachineLearning community 8d ago What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D] We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI Studio quality speech/audio datasets (high fidelity recordings) Egocentric household activity video datasets (first person daily task recordings) One thing that… 37 arXiv — NLP / Computation & Language research 8d ago MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages arXiv:2608.04433v1 Announce Type: new Abstract: We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based… 29 arXiv — NLP / Computation & Language research 8d ago Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders arXiv:2608.04586v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from… 15 arXiv — NLP / Computation & Language research 8d ago A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy arXiv:2608.04808v1 Announce Type: new Abstract: Part-of-speech tagging for low-resource languages remains challenging due to limited annotated data, especially for linguistically complex languages. Gaidhlig (Scottish Gaelic) is a morphologically rich and endangered language with… 24 arXiv — NLP / Computation & Language research 8d ago Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos arXiv:2608.04939v1 Announce Type: new Abstract: Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context,… 20 arXiv — NLP / Computation & Language research 8d ago GEB-Bench: Abstract Structures Told in Many Voices arXiv:2608.04111v1 Announce Type: cross Abstract: Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the… 21 Page 1 of 10 · 500 articles Older →