News / #music Tag Music 370 articles archived under #music · RSS Sign in to follow Hugging Face Daily Papers research 1mo ago UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating Abstract UnityShots is a memory-driven audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. Generated by… 7 arXiv — NLP / Computation & Language research 1mo ago Robustness assessment of large audio language models in multiple-choice evaluation arXiv:2510.04584v2 Announce Type: replace Abstract: Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in… 13 Hugging Face Daily Papers research 1mo ago Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models Abstract Wan-Streamer is a unified, end-to-end multimodal model that enables real-time audio-visual interaction through causal attention mechanisms and integrated processing of visual, audio, and text modalities. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We present… 20 arXiv — NLP / Computation & Language research 1mo ago AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression arXiv:2606.24286v1 Announce Type: new Abstract: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To… 15 arXiv — NLP / Computation & Language research 1mo ago Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams arXiv:2606.24523v1 Announce Type: new Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially… 11 arXiv — NLP / Computation & Language research 1mo ago ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge arXiv:2606.24648v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained… 15 r/LocalLLaMA community 1mo ago GLM 5.2 on Mac Studio Speedup PR Just a heads up for the lucky few 512 gb mac owners: GLM 5.2 is a game changer because prefill speeds stay above 100 t/s at much higher context, and also take less space, so we can run 4 bit quants well above 100k context. See this PR by the oMLX creator:… 5 Hugging Face Daily Papers research 1mo ago Libretto: Giving LLM Agents a Sense of Musical Structure Abstract Libretto provides a structured framework for symbolic music generation and revision using LLM-native grammar and statistical evaluation across musical dimensions. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Generative music systems can now produce impressive audio from… 18 r/LocalLLaMA community 1mo ago CPU-only TTS benchmark: Kokoro 82M vs Supertonic 3 vs Inflect-Nano-v1 (4.6M params), with UTMOS scoring on every sample Ran three open-weight TTS models head to head on CPU. Intel Xeon, 4 cores, 15.6GB RAM, no GPU. Five configs, six text lengths from 12 to 1712 chars, 5 timed reps per cell after warmup, 150 timed runs total. Every audio output scored with UTMOS (utmos22_strong) so quality isn't… 19 Hugging Face Daily Papers research 1mo ago Improving Text-to-Music Generation with Human Preference Rewards Abstract A text-to-music generation system uses reward conditioning, expert iteration, and preference tuning to improve audio quality while maintaining efficiency within a 120M-parameter model framework. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We describe our entry to the… 19 r/LocalLLaMA community 1mo ago EU AI Act requires TEXT from models and providers to be watermarked 2nd August onwards. Everyone here is affected, regardless where you live. Anyone hate the cookie banners ? Those are absolutely nothing in comparison to what is about to come. The AI Act requires lots of things, many people know it requires every AI modified or generated audiofile to be metadata tagged and fingerprint-watermarked from August on (32M$… 9 r/MachineLearning community 1mo ago Recommendations for speech annotation tools [D] I'm looking for human-in-the-loop platforms that allow you to automatically transcribe audio followed by manually fixing the transcriptions and fine tuning the model. Is there a local (not an online service) installable platform for doing this?   submitted by  … 11 r/LocalLLaMA community 1mo ago Qwen code companion on vscode marketplace - thoughts I just came across this extension in vscode few days ago and tried to use with LM studio hosted models and it really is pretty good compared to `continue`, `kilo`, `cline`, `roo` like I felt without much tweaks, gets straight to the point, if any tweaks required u could do… 36 r/LocalLLaMA community 1mo ago Local agent on 4090 - looking for LM Studio settings I have moved on from Ollama to just dink around and instead want to start running a local agent from time to time. With the 24GB of a 4090 (Gigabyte OC edition) that should be quite possible. But no matter what settings I use for context and batching, token generation is slow as… 36 r/LocalLLaMA community 1mo ago Single RTX 3090 (MSI TRio) giving trouble on inference. Hi, I'm having weird issues with my 3090 on inferencerence via lmstudio , it just: unloads the model/ model crashes + nvidia driver resets freezes the pc gives blue/black screen and the computer restarts or straight up restarts everything. I tried running it regularly,… 33 r/LocalLLaMA community 1mo ago Best Harness for Web Searching Looking for opinions on the best software to do web searching resources. What I've tried: LM Studio + plugins Odysseus I think the problem they're both running into is the search engines they're using max out at like, 10 requests per day/hour or something without an api. I don't… 17 Hugging Face Daily Papers research 1mo ago Duration Aware Scheduling for ASR Serving Under Workload Drift Abstract Duration-aware scheduling policies improve ASR serving latency by leveraging audio length as a predictor for processing time, with SJF and HRRN algorithms showing significant median latency reductions while maintaining throughput. Generated by… 26 r/LocalLLaMA community 1mo ago GLM-5.2 can now run locally in llama.cpp and Unsloth Studio. The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Check the graph for the accuracy of each GLM-5.2-GGUF quantization. Full guide:… 35 arXiv — NLP / Computation & Language research 1mo ago ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion arXiv:2606.20179v1 Announce Type: new Abstract: Grapheme-to-phoneme (G2P) conversion for Modern Hebrew is needed for applications like text-to-speech (TTS), but is challenging due to the language's abjad writing system, which leaves vowels largely unwritten, creating substantial… 21 Hugging Face Daily Papers research 1mo ago MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Abstract MaineCoon represents the first real-time audio-visual autoregressive model for social worlds, achieving high frame rates and long-horizon generation through novel training techniques and inference frameworks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct As an increasing… 21 r/LocalLLaMA community 1mo ago I have a M5 Max MacBook Pro with 128gb of ram, what models should I run on it? Yes I know this is a simple question I could just ask Claude or something but I want to see what the community suggests For context it’s a 16in MacBook Pro and i use Hermes agent as a harness connected to LM studio as obviously it’s preferable to be running MLX models especially… 4 arXiv — NLP / Computation & Language research 1mo ago Continuous Audio Thinking for Large Audio Language Models arXiv:2606.18273v1 Announce Type: new Abstract: Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcription to music analysis. However, because LALMs are typically trained to produce text-aligned… 37 arXiv — NLP / Computation & Language research 1mo ago IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages arXiv:2606.19157v1 Announce Type: cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge… 35 arXiv — NLP / Computation & Language research 1mo ago FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs arXiv:2601.13836v2 Announce Type: replace Abstract: Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on… 35 Ars Technica — AI news-outlet 1mo ago The Gemini-powered Google Home Speaker arrives on June 25 for $100 Google's new smart speaker is more about Gemini than audio quality. 27 TechCrunch — AI news-outlet 1mo ago DeepL acquires Mixhalo for live-event audio streaming and translation With this acquisition, DeepL is opening an office in San Francisco to expand its U.S. business. 12 arXiv — NLP / Computation & Language research 1mo ago NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama arXiv:2606.17391v1 Announce Type: new Abstract: Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmark 21 models, spanning classical, fine-tuned,… 12 arXiv — NLP / Computation & Language research 1mo ago ALAS: An Automatic Latent Alignment Score for Audio Language Models arXiv:2505.19937v3 Announce Type: replace Abstract: Large Language Models (LLMs) are extended into Speech-LLMs, and the quality of the audio--text alignment they learn affects most downstream Spoken Language Understanding (SLU) behavior. Yet despite a growth of fusion… 17 r/LocalLLaMA community 1mo ago I didn't know it was possible to compile llamacpp to run cuda + vulkan at the same time.. cmake -B build -G "Visual Studio 17 2022" -A x64 -DCUDAToolkit_ROOT="C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v13.1" -DGGML_CUDA=ON -DGGML_VULKAN=ON -DGGML_FLASH_ATTN=ON -DGGML_BLAS=OFF -DGGML_NATIVE=OFF -DGGML_RPC=ON -DGGML_BACKEND_DL=ON… 31 Hugging Face Daily Papers research 1mo ago MVEB: Massive Video Embedding Benchmark Abstract A large-scale video embedding benchmark evaluates diverse models across multiple video understanding tasks, revealing that different model architectures excel in specific domains and demonstrating the nuanced impact of audio on performance based on dataset… 7 arXiv — Machine Learning research 1mo ago Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models arXiv:2606.15436v1 Announce Type: new Abstract: Respiratory acoustic foundation models (FMs) excel at cough classification, yet their ability to predict continuous health quantities from cough audio remains largely unexplored, despite the clinical value of passive age, BMI, and… 28 arXiv — NLP / Computation & Language research 1mo ago TMASC: Transmasculine Attitude and Speech Corpus arXiv:2606.16351v1 Announce Type: new Abstract: We introduce the Transmasculine Attitudes and Speech Corpus (TMASC), a multimodal corpus of 196 transmasculine individuals, including questionnaire responses and 66 audio recordings. The questionnaire includes items exploring the… 25 Hugging Face Daily Papers research 1mo ago TuneJury: An Open Metric for Improving Music Generation Preference Alignment Abstract A novel open-source pairwise reward model for text-to-music generation that provides calibrated preference scoring and generalizes across multiple downstream applications through a frozen reward mechanism. Generated by Qwen/Qwen2.5-Coder-32B-Instruct We introduce… 5 r/MachineLearning community 1mo ago Embedded/edge ML folks: what actually eats the most time ,getting data, or cleaning/labeling it (time series sensor data, not computer vision/audio)? [D] I'm trying to understand where people doing sensor based ML on microcontrollers (IMU, accelerometer, vibration ,that kind of time-series data) actually lose the most time. When you've built something like this, what was the bottleneck: Getting enough real world data in the first… 6 r/LocalLLaMA community 1mo ago What do you guys think about Unsloth Studio? As a person who has gone through more AI frontend than one goes through socks, I have really appreciated the Unsloth frontend. It's anything I could ever need and it supports Diffusion Gemma! It has easy options to enable tensor parallelism and much more. Have you guys tried it… 33 r/LocalLLaMA community 1mo ago I think we need a /LocalHarnessLLM or something ... LM Studio Hermes Qwen Code Odysseus Open Claw Open Code Claude Code (and then IDEs w/ agentic capabilities) Continue Rider VS Code And a dozen others I'm sure ... Would love a place to discuss these? If not a new subreddit, a new discord section in localllama discord? I've made… 24 arXiv — Machine Learning research 2mo ago Beyond task performance: Decoding bioacoustic embeddings with speech features arXiv:2606.14662v1 Announce Type: new Abstract: Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species… 6 arXiv — NLP / Computation & Language research 2mo ago The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models arXiv:2606.13993v1 Announce Type: new Abstract: A crucial aspect of linguistic capability is the ability to trade off between stored representations and abstract knowledge: one must retrieve learned representations, but also generate novel ones by applying productive rules.… 34 arXiv — NLP / Computation & Language research 2mo ago AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization arXiv:2606.14694v1 Announce Type: new Abstract: Large reasoning models typically follow a read-then-think paradigm: they observe the complete input, reason over a static context, and then produce the answer. Yet many real-world scenarios are inherently dynamic, such as audio and… 4 arXiv — NLP / Computation & Language research 2mo ago Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources arXiv:2606.14141v1 Announce Type: cross Abstract: Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source… 12 arXiv — NLP / Computation & Language research 2mo ago A Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators arXiv:2606.14230v1 Announce Type: cross Abstract: Deepfakes are artificially generated images, audio, or videos that threaten privacy, security, and information integrity. Detecting such content is crucial for countering disinformation, as the latest models generate highly… 19 Simon Willison community 2mo ago OpenAI WebRTC Audio Session, now with document context OpenAI WebRTC Audio Session, now with document context I built the first version of this tool in December 2024 to try out the then-new OpenAI WebRTC API for interacting with their realtime audio models. Last month OpenAI introduced a brand new model to that API called… 9 Hugging Face Daily Papers research 2mo ago PianoKontext: Expressive Performance Rendering from Deadpan Context Abstract PianoKontext generates variable-length piano performances by aligning MIDI scores with audio in latent space using DTW and DiT blocks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Expressive performance rendering (EPR) aims to generate realistic performances constrained… 12 r/LocalLLaMA community 2mo ago Why hasn't any mainstream game integrated LLMs into NPCs yet? tech demos exist but nothing's actually shipped in a real game. Is it a latency problem or are game studios just not interested~   submitted by   /u/Enough-Astronaut9278 [link]   [comments] 29 arXiv — NLP / Computation & Language research 2mo ago Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation arXiv:2606.13322v1 Announce Type: new Abstract: We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is accumulated waiting time; conventional pipelines… 13 arXiv — NLP / Computation & Language research 2mo ago Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data arXiv:2606.13507v1 Announce Type: new Abstract: Large-scale mined corpora provide abundant training data for end-to-end speech-to-speech translation (S2ST) but may contain noise, misalignment, and semantic errors. Filtering noisy data is crucial to maintain robust speech… 30 r/LocalLLaMA community 2mo ago Infinite Music Glitch on my Arduino with Magenta Realtime 2 I built a local voice AI realtime music setup where my ESP32 microcontroller talks to my MacBook over WebSockets. The microcontroller is just a tiny Arduino-based device with a mic and speaker, and the MacBook M4 Pro runs Magenta Realtime 2 locally and streams the audio back to… 38 arXiv — NLP / Computation & Language research 2mo ago Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents arXiv:2606.11219v1 Announce Type: new Abstract: Audio language models (ALMs) are increasingly used for speech-based understanding, yet their ability to perform semantic reasoning beyond transcription, Text-to-Audio Retrieval, Captioning, and Question-Answering accuracy remains… 32 arXiv — NLP / Computation & Language research 2mo ago Pretrained self-supervised speech models can recognize unseen consonants arXiv:2606.11542v1 Announce Type: new Abstract: Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their training data are heavily skewed toward high-resource… 17 r/LocalLLaMA community 2mo ago I wired a fully offline voice loop to Ollama + LM Studio — 100% CPU, no GPU, nothing leaves your machine (Silero VAD + Parakeet STT + Supertonic TTS 3) I kept wanting to talk to my local models instead of typing, but every voice setup wanted a GPU, shipped my audio to the cloud, or was macOS-only. So I built one that's none of those — and I benchmarked it, so these are real measured numbers, not vibes. One command installs the… 12 Page 5 of 8 · 370 articles ← Newer Older →