News / #music Tag Music 368 articles archived under #music · RSS Sign in to follow arXiv — NLP / Computation & Language research 5h ago EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory arXiv:2608.12627v1 Announce Type: cross Abstract: Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are… 11 r/LocalLLaMA community 11h ago dots-studio/dots3-note-prev · Hugging Face dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and… 27 r/LocalLLaMA community 1d ago Minimax Music 3 open weight release soon? Diffusers has a PR with deets: https://github.com/huggingface/diffusers/pull/14456 Minimax is working on this repository right now and put up a bunch of samples: https://github.com/MiniMax-AI/music3-demo/tree/main/assets/audio/tracks Comfy-Org is teasing about a big release in… 30 arXiv — NLP / Computation & Language research 1d ago Easper: An Accessible ASR Pipeline for Language Documentation arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present… 4 arXiv — NLP / Computation & Language research 1d ago Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning arXiv:2608.11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low… 18 arXiv — NLP / Computation & Language research 1d ago Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This… 8 Hugging Face official-blog 1d ago Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Back to Articles a]:hidden"> Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis Enterprise Article Published August 12, 2026 Upvote 1 Kyle Wiggers Ai2Comms allenai 📄 Tech Report: https://allenai.org/papers/olmoearth | 📊… 9 Hacker News — AI on Front Page community 1d ago Show HN: Woxi - Open-source Mathematica / Wolfram Language reimplementation Woxi is an interpreter for the Wolfram Language written in Rust. It comes with Woxi Studio, a Mathematica-like GUI built with iced, but you can also use Woxi through a CLI, Jupyter kernel, Python package, npm package, or WASM module. Compared with wolframscript / Mathematica,… 28 OpenAI Python SDK releases dev-tools 2d ago v2.54.0 2.54.0 (2026-08-11) Features api: Add new Responses model identifiers ( #3595 ) ( 0652787 ) Bug Fixes api: clarify audio upload metadata requirements ( #3596 ) ( 28888f9 ) Chores api: Update generated-file header attribution to Castiron ( #3583 ) ( ea17fda ) 23 r/LocalLLaMA community 2d ago Introducing Unsloth Desktop app Hi LocalLlama, we're super excited to release Unsloth Desktop today! 🦥 It's the first desktop app that enables you to run and train models locally. Open-source. Available on Mac, Windows, and Linux Supports MLX, diffusion image/video models, audio models, and GGUF You can run… 19 r/LocalLLaMA community 3d ago Why have 8B-12B models been dropped? I am a Macbook Pro M4 user with the 16GB of unified ram. The best model I have been able to run on LM Studio is Gemma4 12B QAT, this model is 66 days old. After that the next best thing LM studio suggests is Nemotron 3 Nano 4B and Qwen3.5 9B, which both are 147-161 days old. It… 37 arXiv — NLP / Computation & Language research 3d ago Multilingual Emotion Neurons in Large Audio-Language Models arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion… 23 arXiv — NLP / Computation & Language research 3d ago REFRAMED: Towards Realistic Audio Description Generation for Movies arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into… 35 arXiv — NLP / Computation & Language research 3d ago Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based… 35 arXiv — NLP / Computation & Language research 3d ago Comparing British and American Audio Description of Movies arXiv:2608.09792v1 Announce Type: new Abstract: Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most… 24 Simon Willison community 3d ago Introducing Muse Glimmer Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). Here's a pelican which I generated using LM Studio's 18.16 GB version of the model : I really… 21 Hacker News — AI on Front Page community 3d ago Humanising LLM Outputs Is Dumb Article URL: https://kuber.studio/blog/Reflections/Humanising-LLM-Outputs-is-Actually-Dumb Comments URL: https://news.ycombinator.com/item?id=49243474 Points: 200 # Comments: 131 5 r/LocalLLaMA community 4d ago Chat UIs with native audio input for multimodal models? I've been running Gemma 4 E4B with oMLX and I can't find any chat interfaces that directly send the audio file to the model instead of running the audio through a separate STT layer. I can confirm the audio layers work because I ran a couple of requests through Pydantic AI in… 6 r/LocalLLaMA community 4d ago MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates… 10 Hugging Face Daily Papers research 4d ago Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Abstract Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final… 18 arXiv — NLP / Computation & Language research 4d ago Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a… 4 arXiv — NLP / Computation & Language research 4d ago Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response… 13 Hugging Face Daily Papers research 4d ago StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Abstract Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design… 27 r/LocalLLaMA community 4d ago The Gemma team will host a special event on August 20 Tweet by u/hackerllama Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest template there are still bugs ), higher precision QAT from the start and… 20 r/LocalLLaMA community 4d ago DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this? Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed a strange behavior during longer agentic coding sessions. Once the context gets… 38 r/LocalLLaMA community 4d ago Open-weight video gen that actually delivers. Five days with MiniMax H3 on local hardware. H3 weights went live on HuggingFace August 3rd and I started pulling them immediately. An omni-modal video model with native stereo audio in the same forward pass, where audio can actually drive the video generation? On open weights? I had to try it. Five days in, the quality is… 31 Simon Willison community 6d ago Quoting John Gruber Me, I try to get into the mindset of playing live music, not recording a studio album. Except when I’m writing a piece where I really want it to be an album. Those aren’t rare , per se, but they’re occasional . If I tried to make every post a hall-of-famer I’d never get anything… 12 Simon Willison community 6d ago Quoting John Gruber Me, I try to get into the mindset of playing live music, not recording a studio album. Except when I’m writing a piece where I really want it to be an album. Those aren’t rare , per se, but they’re occasional . If I tried to make every post a hall-of-famer I’d never get anything… 15 llama.cpp releases dev-tools 6d ago b10326 tts: account for the vocoder pass in the timings line ( #26733 ) get_output runs the waveform work the pipeline defers to it, from a single trailing window to a full pass depending on the model. Measuring it keeps the reported total and the audio to process ratio honest.… 6 r/LocalLLaMA community 6d ago parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM High-performance inference of NVIDIA's Parakeet TDT 0.6B V2 English transcription model, in the browser. Check out the live demo: https://parakeet.narcotic.sh/ A fully custom, dependancy-free implementation with raw WebGPU compute shaders and SIMD WebAssembly audio frontend. 1… 19 Simon Willison community 6d ago The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s… 5 Simon Willison community 6d ago The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s… 4 Hugging Face Daily Papers research 6d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Abstract Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities,… 20 Ars Technica — AI news-outlet 7d ago Suno hopes to go legit with watermarks for AI-generated music Suno plans watermarks and download limits to stop "large-scale abuse." 32 r/LocalLLaMA community 7d ago nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it plainly, the vision and audio towers need a runtime that implements the… 20 TechCrunch — AI news-outlet 7d ago Amid legal battles, Suno says it will start watermarking songs Suno's watermarking feature comes as the company is fighting legal battles on several fronts. 35 r/MachineLearning community 8d ago What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D] We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI Studio quality speech/audio datasets (high fidelity recordings) Egocentric household activity video datasets (first person daily task recordings) One thing that… 37 Hugging Face Daily Papers research 8d ago AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Abstract While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on… 15 arXiv — NLP / Computation & Language research 8d ago Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for… 7 r/LocalLLaMA community 8d ago Scenema Audio Comes to ComfyUI, Runs on 8GB VRAM Hey everyone! Scenema Audio is now a native ComfyUI custom node. Same model that powers scenema.ai now quantized so it fits on 8GB VRAM. When we first released it a few months ago as an API and Docker stack, the full precision transformers were too heavy for most people to… 6 Hugging Face Daily Papers research 8d ago Multi-Task Multi-Frame Visual Piano Transcription Abstract Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT)… 38 r/LocalLLaMA community 9d ago Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected pages to an MP3. Everything runs locally: Kokoro for speech, and an on-device… 5 arXiv — NLP / Computation & Language research 9d ago string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms arXiv:2608.03984v1 Announce Type: new Abstract: We present string2string Studio, an interactive in-browser platform for string-to-string analysis across natural language processing, computational biology, and the digital humanities. The system integrates six main modules… 15 arXiv — NLP / Computation & Language research 9d ago Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning arXiv:2608.02831v1 Announce Type: cross Abstract: Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations:… 14 Hugging Face Daily Papers research 9d ago OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Abstract Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token… 6 Hugging Face Daily Papers research 9d ago AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Abstract Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces… 32 llama.cpp releases dev-tools 9d ago b10274 mtmd: correcting duplicate empty audio chunks for short inputs ( #26536 ) correcting duplicate empty audio chunks for short inputs tests.sh code restored Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED… 30 TechCrunch — AI news-outlet 9d ago Meet Wrinkles, an app that uncovers the hidden stories of the places around you Wrinkles, available on both iOS and Android, essentially acts as an AI-powered audio tour guide that reveals hidden history and local stories. 19 Simon Willison community 9d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 28 Simon Willison community 9d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 26 Page 1 of 8 · 368 articles Older →