News / #music Tag Music 452 articles archived under #music · RSS Sign in to follow arXiv — NLP / Computation & Language research 2d ago Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions arXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However,… 11 arXiv — NLP / Computation & Language research 2d ago Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis arXiv:2609.03992v1 Announce Type: new Abstract: We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs… 7 Hugging Face Daily Papers research 2d ago The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation Abstract Temporal Context Routing improves script-aligned timing of shots and dialogue in joint audio-video generation by mapping structured script timing onto shared video-audio temporal axes. Generated by thinkingmachines/Inkling-Small Joint audio-video generation models have… 20 arXiv — NLP / Computation & Language research 3d ago AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking arXiv:2609.01828v1 Announce Type: new Abstract: Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem. A strong per-turn text editor… 29 arXiv — NLP / Computation & Language research 3d ago SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval arXiv:2609.02343v1 Announce Type: cross Abstract: Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and… 7 llama.cpp releases dev-tools 3d ago b10773 server : accept data: URLs for input_video and input_audio ( #27735 ) server : accept data: URLs for input_video and input_audio input_video and input_audio passed accept_base64_uri=false to handle_media(), so data: URLs got treated as raw base64 strings and failed later with a… 24 Ollama releases dev-tools 3d ago v0.33.3: gemma4: image and audio input support Safetensors gemma4 imports served by the MLX engine now answer image and audio chats. Images run through both vision architectures: the transformer tower (26B, 31B, e-series) and the 12B's encoder-free unified embedder. Audio arrives through the same intake the ollama API… 22 r/MachineLearning community 4d ago Where can I find legally usable datasets for advanced audio chord recognition? [D] I’m researching how to build or fine-tune an audio-to-chord-recognition engine comparable in ambition to Song Master Pro / Auralis Sound Prism. The goal is not basic major/minor chord detection. I need reliable recognition of dense harmonic material: jazz, soul, funk, neo-soul,… 26 r/MachineLearning community 4d ago MIR with AudioMuse-AI-SAE [P] Hi all, I recently read this paper: Julien Guinot, Alain Riou, Elio Quinton, Gyorgy Fazekas. Steering dense music retrieval with open-vocabulary concept discovery. https://arxiv.org/abs/2608.08757 There is multiple model where you can get embedding from Song and Text so that you… 35 r/LocalLLaMA community 4d ago Android Studios native Gemma 4 runs on llama.cpp https://preview.redd.it/6e9xb57a42nh1.png?width=787&format=png&auto=webp&s=ffae7996bbf8ab00498cc62c733e7597dc550f24 I'm not sure how many people care about Android Studio, but I think it's cool that Google uses llama.cpp. My guess is that it is Vulkan and the QAT versions of… 23 arXiv — NLP / Computation & Language research 4d ago Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment arXiv:2609.00055v1 Announce Type: new Abstract: Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with… 33 r/LocalLLaMA community 4d ago GB10 price increases. Seriously what is the best bang for the buck now...Mac Studio? It is crazy how fast prices are increasing. I'm pulling my hair out to keep ahead of this for students. Servers aren't even an option any more.   submitted by   /u/geekender [link]   [comments] 23 The Information — AI news-outlet 5d ago Meta Unveils New Audio AI Model Meta Platforms on Tuesday unveiled a new audio transcription model, Muse Voice Transcribe, that CEO Mark Zuckerberg said can transcribe speech to text and segment audio based on who is speaking. According to Zuckerberg’s announcement in a Threads post , the model was trained on… 34 Hugging Face Daily Papers research 5d ago DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Abstract A compact 7B native joint audio-video generator uses cross-modal attention, progressive joint training, reinforcement learning with multimodal feedback, and an autoregressive 2K refinement pipeline to produce synchronized high-resolution outputs. Generated by… 30 arXiv — NLP / Computation & Language research 5d ago VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models arXiv:2608.28932v1 Announce Type: new Abstract: Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed… 16 arXiv — NLP / Computation & Language research 5d ago Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning arXiv:2608.29278v1 Announce Type: new Abstract: Omni-modal large language models are increasingly evaluated on clean text--vision--audio inputs, where every channel is present, synchronized, and readily interpretable. Such scores are often taken as evidence of robust cross-modal… 21 Hacker News — AI on Front Page community 6d ago Apple caught off guard by AI demand for Mac Mini and Mac Studio Article URL: https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mini-and-studio-demand/ Comments URL: https://news.ycombinator.com/item?id=49508982 Points: 250 # Comments: 280 4 arXiv — NLP / Computation & Language research 6d ago Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict arXiv:2608.27785v1 Announce Type: new Abstract: We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA… 35 arXiv — NLP / Computation & Language research 6d ago Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluation arXiv:2608.27817v1 Announce Type: cross Abstract: Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence… 33 r/LocalLLaMA community 6d ago snkii/Sori-1B: Audio-Grounded LM Trained From Scratch (No Text-Only Pretraining) Sori-1B is a 1B-parameter audio-language model built by a single SNU researcher whose core claim to fame is that its decoder is trained entirely from scratch on audio-paired text — no text-only pretraining, no pretrained-LM initialization — with the idea being that a language… 9 The Information — AI news-outlet 7d ago How Apple Stumbled Into AI Hardware Success With the Mac The hottest products at Apple right now are not the iPhone, the iPad or a buzzy new show on the company’s streaming service. They are two of the lowliest members of Apple’s venerable Mac product line: its boxy Mac mini and Mac Studio computers, which come without monitors,… 26 r/LocalLLaMA community 7d ago When you say, because I can. Limits of X870e As the heading goes, at some point it stopped being about improvements and just whether I can. So check out my abomination. GLM-5.3-Flash at IQ3_XXS gets about 20t/s generation in Unsloth Studio. Now if only I can make my second 2x48GB DDR5 ram kit play nice, but computer just… 8 r/LocalLLaMA community 7d ago Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one TL;DR: Every public low-bit GGUF of this model is secretly ~4.70 bpw. Shim the rows to 256 and it becomes a real 3.07 bpw / 11.77 GiB file that runs 262K context on 16GB. Needs patched llama.cpp — not LM Studio or Ollama. In the AtomicChat HuggingFace repo it says "There is… 4 r/LocalLLaMA community 8d ago Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering Exo labs making some very exciting and interesting claims. The headline is bandwidth scales linearly on Mac Studio clusters with their solution. There is a thread over at localllm subreddit ( https://www.reddit.com/r/LocalLLM/s/qEYLOFaYwc ) where one of their employees speaks… 17 r/LocalLLaMA community 8d ago AtomicChat/Qwen3.8-Flash-Next-GGUF is Really Good specs hardware: M4 Max 128GB Studio inference engine: llama.cpp (qwen4exp branch) judge: claude-opus-4-6 AtomicChat/Qwen3.8-Flash-Next-GGUF Qwen3.8-Flash-Next is a great model I benched in my previous post , but it is very tight, since all n-grams / PLE are loaded along with the… 17 arXiv — NLP / Computation & Language research 9d ago When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue arXiv:2608.27176v1 Announce Type: new Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or… 32 r/LocalLLaMA community 10d ago 5090 now officially cost 5090 I was planning on another 5090, but then I realize... perhaps I am much better off getting an M5 Ultra Mac Studio with 256gb of ram. We are so genuinely cooked.   submitted by   /u/Sadge404 [link]   [comments] 5 r/LocalLLaMA community 10d ago Let’s be real, memory and gpus price will continue go up next year and the year Models will continue improve for open and closed labs, the demand for compute and memory will continue to increase. Expect to pay double or more for ddr 5 and 6 ram and +60% plus for new consumer gpus . Even a 512 gb mac studio will likely be over 22k .   submitted by  … 8 r/LocalLLaMA community 10d ago [audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more audio.cpp 0.7 is out :) This release adds a lot of new audio models and a new way to compare them locally. Audio.cpp is now at 62 model families and 85+ model variants. And it keeps growing! The biggest user-facing change is the new Arena UI . Instead of testing one model at a… 22 Google DeepMind official-blog 10d ago Gemini Omni 1.1 Flash lets you build with more control Gemini Omni 1.1 Flash lets you build with more control Aug 27, 2026 | x.com Facebook LinkedIn Mail Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.… 26 r/LocalLLaMA community 10d ago Qwen3.8-Flash-Next: Time to Update Those Benchmarks specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless,… 35 Hugging Face Daily Papers research 10d ago Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Abstract JoyAI-Echo-1.5 unifies long-form video and interactive world generation through cross-shot memory, geometry-aware camera control, and rollout-aware training to maintain identity and coherence over extended sequences. Generated by thinkingmachines/Inkling-Small Video… 36 Hugging Face Daily Papers research 10d ago Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Abstract A new benchmark evaluates how well multimodal language models follow diverse video-based instructions with visual, audio, and structural constraints. Generated by thinkingmachines/Inkling-Small Multimodal Large Language Models (MLLMs) have shown strong performance in… 14 arXiv — NLP / Computation & Language research 10d ago Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace arXiv:2608.24958v1 Announce Type: cross Abstract: An audio language model is a black box in a specific way: we see what it says, never what it works out on the way there, and chain-of-thought monitoring helps only if the model writes its reasoning down. Reading a base Qwen3-Omni… 8 Stratechery (Ben Thompson) community 11d ago Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia. 30 Hugging Face Daily Papers research 11d ago LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Abstract LAION-BVD is a large-scale open video dataset enabling multimodal pre-training across video, audio, and image modalities with synthetic captions and strong benchmark performance. Generated by thinkingmachines/Inkling-Small We present LAION-BVD, a large-scale open video… 22 Vercel — AI dev-tools 11d ago Gemini 3.5 Transcribe now available on AI Gateway Gemini 3.5 Transcribe from Google is now available on AI Gateway. It takes audio and returns text, in two variants: google/gemini-3.5-transcribe transcribes a complete recording in a single request. google/gemini-3.5-transcribe-live transcribes audio over a WebSocket, returning… 18 r/LocalLLaMA community 11d ago M5 Ultra 96GB vs M5 Max 128GB — is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow? I’ve been going back and forth on this for a week and I can’t settle it, so I’m hoping someone here has hands-on numbers. The two configs (German prices, dealer quote, incl. VAT): Config Price Mac Studio M5 Max, 128GB / 512GB SSD €5,859 Mac Studio M5 Max, 128GB / 1TB SSD €6,189… 32 r/LocalLLaMA community 12d ago Mac Studio M5 Max Cost Analysis At $10k, you could get - 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan) - 5.7B tokens with DeepSeek V4 Pro OpenRouter - 100B tokens with DeepSeek V4 Flash OpenRouter As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait… 26 r/LocalLLaMA community 12d ago Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory   submitted by   /u/themixtergames [link]   [comments] 18 Hacker News — AI on Front Page community 12d ago Apple Introduces New Mac Studio with M5 Max and M5 Ultra Article URL: https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/ Comments URL: https://news.ycombinator.com/item?id=49433316 Points: 277 # Comments: 162 26 r/LocalLLaMA community 12d ago tencent/WeMM-Embedding 9B/4B/2B WeMM-Embedding-9B is a universal multimodal embedding model built on Qwen3.5. It accepts text, images, videos, visual documents, and interleaved multimodal inputs, and returns a 4,096-dimensional L2-normalized embedding. Audio input is not supported.… 10 arXiv — NLP / Computation & Language research 12d ago PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing… 12 arXiv — NLP / Computation & Language research 12d ago WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs arXiv:2608.22704v1 Announce Type: new Abstract: Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no pathway to recover them during decoding. We show… 11 Hugging Face Daily Papers research 12d ago TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Abstract TLive-Omni is an omni-modal model for live-commerce that unifies image, video, audio, and text via timestamped token grouping, staged supervised training, and reinforcement fine-tuning with verifiable feedback to enable accurate real-time understanding. Generated by… 24 Vercel — AI dev-tools 12d ago Wan 3.0 now available on AI Gateway Wan 3.0 from Alibaba is now available on AI Gateway as alibaba/wan-v3.0-video . One model covers text to video, image to video, first and last frame conditioning, and reference-based generation, and it takes image, video, and audio as references. Clips run up to 30 seconds at… 29 r/LocalLLaMA community 13d ago What's the best local model you've found for 8 GB of VRAM? I'm curious what other people are using for local LLM coding / agentic coding with only 8 GB of VRAM . My current setup is: Intel Core i7-11800H RTX 3070 Laptop , 8 GB VRAM 32 GB DDR4 RAM openSUSE Tumbleweed / KDE Unsloth Studio pi.dev as the coding agent After testing quite a… 33 r/LocalLLaMA community 13d ago iPhone Local TTS EPUB Reading - Audiobookify I've been working on this project (been a developer for a few years) for a few months now, and it has been in active testing for ~2 months. Its an EPUB reader that also offers local offline TTS, so its not just TTS focused, its meant to be a good regular reading app as well. I'd… 38 arXiv — Machine Learning research 13d ago AudioWorldSim: Realistic Binaural Audio Datasets For World Models arXiv:2608.21075v1 Announce Type: cross Abstract: This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a custom… 37 arXiv — NLP / Computation & Language research 13d ago Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care arXiv:2608.20346v1 Announce Type: new Abstract: Speech systems used in customer-facing applications often require domain-specific language coverage. We present a synthetic Bengali speech dataset for telecom customer-care scenarios. The dataset contains 10,000 audio-text pairs,… 8 Page 1 of 10 · 452 articles Older →