News / #music Tag Music 368 articles archived under #music · RSS Sign in to follow r/LocalLLaMA community 10d ago Why are Chinese models better* at Frontend than the western top labs? I use A LOT both openAI and Anthropic products. When I need some frontend work (pure web dev) (or answer that feel less verbose and more to the point) I use Anthropic. For multimodality openAI feels better (understanding audio, screenshots, generating images, etc). But openAI… 17 r/LocalLLaMA community 10d ago Time to finally migrate from LM Studio -> llama.cpp, your experience? Has anyone moved from LM Studio to llama.cpp? What was your experience like? What did you have to learn in order to recreate your experience? Which harness/GUI did you switch to? Thanks in advance!   submitted by   /u/CSEliot [link]   [comments] 19 r/LocalLLaMA community 10d ago Is LM Studio abandoning their core product? Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand… 7 arXiv — Machine Learning research 10d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map… 38 arXiv — NLP / Computation & Language research 10d ago Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding arXiv:2608.01560v1 Announce Type: new Abstract: Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding… 29 Hugging Face Daily Papers research 10d ago SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Abstract Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural… 4 Hacker News — AI on Front Page community 10d ago MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video Article URL: https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui Comments URL: https://news.ycombinator.com/item?id=49155629 Points: 202 # Comments: 58 18 arXiv — NLP / Computation & Language research 11d ago TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models arXiv:2607.28896v1 Announce Type: cross Abstract: Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about… 26 r/LocalLLaMA community 11d ago MiniMax-H3 now on huggingface MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks… 16 r/LocalLLaMA community 11d ago DeepSeek V4 @ IQ3XXS on M1 Ultra 128GB- 16 tok/s in LM Studio after patch M1 Ultra 128GB, Unsloth UD-IQ3_XXS, wired limit at 120GB. I was at 5-6 tok/s before the patch. Getting 15-16 tok/s now with the patched engine, and the output seems to have improved. Big thanks to this guy.   submitted by   /u/mil_phickelson [link]   [comments] 9 r/LocalLLaMA community 12d ago DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s. Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was done in there. Not a proper benchmark (used PC in parallel as well) but it gives an idea of the performance from dual 3060 with… 16 r/LocalLLaMA community 12d ago EU AI Act takes effect tomorrow, August 2, 2026. 🤡 Basically you now have to mark all AI generated images, audio, video and text as AI generated. :P   submitted by   /u/xoxaxo [link]   [comments] 32 r/LocalLLaMA community 13d ago Deepseek V4 Flash 0731. LM Studio loading only into RAM. The model refuses to load into VRAM and uses only RAM. What can be an issue? Q2_K_XL from Unsloth if that changes something.   submitted by   /u/esw123 [link]   [comments] 29 r/LocalLLaMA community 13d ago [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox . It is closer to prompt-directed voice acting. DramaBox is built on the LTX-2.3 audio architecture, and prompts can control emotion, delivery, laughs, sighs, pauses, transitions, and speaker behavior. Example… 15 Hugging Face Daily Papers research 13d ago OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Abstract Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different… 27 Hacker News — AI on Front Page community 13d ago Getting 25 Gbps Thunderbolt Ethernet on My Mac Studio Article URL: https://www.jeffgeerling.com/blog/2026/getting-25g-ethernet-mac-thunderbolt/ Comments URL: https://news.ycombinator.com/item?id=49125034 Points: 200 # Comments: 100 24 r/LocalLLaMA community 14d ago Minimax-H3 video model released, open weights coming in the next few days https://x.com/MiniMax_AI/status/2083006198828417501?s=20 Quote from their article: Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound,… 12 arXiv — Machine Learning research 15d ago Journey Operators for Structured Multi-Axis Composition arXiv:2607.26775v1 Announce Type: new Abstract: Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from… 10 arXiv — NLP / Computation & Language research 15d ago Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens arXiv:2607.26350v1 Announce Type: cross Abstract: Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as… 6 arXiv — NLP / Computation & Language research 15d ago Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in… 29 Vercel — AI dev-tools 15d ago Inkling Small from Thinking Machines is now available on AI Gateway Inkling Small from Thinking Machines is now available on AI Gateway. Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task. It is a broad generalist with native reasoning over audio and images,… 30 Hugging Face Daily Papers research 16d ago OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs Abstract Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important… 7 Vercel — AI dev-tools 16d ago Grok Voice Think Fast 2.0 now available on AI Gateway Grok Voice Think Fast 2.0 from xAI is now available on AI Gateway. It is a speech-to-speech voice model that takes audio in and audio out, improving on the previous Grok Voice model in reasoning, transcription accuracy, and conversation. The model reasons in parallel with… 38 TechCrunch — AI news-outlet 16d ago Fish Audio raises $50M seed to build AI voice models for creators and enterprises Since launching last year, the startup today has more than 8 million people using the open-source or hosted version of its models, and now generates annual recurring revenue of $21 million. 4 arXiv — NLP / Computation & Language research 17d ago The JEPA Paradox in Language: The Geometry of Linguistic Alternatives arXiv:2607.23531v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a… 37 arXiv — NLP / Computation & Language research 17d ago Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages arXiv:2607.23808v1 Announce Type: new Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field… 5 arXiv — NLP / Computation & Language research 17d ago Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction arXiv:2607.22729v1 Announce Type: cross Abstract: Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial… 18 Hugging Face Daily Papers research 17d ago JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Abstract Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative… 4 Hugging Face Daily Papers research 17d ago OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Abstract Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging… 12 r/LocalLLaMA community 17d ago Viable ways to run K3 locally just curious how would people run it cheap if they really want kimi k3. dgx spark / strix halo clusters optane persistent memory platform + some gpus mac studio clusters orange pi 6 clusters ssd streaming + gpus multiple ddr3 + connectx 5 rdma clients two dgx stations power 10… 37 arXiv — Machine Learning research 18d ago Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining arXiv:2607.22458v1 Announce Type: new Abstract: Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio models (AST, CLAP, BEATs-bio and BirdNET) with a downstream task none of them saw during training: recovering… 17 arXiv — Machine Learning research 18d ago Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv:2607.21820v1 Announce Type: cross Abstract: Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate… 31 r/LocalLLaMA community 19d ago ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding… 36 r/LocalLLaMA community 19d ago Benchmarks: TensorSharp vs. llama.cpp Cuda and Vulkan Benchmark: TensorSharp vs. llama.cpp I would like to share my latest open source local Unsloth (GGUF) LLM inference engine and applications. It supports many models from Unsloth, like Gemma4, DiffusionGemma, Qwen3.6 with multi-modal (image, vision, audio), Qwen… 38 r/LocalLLaMA community 19d ago My GX10 died Everything ran fine, I was using UD 3.6 Q6 for 35 and 27B, each 4 concurrent requests at 200K context. I had Dify and Mastra to play around with, Unsloth studio to get around to and vLLM ready for whenever I decided to do some more testing. LLama-swap above lama.cpp and liteLLM… 30 r/LocalLLaMA community 20d ago OrangePi AI Studio Pro - Qwen3.5-122B-A10B https://preview.redd.it/wbq8ullnbafh1.png?width=1409&format=png&auto=webp&s=e6d2fe2b1c87c724bc64003c25f917dcee53260f I finally got round to tweaking this, with a bit of help from GLM5.2. The trick to getting it running with vLLM (which I couldn't get anything really out of… 24 r/MachineLearning community 20d ago I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P] Built an open-source AI coding agent that was 7%–75% cheaper than a cold "claude -p" run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: $6.83, 207 turns AutoDev Studio: ~$1.70 for the same bug The full benchmark (including… 14 r/LocalLLaMA community 21d ago FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. Blog Post : https://bfl.ai/blog/flux-3   submitted by   /u/pmttyji [link]   [comments] 23 arXiv — NLP / Computation & Language research 21d ago An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this… 8 r/LocalLLaMA community 21d ago [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR and two community models OuteTTS TTS and… 7 arXiv — NLP / Computation & Language research 22d ago Abstraction Induces the Brain Alignment of Language and Speech Models arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very little is known about the… 13 Hugging Face official-blog 22d ago Bringing Nunchaku 4-bit Diffusion Inference to Diffusers Back to Articles a]:hidden"> Bringing Nunchaku 4-bit Diffusion Inference to Diffusers Published July 23, 2026 Update on GitHub Upvote 6 Pham Hong Vinh rootonchair guest Sayak Paul sayakpaul Large diffusion transformers can create stunning images (or even videos, audio snippets,… 22 arXiv — NLP / Computation & Language research 23d ago Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv:2607.18666v1 Announce Type: new Abstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio… 36 arXiv — NLP / Computation & Language research 23d ago From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin arXiv:2607.18912v1 Announce Type: new Abstract: Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We… 24 Vercel — AI dev-tools 23d ago AI Gateway now supports streaming transcription AI Gateway now supports streaming transcription . Previously, transcription required a complete audio file and returned the full transcript in a single response. Now you can stream audio in as it's captured and get transcript updates back as the model produces them, keeping… 20 TechCrunch — AI news-outlet 23d ago AI and the rise of the universal entertainment app Over the past decade, streaming platforms competed by dominating individual formats like music, video, podcasts, or audiobooks. Now, as AI makes it easier to create, organize, and recommend content, those distinctions are fading, pushing companies like Spotify, Netflix, YouTube,… 17 r/LocalLLaMA community 24d ago Today I learnt the power of LocalLlama DISCLAIMER: No Ai was prompted in the creation of this post. Today I had an experience that complely blew my mind, I just had to write it down. As a bit of background I have been dabbling prompting local models using LM studio for the better part of 18 months now, keeping up to… 38 arXiv — NLP / Computation & Language research 24d ago ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions arXiv:2607.17812v1 Announce Type: new Abstract: As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We… 6 arXiv — NLP / Computation & Language research 24d ago Modeling turn-taking with distant viewing: investigating silence thresholds in human and AI-generated discourse arXiv:2607.18076v1 Announce Type: new Abstract: This study investigates silence gaps in two kinds of audiovisual material. We analysed thirty US situational comedies and fifty-one synthetic podcasts generated with Google NotebookLM. Gaps were compared across speaker gender,… 4 r/LocalLLaMA community 24d ago Trellis.cpp now has a studio! When Trellis.cpp released, people were rightly complaining that while the port was nice, the usability barrier was still high since you had to navigate the command line and fetch all the weights manually. So now, Trellis.cpp has a built-in simple Studio binary: picks the proper… 19 Page 2 of 8 · 368 articles ← Newer Older →