News / #music Tag Music 370 articles archived under #music · RSS Sign in to follow TechCrunch — AI news-outlet 2mo ago Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start Google's Gemini Omni is a new multimodal model that reasons across text, images, audio, and video to generate and edit videos through simple conversation — starting with Omni Flash. 4 r/LocalLLaMA community 2mo ago Floor for local meeting summarization on a 6GB GPU: qwen3.5:0.8b works at 57s, Granite 4 350M hallucinates Disclosure: I made this. Open-source, MIT, Windows + Linux. Not affiliated with voiceflow.com (the chatbot SaaS, name collision, sorry). Why this exists: I wanted local-only dictation and meeting transcription, because audio shouldn't have to leave the machine just to become… 13 r/LocalLLaMA community 2mo ago Audio upscaling, cleanup, or improvement models? I never see this type of model talked about. Are there many open models in the category? I do a lot of audio cleanup and end up using auphonic but would like to be using a local model. Edit: e.g like voice recovery, reverb removal, auto-EQ type stuff   submitted by  … 5 arXiv — NLP / Computation & Language research 2mo ago Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech arXiv:2605.17652v1 Announce Type: new Abstract: There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows… 10 r/MachineLearning community 2mo ago Architecture advice: Real-time pipeline for YouTube Audio -> Whisper -> LLM -> SSE (Sub-10s latency) [D] Hey everyone, I’m building a backend that analyzes long YouTube videos using an LLM. Currently, my flow is a slow waterfall: Download full audio -> Whisper -> LLM -> Return results . For a 30-minute video, the user waits forever. I want to pipeline this for real-time SSE… 5 Hugging Face Daily Papers research 2mo ago AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting Abstract AuralSAM2 integrates audio into SAM2 through an AuralFuser module that generates sparse and dense prompts, enhancing cross-modal influence while maintaining interactive segmentation efficiency. AI-generated summary Segment Anything Model 2 (SAM2) exhibits strong… 18 arXiv — NLP / Computation & Language research 2mo ago Smoothie: Smoothing Diffusion on Token Embeddings for Text Generation arXiv:2505.18853v2 Announce Type: replace Abstract: Diffusion models have achieved state-of-the-art performance in generating images, audio, and video, but their adaptation to text remains challenging due to its discrete nature. Prior approaches either apply Gaussian diffusion… 23 r/LocalLLaMA community 2mo ago Looking to migrate off of Ollama and LMStudio Hello, I'm currently using Ollama / lm studio for things like code inference and proof reading emails, etc. Definitely not experienced in this space but looking to grow. It's been working great but it's a bit slow at times. I use Gemma 4 / Qwen, I also recently tried using… 22 r/LocalLLaMA community 2mo ago GitHub - richardr1126/openreader: An open-source read-along document reader server with high-quality TTS options, synchronized highlighting, and audiobook export for EPUB, PDF, DOCX, TXT, and MD. Sharing my latest release of OpenReader v3.0.0, an open-source text-to-speech document reader and audiobook exporter. It has been live for over a year now, and slowly has gained 300+ GitHub stars. What is OpenReader? A Next.js web app for reading and listening to EPUB, PDF, TXT,… 9 r/LocalLLaMA community 3mo ago Audio input not accepted with llamacpp for Nemotron 3 nano Omni ? Llama-server does not accept audio input (or video for that matter) with Nemotron 3 nano omni (unsloth). I’m on a recent build of llamacpp and I redownloaded Nemotron, and I have the mmproj loaded too. It still accepts images, but not audio, in fact the audio input option on the… 35 llama.cpp releases dev-tools 3mo ago b9169 mtmd: add chunks and fix preproc for qwen3a ( #23073 ) mtmd: add chunks and fix preproc for qwen3a add attn_mask limit mtmd_chunk size (avoid blow up memory) correct audio tokens re-order the set_input case remove attn_mask macOS/iOS: macOS Apple Silicon (arm64) macOS Apple… 7 r/LocalLLaMA community 3mo ago Adding E4B audio encoder to larger models I am curious if anyone here has tried doing this, I did a bit of digging and it seems like it would be easier to do then I first thought and would like to ask ask for correction if my assumptions are wrong. Here is how I would go about it: Extract the 300mb audio encoder from… 22 r/LocalLLaMA community 3mo ago Qwen 3.6 27B: IQ3XXS KV Q8 vs Q4XL KV Q4 (262K context) hey yall. So I have a 24GB gpu. What do you think is better? I am using unsloth quants. Both are UD quants. I need 262K context for my hermes agent and use case. Both setups fit perfectly in vram. I have heard that Qwen 3.6 27B is quite good even with Q4 KV. I am using LM studio… 27 arXiv — Machine Learning research 3mo ago AudioMosaic: Contrastive Masked Audio Representation Learning arXiv:2605.14231v1 Announce Type: new Abstract: Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generative reconstruction objectives, contrastive approaches… 11 arXiv — NLP / Computation & Language research 3mo ago From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents arXiv:2605.15104v1 Announce Type: new Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling… 19 r/LocalLLaMA community 3mo ago Llama-Studio, WebUI for llama-server Management Hey all, I have built myself a WebUI for configuring and managing llama-server sessions, and want to share the code and concept. Python and a bit of JS. Hack away! Local only. https://github.com/m94301/llama-studio The major use case is running various instances of llama-server… 11 r/LocalLLaMA community 3mo ago Scenema Audio: Zero-shot expressive voice cloning and speech generation We've been building Scenema Audio as part of our video production platform at scenema.ai, and we're releasing the model weights and inference code. The core idea: emotional performance and voice identity are independent. You describe how the speech should be performed (rage,… 17 Hugging Face Daily Papers research 3mo ago Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition Abstract Research identifies studio-bias in multilingual ASR fine-tuning and proposes R-MFT method to improve spontaneous speech performance while maintaining efficiency. AI-generated summary Fine-tuning multilingual ASR models like Whisper for low-resource languages often… 20 r/LocalLLaMA community 3mo ago running Qwen 3.6 35b A3B on 2x 5060TI i ran Qwen 3.6 35b A3B two 5060TI 16gb ( 32 gb vram also i have 32gb dram but i don't like offloading ) i used Q4 on LM Studio to get full context and i get 90t/s any tricks to optimze this more to upgrade to Q6 or Q8 ? thanks ! another thing if you recommend somthing for… 11 r/MachineLearning community 3mo ago Scenema Audio: Zero-shot expressive voice cloning and speech generation [N] We've been building Scenema Audio as part of our video production platform at scenema.ai, and we're releasing the model weights and inference code. The core idea: emotional performance and voice identity are independent. You describe how the speech should be performed (rage,… 37 Page 8 of 8 · 370 articles ← Newer