News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow r/LocalLLaMA community 21d ago UPDATE - HuggingHack Is Now On Github Due to encouragement to move my local huggingface project, it is now available on Github! Let me know your thoughts! Repo: https://github.com/tyedalwaves/HuggingHack/   submitted by   /u/TyedalWaves [link]   [comments] 38 r/LocalLLaMA community 21d ago [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio S2 Pro, Voxtral Realtime ASR and two community models OuteTTS TTS and… 7 r/MachineLearning community 21d ago NeurIPS E and D, Average rating 3 and average confidence 4, I can rebuttal and address all their concerns? Do I still have a decent shot or unlikely ?[R] NeurIPS E and D track review are out today and the average rating I received is a 3 and confidence is a 4. I can correct and address all their concerns. Do I still have a genuine shot of getting in or is it basically impossible at this point since none of my scores are a 4 or 5?… 37 r/MachineLearning community 22d ago Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R] Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier. Total cost spread was 10.6x despite published rates… 15 arXiv — Machine Learning research 22d ago Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models arXiv:2607.19399v1 Announce Type: new Abstract: It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger… 34 arXiv — Machine Learning research 22d ago Zero-Shot Heart Rate Variability Forecasting from Consumer Wearables Using Time Series Foundation Models arXiv:2607.20027v1 Announce Type: new Abstract: Short-term Heart Rate Variability (HRV) forecasting could provide clinicians with actionable lead time for detecting autonomic dysfunction and adverse cardiac events. Consumer wearable devices generate fragmented, artifact-rich HRV… 18 arXiv — Machine Learning research 22d ago Evaluating and Mitigating Gender Bias in Pre-trained Embeddings for ML-based Recruitment arXiv:2607.20073v1 Announce Type: new Abstract: AI-based recruitment systems that rely on machine learning models trained on historical CV data, risk perpetuating and amplifying social biases. A key challenge arises in unstructured CV text, where pre-trained language model… 16 arXiv — Machine Learning research 22d ago Active Inference as a Convex Markov Decision Process arXiv:2607.20152v1 Announce Type: new Abstract: Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show… 16 arXiv — Machine Learning research 22d ago Local Stability and Gaussian Smoothing of Quantized Neural Networks arXiv:2607.20153v1 Announce Type: new Abstract: We study Gaussian averaging as a smooth surrogate for quantized neural models. Under bounded local oscillation, we derive a local dimension-dependent bound on |f-g|, linking Gaussian smoothing to the stability analysis of… 38 arXiv — Machine Learning research 22d ago PIER: Physics-Informed Environmental Retrieval for Time-Series Modeling arXiv:2607.20230v1 Announce Type: new Abstract: Accurate modeling of environmental systems is fundamental to scientific understanding and decision-making, yet remains challenging because observations are limited and physical dynamics vary across systems. Retrieval-augmented… 29 arXiv — Machine Learning research 22d ago Classical Hardware Acceleration of Quantum Autoencoders for Real-Time Anomaly Detection in Collider Experiments arXiv:2607.20302v1 Announce Type: new Abstract: Quantum machine learning (QML) algorithms in high energy physics (HEP) can efficiently represent and leverage long-range, high-order correlations in high-dimensional collider data, potentially with fewer parameters and favorable… 16 arXiv — NLP / Computation & Language research 22d ago VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization arXiv:2607.19830v1 Announce Type: new Abstract: Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic facts among entities, rather than relying solely on binary relationships. Despite the advancements in multimodal large… 4 arXiv — NLP / Computation & Language research 22d ago D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios arXiv:2607.19834v1 Announce Type: new Abstract: With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in… 34 arXiv — NLP / Computation & Language research 22d ago emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity arXiv:2607.19848v1 Announce Type: new Abstract: There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools for… 34 arXiv — NLP / Computation & Language research 22d ago Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results arXiv:2607.20090v1 Announce Type: new Abstract: Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields… 15 arXiv — NLP / Computation & Language research 22d ago OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety… 12 arXiv — NLP / Computation & Language research 22d ago AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally arXiv:2607.19363v1 Announce Type: cross Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified… 10 arXiv — NLP / Computation & Language research 22d ago HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions arXiv:2607.19398v1 Announce Type: cross Abstract: Multi-entity compositional questions pose significant challenges to existing retrieval-augmented language models. Conventional methods fall into a dilemma: standard RAG lacks dynamic reasoning, traditional Graph-RAG is limited by… 15 arXiv — NLP / Computation & Language research 22d ago Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology arXiv:2607.19415v1 Announce Type: cross Abstract: High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially… 25 arXiv — NLP / Computation & Language research 22d ago The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation arXiv:2604.26347v2 Announce Type: replace-cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on… 20 llama.cpp releases dev-tools 22d ago b10089 cuda: GET_ROWS quants ( #25962 ) cuda: add k-quant support to GET_ROWS Device-side embedding lookups require GET_ROWS to handle the k-quants used by common GGUF recipes (Q4_K_M stores token_embd as q6_K). Without it the backend rejects the op and the scheduler falls back to the… 31 Stratechery (Ben Thompson) community 23d ago OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips OpenAI accidentally hacked Hugging Face, but the takeaways are more encouraging than people realize. 10 llama.cpp releases dev-tools 23d ago b10085 mtmd : use align_corners for qwen3vl vision position embedding interpolation ( #25781 ) The Qwen3-VL learned position embedding is interpolated to the runtime patch grid with the default bilinear+antialias (align_corners=False) sampling, while the transformers reference uses… 32 arXiv — Machine Learning research 23d ago The Information Shadow: Measuring Structural Limits on What Language Models Can Learn arXiv:2607.18305v1 Announce Type: new Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the information shadow: the region of phenomena that a text-trained learner cannot acquire… 28 arXiv — Machine Learning research 23d ago Physics-Guided Masked Multi-Task Network for Edge-Friendly Battery Health Diagnostics from Sto-chastically Fragmented Charging Profiles arXiv:2607.18330v1 Announce Type: new Abstract: The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of State of Health (SOH) and Remaining Useful Life (RUL) remains severely hindered by task… 26 arXiv — NLP / Computation & Language research 23d ago A Controlled Study of Attention-Only Transformers arXiv:2607.18363v1 Announce Type: cross Abstract: Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder… 21 arXiv — Machine Learning research 23d ago Scalable and Efficient Joint Spiking Embedding Predictive Architecture for Large-Scale Dynamic Graphs arXiv:2607.18412v1 Announce Type: new Abstract: Dynamic graph learning aims to capture evolving structural and semantic patterns in real-world systems, such as fraud detection and recommender systems. Due to the scarcity of labeled data in real-world dynamic graphs, recent… 5 arXiv — Machine Learning research 23d ago Decafs: Disentangled Conditional adversarial Flows arXiv:2607.18755v1 Announce Type: new Abstract: Flow-based models have established state-of-the-art performance in generative modeling across domains, but are hard to interpret due to their complex latent embeddings. In particular, the entanglement of generative factors in the… 5 arXiv — Machine Learning research 23d ago Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals arXiv:2607.18860v1 Announce Type: new Abstract: Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning methods depend mainly on temperature and may therefore miss mechanical precursors that emerge… 29 arXiv — Machine Learning research 23d ago Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation arXiv:2607.19044v1 Announce Type: new Abstract: Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design. Current methods primarily rely on supervised training or fine-tuning with limited datasets, which are… 16 arXiv — NLP / Computation & Language research 23d ago Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives arXiv:2607.18443v1 Announce Type: new Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics… 25 arXiv — NLP / Computation & Language research 23d ago Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv:2607.18666v1 Announce Type: new Abstract: A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead text/image/video retrieval benchmarks but lack audio… 36 arXiv — NLP / Computation & Language research 23d ago AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System arXiv:2607.18825v1 Announce Type: new Abstract: This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the Indian legal context. AILQA leverages a variety of embedding and generative models,… 35 arXiv — NLP / Computation & Language research 23d ago EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation arXiv:2607.18336v1 Announce Type: cross Abstract: Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across… 14 arXiv — NLP / Computation & Language research 23d ago RAGAL: A Frugal, Fully Local Retrieval-Augmented Assistant for Technical Support at a Government Agency arXiv:2607.18756v1 Announce Type: cross Abstract: Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the… 9 r/LocalLLaMA community 23d ago Felix Rieseberg (Anthropic, ElectronJS) has released a free Mac app designed to help people build their own LLMs from scratch. Announcement tweet here. Direct link: Language Model Builder From the site: "Using the default settings, you’ll get a model that writes coherent, grammatical multi-paragraph text in as little as a day. On a MacBook Pro M5 Max you could train a GPT-2-small-class model (~100–150M… 10 Ollama releases dev-tools 23d ago v0.32.2: test: revamp integration test entrpoints (#16560) This refactors the existing integration tests into 3 priumary groups: fast, release, and library. It also refines some of the release tests to drop some of the older models and pick up newer models, while retaining the broad coverage in the library group. 25 r/MachineLearning community 23d ago Looking for feedback on my GPU-accelerated Snake AI project [P] I've been building an AI that learns to play the classic Snake game through reinforcement learning. The goal is to reach high scores while keeping training time as low as possible. The current version averages 86 points (87 is the maximum) after less than 10 hours of training on… 31 arXiv — Machine Learning research 24d ago Fully-sensorized smart-eyewear platform for on-device Machine Learning arXiv:2607.16222v1 Announce Type: new Abstract: This paper presents ARGO, a smart eyewear platform designed to bridge ergonomic comfort, high computational throughput, and energy efficiency. Unlike cloud-dependent solutions, ARGO leverages the STM32N6 microcontroller and its… 11 arXiv — NLP / Computation & Language research 24d ago From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training arXiv:2607.16257v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with… 16 arXiv — Machine Learning research 24d ago Feedback Attribution and Representation Geometry: Metrics for Comparing Individual and Shared Rewards in MARL arXiv:2607.16524v1 Announce Type: new Abstract: Cooperative multi-agent RL systems routinely use team-averaged rewards, a feedback-attribution choice that gives each agent the team outcome regardless of its individual contribution. We ask whether this leaves a measurable… 26 arXiv — Machine Learning research 24d ago Discrete Ricci Curvature on Protein Contact Graphs for Lightweight Fold Classification arXiv:2607.16553v1 Announce Type: new Abstract: Protein fold classification can be approached via sequence-based representations or structural descriptors, but direct comparisons between lightweight handcrafted descriptors and pretrained protein language model embeddings remain… 22 arXiv — Machine Learning research 24d ago A Framework for Early Sepsis Prediction via Self-Supervised (JEPA) and Federated Representation Learning arXiv:2607.16681v1 Announce Type: new Abstract: Early sepsis prediction from electronic health records is challenged by irregular sampling, high missingness, and class imbalance. We systematically compare four modeling paradigms -- self-supervised Joint Embedding Predictive… 28 arXiv — Machine Learning research 24d ago First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers arXiv:2607.16821v1 Announce Type: new Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations:… 11 arXiv — Machine Learning research 24d ago CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation arXiv:2607.16955v1 Announce Type: new Abstract: On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to… 16 arXiv — Machine Learning research 24d ago TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization arXiv:2607.16973v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two underexplored challenges: (1) trained codebook quantizers may expose corpus statistics… 38 arXiv — NLP / Computation & Language research 24d ago RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation arXiv:2607.16431v1 Announce Type: new Abstract: Small-scale language models (SLMs) are attractive for retrieval-augmented generation (RAG) in resource-constrained settings, but their limited capacity makes them highly sensitive to noisy or spurious retrieved evidence. Existing… 14 arXiv — NLP / Computation & Language research 24d ago Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit arXiv:2607.17466v1 Announce Type: new Abstract: Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability… 32 arXiv — NLP / Computation & Language research 24d ago C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference arXiv:2607.17715v1 Announce Type: new Abstract: Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value… 20 arXiv — NLP / Computation & Language research 24d ago ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG arXiv:2607.16208v1 Announce Type: cross Abstract: Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy depends on which multimodal assets are ranked highly enough to enter downstream reasoning;… 6 Page 7 of 10 · 500 articles ← Newer Older →