News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — Machine Learning research 14d ago S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring arXiv:2607.27913v1 Announce Type: new Abstract: Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global… 9 arXiv — Machine Learning research 14d ago Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction arXiv:2607.28079v1 Announce Type: new Abstract: Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However, existing benchmarks often suffer from limited task diversity, fragmented… 24 arXiv — NLP / Computation & Language research 14d ago LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer… 36 arXiv — NLP / Computation & Language research 14d ago Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built… 36 arXiv — NLP / Computation & Language research 14d ago Correlation between prosody and pragmatics: A case study of the discourse marker h\=al\=a `now' in Persian arXiv:2607.28359v1 Announce Type: new Abstract: The Persian discourse marker h\=al\=a ('now') exhibits remarkable multifunctionality, extending far beyond its temporal adverbial role to encompass a variety of pragmatic functions. This study presents a pragmatic and acoustic… 35 arXiv — NLP / Computation & Language research 14d ago MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist and corrupt future behavior. Existing systems improve storage and retrieval,… 18 arXiv — NLP / Computation & Language research 14d ago CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising arXiv:2607.28236v1 Announce Type: cross Abstract: Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout.… 19 arXiv — NLP / Computation & Language research 14d ago GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation arXiv:2607.28397v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based… 35 Hugging Face Daily Papers research 14d ago BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Abstract Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we… 29 Hugging Face Daily Papers research 14d ago ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Abstract Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this… 24 TechCrunch — AI news-outlet 14d ago AI hedge fund Situational Awareness may have sold its public portfolio, but it still has its Anthropic shares The former OpenAI researcher’s fund was forced to unwind public equities after leveraged public bets plummeted. But he still has cards to play. 31 llama.cpp releases dev-tools 14d ago b10194 ggml-cuda: Allow transpose-free gemmv computation ( #26171 ) When matrix's weights are shaped 1xK is leverage a transpose-free computation to use mat_mul_vec_f. Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled)… 20 Hugging Face Daily Papers research 15d ago CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation Abstract On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii)… 18 arXiv — Machine Learning research 15d ago FloDR: An invertible dimensionality reduction method based on a normalising flow arXiv:2607.26278v1 Announce Type: new Abstract: It is common for two-dimensional embeddings of high-dimensional data to be read far beyond what they can support. Distances in and between clusters, the meaning behind empty spaces, and the amount of structure hidden at each point… 22 arXiv — Machine Learning research 15d ago RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning arXiv:2607.26339v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. We… 28 arXiv — Machine Learning research 15d ago ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling arXiv:2607.26369v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) has been widely adopted in transformer-based large language models. However, its log-linear frequency schedule, originally designed to produce long-term attention decay, limits its adoption in… 38 arXiv — Machine Learning research 15d ago From Conceptual Hydrologic Models to Conceptually Interpretable Neural Networks: A Snow-Water Mass-Conserving-Perceptron Framework for Discovering Catchment-Scale Precipitation-Storage-Runoff Representations arXiv:2607.26492v1 Announce Type: new Abstract: The Mass-Conserving Perceptron (MCP) establishes a modeling paradigm in which conceptual hydrologic models can be reformulated as physically constrained, conceptually interpretable neural networks. Here, we develop a snow-water MCP… 35 arXiv — Machine Learning research 15d ago Simultaneous Coverage and Efficiency Guarantee in Online Conformal Prediction arXiv:2607.26577v1 Announce Type: new Abstract: Adaptive conformal inference (ACI) of Gibbs and Cand{\`e}s and its variants are the standard approach to online conformal prediction under distribution shift, but they suffer from three fundamental limitations. First, their… 30 arXiv — Machine Learning research 15d ago Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation arXiv:2607.26599v1 Announce Type: new Abstract: Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine. We focus on the conditional average treatment effect (CATE), a standard estimand for… 6 arXiv — Machine Learning research 15d ago RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment arXiv:2607.26631v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large… 24 arXiv — NLP / Computation & Language research 15d ago Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight arXiv:2607.26250v1 Announce Type: new Abstract: Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address only one component at a time. This paper presents Robostreet Flow, a freight… 38 arXiv — NLP / Computation & Language research 15d ago CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG arXiv:2607.26470v1 Announce Type: new Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history,… 32 arXiv — NLP / Computation & Language research 15d ago Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms arXiv:2607.26497v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost… 35 arXiv — NLP / Computation & Language research 15d ago WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained index, whereas construction receives no signal from downstream use. We present… 15 arXiv — NLP / Computation & Language research 15d ago ARC-Encoder: learning compressed text representations for large language models arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective… 8 arXiv — NLP / Computation & Language research 15d ago Ensembling LLM-Induced Decision Trees for Explainable and Robust Error Detection arXiv:2512.07246v3 Announce Type: replace Abstract: Error detection (ED), which aims to identify incorrect or inconsistent cell values in tabular data, is important for ensuring data quality. Recent state-of-the-art ED methods leverage the pre-trained knowledge and semantic… 21 Hacker News — AI on Front Page community 15d ago The Productivity Mirage Article URL: https://frantic.im/mirage/ Comments URL: https://news.ycombinator.com/item?id=49104335 Points: 216 # Comments: 71 23 r/LocalLLaMA community 15d ago Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar? I bought an RTX 5090 last year just to run 27B models natively. I even fine-tuned it with my own data using LoRA, building RAGs and was pretty damn happy with the results at first. But, Q8 quantization 130k context was barely squeezing through. Naturally, I bought two RTX 6000… 37 r/LocalLLaMA community 15d ago Budget Inference: A GPU for dense models vs. More RAM for MoE models? Hi all, I’m building a budget inference machine primarily for personal use (chat/assistant tasks, possibly some RAG). I'm torn between two hardware paths and would love input from anyone who has actually benchmarked these setups. The Dilemma: Option A (GPU for dense models): Buy… 35 r/MachineLearning community 16d ago Vendor-agnostic ML inference on production edge devices [R] I work on PostSlate, a video editing tool, and this comes out of our own work. We run ML models on-device, face detection and embedding among other things, which means we can't assume anything about the user's GPU. NVIDIA discrete, AMD, Intel integrated, Apple Silicon, all of… 7 Hugging Face Daily Papers research 16d ago Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control Abstract Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes… 30 arXiv — NLP / Computation & Language research 16d ago Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control arXiv:2607.25337v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline… 36 arXiv — NLP / Computation & Language research 16d ago Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs arXiv:2607.25959v1 Announce Type: new Abstract: Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a… 28 arXiv — NLP / Computation & Language research 16d ago VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents arXiv:2607.24748v1 Announce Type: cross Abstract: Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple pages, mixing text with layout cues, tables, charts, and figures. This work studies multimodal… 36 arXiv — NLP / Computation & Language research 16d ago The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance arXiv:2607.24767v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relevant information to use as source material during text generation. A critical yet… 11 arXiv — NLP / Computation & Language research 16d ago LLM Scheming Inversely Scales with Pretraining Language Coverage arXiv:2607.24769v1 Announce Type: cross Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned… 36 arXiv — NLP / Computation & Language research 16d ago From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance arXiv:2607.24791v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow.… 23 arXiv — NLP / Computation & Language research 16d ago Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging. Pure parametric Large Language models (LLMs) do not contain… 10 arXiv — NLP / Computation & Language research 16d ago Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation arXiv:2607.24884v1 Announce Type: cross Abstract: Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-code examples, repository context, and project-specific APIs may provide… 14 arXiv — NLP / Computation & Language research 16d ago Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs arXiv:2607.25600v1 Announce Type: cross Abstract: Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evidence and unnecessary computation. We investigate whether verbalized confidence from… 14 arXiv — NLP / Computation & Language research 16d ago Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despite their potential, existing work has predominantly focused on enhancing either… 32 arXiv — NLP / Computation & Language research 16d ago VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation arXiv:2510.09733v2 Announce Type: replace Abstract: Visual Retrieval-Augmented Generation (VRAG) has emerged as a promising paradigm for equipping Vision-Language Models (VLMs) with external visual evidence, enabling them to go beyond parametric knowledge when answering visually… 36 arXiv — NLP / Computation & Language research 16d ago Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore arXiv:2601.15050v5 Announce Type: replace Abstract: Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. This… 28 r/MachineLearning community 16d ago How to deal with text only vector search across multimodal embedding space? [D] My data set is a list of images, each equipped with a a couple sentences of text. A user would search primarily with text only. My default approach is using BM25, but how would I facilitate searching with a vector DB and a model that embeds vectors in a multimodal combined… 37 arXiv — Machine Learning research 17d ago Hierarchical Grading in Large Language Models arXiv:2607.22757v1 Announce Type: new Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention,… 30 arXiv — Machine Learning research 17d ago LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers arXiv:2607.22804v1 Announce Type: new Abstract: Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture and storage (CCS), geothermal development, and extraction of natural resources. Existing… 38 arXiv — Machine Learning research 17d ago OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence arXiv:2607.22805v1 Announce Type: new Abstract: We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Service to automatically design service-adaptive models for heterogeneous edge environments. The… 32 arXiv — Machine Learning research 17d ago PerturbPFN: Probing the Limits of Synthetic Priors in Drug Perturbation Modelling arXiv:2607.23447v1 Announce Type: new Abstract: Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space.… 29 arXiv — NLP / Computation & Language research 17d ago Not All LLM Reasoning is Visible in the Chain-of-Thought arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically… 16 arXiv — NLP / Computation & Language research 17d ago IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems arXiv:2607.23322v1 Announce Type: new Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of… 10 Page 5 of 10 · 500 articles ← Newer Older →