News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — Machine Learning research 1mo ago Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control arXiv:2607.12856v1 Announce Type: new Abstract: Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well requires planning hours ahead under storage constraints. Model predictive control… 9 arXiv — Machine Learning research 1mo ago Contrastive-Collapsed Loss for Flexible and Geometrically Optimal Embeddings and Faster Convergence arXiv:2607.12916v1 Announce Type: new Abstract: In this work, we introduce CoCo, a loss function aimed at learning normalized and well-structured representations. The proposed loss encourages intra-class collapse and inter-class contrast while preserving sufficient flexibility… 18 arXiv — NLP / Computation & Language research 1mo ago TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation arXiv:2607.11898v1 Announce Type: new Abstract: Large-scale text corpora have become a quiet bottleneck in modern NLP, not just in storage, but in the accumulated cost of training, fine-tuning, and continual learning. We propose a text dataset distillation framework that reduces… 13 arXiv — Machine Learning research 1mo ago Predictive Modeling of High-Altitude Clear Air Turbulence in the United States: A Machine Learning Approach arXiv:2607.11899v1 Announce Type: cross Abstract: High-altitude Clear Air Turbulence (CAT) poses significant risks to aviation safety due to its unpredictability and challenges in detection. This study leverages machine learning models to improve CAT prediction within U.S.… 24 arXiv — NLP / Computation & Language research 1mo ago Transforming LLMs into Efficient Cross-Encoders via Knowledge Distillation for RAG Reranking arXiv:2607.11933v1 Announce Type: new Abstract: Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-time deployment. We address this by fine-tuning LLaMA 3 (8B) as a drop-in… 25 arXiv — Machine Learning research 1mo ago Contrastive Joint-Embedding Prediction for Representation Learning in Structural MRI arXiv:2607.11962v1 Announce Type: cross Abstract: Self-supervised learning offers a compelling approach for medical imaging, where labeled data are scarce and acquisition costs are high. We present COJEPA, a self-supervised framework for volumetric brain MRI that combines a… 28 arXiv — Machine Learning research 1mo ago HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning arXiv:2607.11998v1 Announce Type: cross Abstract: High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper presents a video-based and high performance computing (HPC) enabled deep… 23 arXiv — NLP / Computation & Language research 1mo ago QUBO-Optimized Evidence Selection for Retrieval-Augmented Question Answering with Unconventional Solvers arXiv:2607.12334v1 Announce Type: new Abstract: Retrieval-augmented question answering depends on selecting evidence passages that jointly support answer generation. However, many RAG pipelines rely on top-\(k\) ranking, where passages are selected mainly by individual relevance… 12 arXiv — NLP / Computation & Language research 1mo ago Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL arXiv:2607.12341v1 Announce Type: new Abstract: Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches… 4 arXiv — NLP / Computation & Language research 1mo ago FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis arXiv:2607.11464v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific questions. Graph-based RAG approaches, such as GraphRAG, enhance retrieval by capturing… 30 arXiv — NLP / Computation & Language research 1mo ago On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage arXiv:2607.12257v1 Announce Type: cross Abstract: On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful, and at what cost, is unmeasured for a deployable small model. This study fixes one 4B… 13 arXiv — NLP / Computation & Language research 1mo ago The Sound of Absence: Audio-Language Embedding Models Struggle with Negation arXiv:2607.12290v1 Announce Type: cross Abstract: Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated… 5 arXiv — NLP / Computation & Language research 1mo ago Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks arXiv:2511.04689v3 Announce Type: replace Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over… 28 arXiv — NLP / Computation & Language research 1mo ago Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation arXiv:2601.11443v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful approach for enhancing large language models' question-answering capabilities through the integration of external knowledge. However, when adapting RAG systems to… 14 arXiv — NLP / Computation & Language research 1mo ago Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective arXiv:2603.14217v3 Announce Type: replace Abstract: In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems for… 21 Vercel — AI dev-tools 1mo ago Vercel Workflows trace viewer now has a minimap The trace viewer for Vercel Workflows now has a minimap that shows the entire run in one view. Drag the viewport to pan, resize it to zoom, or drag across the minimap to select a new section. Click anywhere on the minimap to jump to that point, or scroll and use arrow keys to… 7 r/LocalLLaMA community 1mo ago built a memory pipeline on Qwen3 235B A22B Instruct 2507 that scored #1 on LongMemEval-S (470/500) while being ~10x more token efficient than the next best system Over past ~10 months I've been iterating on my memory system so I can make a proper assistant, like Rick's garage from Rick and Morty. I benched my latest iteration and it scored top out of any system I know of (470/500 on LongMemEval-S), while being way more token efficient and… 11 r/LocalLLaMA community 1mo ago Kimi K3 in the next few hours. Deepseek V4 GA later in the week. New Liquid models. New Mistral models sometime this month. And some rumours suggest GLM 5.5 is coming in August. Openweight AI is eating good. dam bois we eating good this week ngl, The velocity of the open_weight ecosystem right now is hitting a point where proprietary, closed-source APIs are losing their leverage on compute intelligence. When you have DeepSeek V4 dropping native MXFP4 mixtures of experts with massive… 10 r/MachineLearning community 1mo ago New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R] Can LLM agents coordinate in long-horizon, open-ended worlds? We evaluate 13 modern LLMs in a new benchmark where agents must work together to explore, communicate, trade resources, craft tools, build structures, and fight mobs. TL;DR : Most agents struggle, averaging only ~6%… 27 r/LocalLLaMA community 1mo ago Nemotron-3-Embed 1B/8B https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16 https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16 Nemotron-3-Embed-8B-BF16 is a versatile text embedding model trained by NVIDIA and optimized for retrieval and semantic similarity tasks. It provides strong multilingual… 24 arXiv — Machine Learning research 1mo ago Multimodal Routing for Interpretable, Robust, and Auditable Clinical Prediction arXiv:2607.09982v1 Announce Type: new Abstract: Electronic health record (EHR) data are inherently multimodal, and leveraging multiple modalities can improve predictive performance. However, most existing approaches rely on deep fusion, which obscures how individual modalities… 25 arXiv — Machine Learning research 1mo ago MLPs are Hebbians: Constructing Efficient Fact-Storing MLPs for Transformers arXiv:2607.10034v1 Announce Type: new Abstract: Large language models (LLMs) store factual knowledge in their parameters. While recent work has shown that this knowledge resides in MLP layers, existing constructive and mechanistic interpretability models of fact-storage in LLMs… 6 arXiv — Machine Learning research 1mo ago Distance-Preserving Embeddings in Inhomogeneous Random Graphs arXiv:2607.10074v1 Announce Type: new Abstract: Graph machine learning provides powerful tools for understanding complex networks and learning meaningful node representations. A central challenge, however, is designing embeddings with minimal distortion of both local and global… 10 arXiv — Machine Learning research 1mo ago Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model arXiv:2607.10285v1 Announce Type: new Abstract: We study how unsupervised autoencoders trained on microscopic spin configurations from the Ising model learn macroscopic, theory-relevant variables underlying the data-generating process. Without embedding domain knowledge, we… 34 arXiv — Machine Learning research 1mo ago ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv:2607.10481v1 Announce Type: new Abstract: Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this work, we investigate a critical source of this… 6 arXiv — Machine Learning research 1mo ago EvidentialRAG: Quantifying and Mitigating Information Conflict in Multi-Source Retrieval-Augmented Generation via Evidential Deep Learning arXiv:2607.10491v1 Announce Type: new Abstract: Retrieval-augmented generation grounds large language models in external evidence, but most pipelines still treat retrieved passages as deterministic and mutually consistent context. In open information environments, retrieved… 24 arXiv — Machine Learning research 1mo ago Learning from Noise: Effective-Rank Collapse and Out-of-Distribution Rejection in Restricted Boltzmann Machines arXiv:2607.10506v1 Announce Type: new Abstract: Restricted Boltzmann machines (RBMs) represent data by shaping an energy landscape over visible and hidden configurations, but their discriminative use is fragile under out-of-distribution (OOD) inputs: samples outside the training… 25 arXiv — Machine Learning research 1mo ago On the modality gap and the contrastive loss in multi-modal representation learning arXiv:2607.10698v1 Announce Type: new Abstract: We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE… 37 arXiv — NLP / Computation & Language research 1mo ago Index SLM Technical Report arXiv:2607.09885v1 Announce Type: new Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion… 27 arXiv — NLP / Computation & Language research 1mo ago Global Merger-Arbitrage Forecasting with Language Models arXiv:2607.09921v1 Announce Type: new Abstract: We present a language-model forecasting system for merger arbitrage, a specialized high-stakes financial setting in which the task is to predict the outcome of announced M\&A deals. Unlike prior work on judgmental forecasting with… 5 arXiv — NLP / Computation & Language research 1mo ago Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG arXiv:2607.10626v1 Announce Type: new Abstract: LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled… 16 arXiv — NLP / Computation & Language research 1mo ago Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG arXiv:2607.10798v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) is often evaluated with clean evidence, yet real retrieval can return topically relevant but unreliable content: false text and misleading images from corrupted metadata, entity… 14 arXiv — NLP / Computation & Language research 1mo ago Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry arXiv:2607.11053v1 Announce Type: new Abstract: Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use of LLMs in collaborative and agentic pipelines raises questions about the extent to which… 35 arXiv — NLP / Computation & Language research 1mo ago Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing arXiv:2607.11400v1 Announce Type: new Abstract: Legal information processing spans retrieval, entailment and judgment prediction problems, requiring text matching, reasoning and robust generalisation with limited supervision. We report Team DU's participation in all five tasks… 4 arXiv — NLP / Computation & Language research 1mo ago GEIS: A Generation-Evaluation-Improvement Loop of Agent Skills for Long-Form Article Generation arXiv:2607.11503v1 Announce Type: new Abstract: Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating… 19 arXiv — NLP / Computation & Language research 1mo ago RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM arXiv:2607.11683v1 Announce Type: new Abstract: Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval.… 38 arXiv — NLP / Computation & Language research 1mo ago How Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation? arXiv:2607.11783v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has been increasingly adopted to reduce hallucinations and strengthen the factual grounding of large language models (LLMs). While robustness to errors in the retrieval process has been… 16 arXiv — NLP / Computation & Language research 1mo ago Improved Answer Selection with Pre-Trained Word Embeddings arXiv:1708.04326v1 Announce Type: cross Abstract: This paper evaluates existing and newly proposed answer selection methods based on pre-trained word embeddings. Word embeddings are highly effective in various natural language processing tasks and their integration into… 37 Vercel — AI dev-tools 1mo ago Vercel Blob now supports consistent reads on private storage Vercel Blob now supports consistent reads on private storage. Pass useCache: false to get() or presignUrl() for a read that reflects the latest write. A blob written to a fresh pathname has no existing cached entry, so reads reflect the latest write immediately. When you… 26 arXiv — NLP / Computation & Language research 1mo ago Sticky Routing: Training MoE Models for Memory-Efficient Inference arXiv:2607.08780v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping between slow storage and fast memory on edge… 25 arXiv — Machine Learning research 1mo ago Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal arXiv:2607.08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions… 13 arXiv — Machine Learning research 1mo ago Group Invariant Spectral Embedding arXiv:2607.08987v1 Announce Type: new Abstract: Spectral embedding methods are widely used for dimensionality reduction and clustering of high-dimensional datasets with intrinsic low-dimensional structures. Although many datasets of practical interest exhibit invariance under… 19 arXiv — Machine Learning research 1mo ago Quantum Circuits in Diffusion Models: A Fair-Comparison Study and a Mechanistic Analysis of Angle-Embedding Failures arXiv:2607.09108v1 Announce Type: new Abstract: We study the integration of variational quantum circuits (VQCs) into diffusion models through a squeeze-and-excitation (SE) channel-modulation scaffold that isolates the quantum contribution. Using a role-matched classical control… 10 arXiv — NLP / Computation & Language research 1mo ago Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning arXiv:2607.09287v1 Announce Type: cross Abstract: Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused… 37 arXiv — Machine Learning research 1mo ago Similarity search generalisation in contrastive learning with InfoNCE loss arXiv:2607.09405v1 Announce Type: new Abstract: Similarity search is a primary application of embedding models trained by contrastive learning. For one of the most popular contrastive learning loss functions, InfoNCE, we show that the population risk with $k$ negative samples is… 7 arXiv — NLP / Computation & Language research 1mo ago Neural Collapse Is Forbidden: Information Floors in Language Models arXiv:2607.09487v1 Announce Type: cross Abstract: Within-class variance in language-model representations is commonly read as incomplete neural collapse. We argue it is allocated information storage, and that the allocation obeys a law. A one-line centering identity voids a… 25 arXiv — Machine Learning research 1mo ago Leveraging Interpretable Tsetlin Machine for PDF Malware Detection arXiv:2607.09290v1 Announce Type: cross Abstract: In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality. However, these same capabilities… 10 arXiv — NLP / Computation & Language research 1mo ago Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation arXiv:2607.09349v1 Announce Type: new Abstract: Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass… 38 arXiv — NLP / Computation & Language research 1mo ago An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? arXiv:2607.09053v1 Announce Type: new Abstract: Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be… 14 arXiv — NLP / Computation & Language research 1mo ago AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs arXiv:2607.09092v1 Announce Type: new Abstract: Knowledge graphs (KGs) are often automatically constructed from large-scale corpora, but they inevitably contain factual errors due to noisy sources and extraction failures, and verifying them reliably at industrial scale remains a… 34 Page 9 of 10 · 500 articles ← Newer Older →