News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — NLP / Computation & Language research 24d ago Thinking in Video: Can Video Generators Really Reason About the Real World? arXiv:2607.17523v1 Announce Type: cross Abstract: Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm… 25 arXiv — NLP / Computation & Language research 24d ago Salience Induction against Multi-Hop RAG Agents: Threat and Defense arXiv:2607.17535v1 Announce Type: cross Abstract: Agentic retrieval-augmented generation (RAG) systems increasingly retrieve external evidence and orchestrate tools for knowledge-intensive applications. In Multi-Hop question answering, agents chain facts across documents.… 12 arXiv — NLP / Computation & Language research 24d ago D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation arXiv:2607.17538v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces… 4 arXiv — NLP / Computation & Language research 24d ago AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning arXiv:2607.17849v1 Announce Type: cross Abstract: Approximately 3,000 of the 4,500 oracle bone script (OBS) characters remain undeciphered due to fragmentary inscriptions and sparse evidence. Current AI approaches fail to replicate expert workflows that integrate form analysis,… 26 arXiv — NLP / Computation & Language research 24d ago FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering arXiv:2607.18102v1 Announce Type: cross Abstract: Financial question answering over U.S. Securities and Exchange Commission (SEC) filings requires retrieving and synthesizing heterogeneous evidence dispersed across long, standardized, and highly redundant disclosures. Existing… 33 Hugging Face Daily Papers research 25d ago RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM Abstract Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an open-source modular GraphRAG… 4 arXiv — Machine Learning research 25d ago DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging arXiv:2607.15986v1 Announce Type: new Abstract: This application paper presents DebrisTracer, a framework for the reliable tracking of debris in hypervelocity impact fast imaging. These noisy and highly specific datasets capture the ejection of a large number of debris fragments… 12 arXiv — Machine Learning research 25d ago DELUGE: Towards Continental-Scale Daily Pluvial Flood Damage Prediction via Interpretable Conditioning on Foundation Model Embeddings arXiv:2607.16050v1 Announce Type: new Abstract: Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to… 27 arXiv — Machine Learning research 25d ago A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing arXiv:2607.16183v1 Announce Type: new Abstract: To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical… 15 arXiv — Machine Learning research 25d ago AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of… 31 arXiv — Machine Learning research 25d ago E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding arXiv:2607.15536v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) captures scenes by coupling explicit geometry (position, covariance) with view-dependent photometry (Spherical Harmonics). However, building $\mathrm{SE}(3)$-equivariant architectures on these… 37 arXiv — NLP / Computation & Language research 25d ago Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory arXiv:2303.01421v2 Announce Type: replace Abstract: Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected… 15 arXiv — NLP / Computation & Language research 25d ago Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization arXiv:2601.22402v2 Announce Type: replace Abstract: Rotary Positional Embeddings (RoPE) have become the standard for Large Language Models (LLMs) due to their ability to encode relative positions through geometric rotation. However, we identify a significant limitation we term… 23 arXiv — NLP / Computation & Language research 25d ago Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering arXiv:2604.04593v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surface hard negatives that are semantically close to the query but describe clinically… 36 r/LocalLLaMA community 26d ago Thinking Machines' best public Tinker result used Qwen3-235B, not Inkling. is the base model actually that important? I Went through the Inkling model card and the Bridgewater/Tinker case study instead of the press coverage. Coverage mostly quoted the 97.1% AIME 2026 number; the rest of the table tells a more mixed story. AIME 2026: Inkling 97.1%, GLM 5.2 99.2%, Fable 5 and GPT-5.6 Sol both… 5 r/MachineLearning community 26d ago Interactive map of GPT-2's token embedding space - tap any token and explore [P] 32,070 alphabetic tokens from GPT-2-small's WTE, no forward pass and no context. Works on mobile. Pinch to zoom, tap a token to see its nearest connections, tap a neighbour to walk the graph. Search box to jump anywhere. Layout is t-SNE over a compressed representation of the… 16 r/MachineLearning community 26d ago GPT-2 Small’s embedding geometry around “Trump”: discretized vs. continuous nearest neighbours [P] This visualization looks at the token “Trump” in GPT-2 Small’s static embedding table, before attention or context is applied. The top plot is a t-SNE projection of 32,070 alphabetic tokens with at least two characters. The two graphs below compare Trump’s nearest neighbours… 22 Ars Technica — AI news-outlet 27d ago Will AI fix prior authorization—or make it worse? The government is piloting a program that uses AI for insurance-coverage decisions. 31 r/LocalLLaMA community 27d ago SigLIP 2 text embedding on CPU with Rust + ONNX We’re building a robotics data platform with a lot of images, video, and text metadata. For search, we use SigLIP 2. GPUs handle batched asynchronous image/video embedding and indexing, while this small Rust + ONNX Runtime service handles live text queries on CPU. Both land in… 37 r/LocalLLaMA community 27d ago [RESEARCH] Breaking the 1-bit Floor: Achieving "Negative-Bit Quantization" (NBQ) via Phase-Inverted Tensor Embedding (satire) Hey everyone, I’ve spent the last three weeks compiling custom llama.cpp forks and running imatrix maps on a modified CUDA kernel setup, and the numbers don’t lie. We’ve been looking at model compression completely wrong. Everyone in the community has assumed that 1-bit… 37 Hugging Face Daily Papers research 27d ago Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving Abstract Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open challenge, and… 20 r/MachineLearning community 28d ago EU AI Act OpenRAG: 933 legally structured chunks and BGE-M3 embeddings in one SQLite file [P] I have released EU AI Act OpenRAG, a downloadable corpus of Regulation (EU) 2024/1689 designed for RAG and legal-NLP experimentation. Instead of sliding character windows, the corpus chunks on the Regulation’s legal structure: one chunk per article paragraph one per recital one… 32 Hugging Face Daily Papers research 28d ago GRASP: GRanularity-Aware Search Policy for Agentic RAG Abstract Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when to retrieve, whether to use… 22 Hugging Face Daily Papers research 28d ago MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Abstract MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow models with human preferences and… 8 arXiv — Machine Learning research 28d ago RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences arXiv:2607.14180v1 Announce Type: new Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset. However, they are vulnerable to model exploitation where data coverage is thin. Prior… 9 arXiv — Machine Learning research 28d ago NeuroGRIP: Retrieval-Augmented Graph Refinement for Knowledge-Grounded EEG Seizure Diagnosis arXiv:2607.14314v1 Announce Type: new Abstract: Seizure diagnosis from EEG signals is a critical yet persistently challenging task, due to the complicated neural dynamics and the spurious connections in inter-channel modeling. While spatial-temporal graph neural networks… 32 arXiv — Machine Learning research 28d ago What's in a Smoothness Constant? Tighter Rates for Local SGD with Bounded Second-order Heterogeneity arXiv:2607.14731v1 Announce Type: new Abstract: Local SGD, also known as Federated Averaging, is a widely used distributed optimization algorithm. Although Local SGD often outperforms alternatives such as Mini-batch SGD in practice, theory still only partially explains when and… 29 arXiv — NLP / Computation & Language research 28d ago Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation arXiv:2607.14895v1 Announce Type: cross Abstract: Reasoning language models (RLMs) have demonstrated impressive performance in domains such as mathematics and coding. These domains permit reliable verification of model outputs, which is important for enabling the reinforcement… 35 arXiv — Machine Learning research 28d ago Mutable Low-Rank Sketches for Retrain-Free Recommendation arXiv:2607.15242v1 Announce Type: new Abstract: A common bottleneck in two-stage recommendation is embedding staleness: when a user rates a new item, their embedding remains fixed until the next retrain cycle. We propose mutable sketches, which store each user's preferences in a… 19 arXiv — Machine Learning research 28d ago NexForge: Scaling Executable Agent Tasks via Requirement-First Synthesis arXiv:2607.14186v1 Announce Type: cross Abstract: Scaling executable agent training data is bottlenecked by substrate-first methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual expansion of the substrate, each… 19 arXiv — NLP / Computation & Language research 28d ago Implicit Reasoning Steering via Concept Chaining arXiv:2607.14242v1 Announce Type: new Abstract: Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how final decisions are formed. We study whether this… 30 arXiv — NLP / Computation & Language research 28d ago DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA arXiv:2607.14400v1 Announce Type: new Abstract: This paper describes DS@GT ARC's submission to the CLEF 2026 LongEval Task 4 on Retrieval-Augmented Generation (RAG). In this submission, we examine a divergence between traditional natural language evaluation metrics and citation… 23 arXiv — NLP / Computation & Language research 28d ago Latent Trajectory Discrimination for AI-Generated Text Detection arXiv:2607.14967v1 Announce Type: new Abstract: Most existing approaches to AI-Generated Text Detection (AIGTD) treat documents as static objects and base their decisions on aggregate statistics or globally compressed embeddings. However, this perspective overlooks the… 9 arXiv — NLP / Computation & Language research 28d ago Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya arXiv:2607.15209v1 Announce Type: new Abstract: Multilingual pre-trained language models (PLMs) exhibit degraded performance on low-resource, non-Latin-script languages, driven by high out-of-vocabulary (OOV) rates and excessive subword fragmentation that result from… 32 NVIDIA Developer Blog official-blog 28d ago Q&A: How Capcom Brought Path Tracing to RE ENGINE Across PRAGMATA and Resident Evil Requiem Capcom's RE ENGINE team set out to bring path tracing into two shipping titles at once, Resident Evil Requiem and PRAGMATA, each with a different visual... 13 VentureBeat — AI news-outlet 28d ago The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector… 27 VentureBeat — AI news-outlet 28d ago The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty… 26 NVIDIA Developer Blog official-blog 28d ago Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage... 16 Hugging Face Daily Papers research 29d ago AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Abstract As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant… 20 arXiv — Machine Learning research 29d ago The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models arXiv:2607.13612v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performance rather than a normative principle. We show that the choice of anti-collapse… 12 arXiv — Machine Learning research 29d ago The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model arXiv:2607.13660v1 Announce Type: new Abstract: Contrastive Language-Image Pretraining (CLIP) representations form a semantic embedding space governed by cosine similarity, reflecting an intrinsic hyperspherical geometry. However, existing probabilistic interpretations typically… 28 arXiv — Machine Learning research 29d ago Leveraging unlabelled data for generalizable neural population decoding arXiv:2607.14086v1 Announce Type: new Abstract: Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session… 27 arXiv — Machine Learning research 29d ago FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents arXiv:2607.13035v1 Announce Type: cross Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage… 28 arXiv — Machine Learning research 29d ago Precomputing the Future-Offset Average in TriAttention arXiv:2607.13051v1 Announce Type: cross Abstract: TriAttention is a recent method for shrinking the KV cache of long-reasoning LLMs: it scores each cached key by how much attention it is likely to receive and evicts the lowest-scoring ones. Because a key does not know how far… 38 Hugging Face Daily Papers research 29d ago Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering Abstract Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework designed to unmask visual deception… 30 llama.cpp releases dev-tools 29d ago b10032 cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) ( #25545 ) cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) chore : remove indentation of #pragma unroll cuda : remove unnecessary kernel template… 37 TechCrunch — AI news-outlet 29d ago Inside Ode with Anthropic, the startup betting AI services are the future of enterprise Can a handful of engineers really do the work of an army of consultants? That’s the bet behind Ode with Anthropic — the joint venture dedicated to embedding forward-deployed engineers in enterprise firms, backed by Anthropic, Blackstone, Hellman &… 36 TechCrunch — AI news-outlet 1mo ago Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not models Anthropic-backed Ode launches as AI labs bet that embedding forward-deployed engineers inside enterprises is the key to accelerating enterprise AI adoption. 27 arXiv — Machine Learning research 1mo ago Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate arXiv:2607.12266v1 Announce Type: new Abstract: Mixed-precision quantization must decide which parts of a model to keep at higher precision. A common premise, shared by sensitivity-based methods such as HAWQ and CoopQ, is that the loss from quantizing a set of layers can be… 38 arXiv — Machine Learning research 1mo ago SinAE: A Single-Architecture Flow-Matching Autoencoder for Cross-Domain Atomic Systems arXiv:2607.12380v1 Announce Type: new Abstract: Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet their generative pipelines remain fragmented across domains, each with its Small molecules, crystals, and proteins all reduce to atoms in 3D space, yet… 9 Page 8 of 10 · 500 articles ← Newer Older →