News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — NLP / Computation & Language research 4d ago Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation arXiv:2609.27844v1 Announce Type: cross Abstract: Claim denial management costs U.S. healthcare approximately $260 billion annually in administrative overhead. Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) can produce fluent clinical text, but… 11 Hacker News — AI on Front Page community 4d ago Linux support is coming to Snapdragon X2 Series Article URL: https://www.qualcomm.com/news/onq/2026/09/snapdragon-summit-agentic-ai-pcs-linux Comments URL: https://news.ycombinator.com/item?id=49823582 Points: 322 # Comments: 138 12 llama.cpp releases dev-tools 4d ago v0.5.0 Overview This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP… 26 NVIDIA Developer Blog official-blog 4d ago Manage Kubernetes Node Fleets with NodeWright Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents,... 18 arXiv — Machine Learning research 5d ago Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining arXiv:2609.25482v1 Announce Type: new Abstract: Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed… 14 arXiv — Machine Learning research 5d ago From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems arXiv:2609.26146v1 Announce Type: new Abstract: Risk monitoring in multi-agent systems is commonly built on a per-state primitive that scores each state independently and selects the top K. Under crowding, where many agents share the same fragility, this approach picks redundant… 13 arXiv — NLP / Computation & Language research 5d ago A Computational Approach to Measuring Semantic Change in Sanskrit Literature arXiv:2609.25012v1 Announce Type: new Abstract: Diachronic word embeddings have become the modern standard for tracking semantic change, yet they have been largely validated on modern, high-resource, and well-segmented languages. This paper tests whether the paradigm transfers… 7 arXiv — NLP / Computation & Language research 5d ago From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication arXiv:2609.25034v1 Announce Type: new Abstract: Central bank press conferences are not merely information releases --- they are structured narratives. We study whether the shape of sentiment within a statement, not just its average tone, carries policy-relevant signals.… 13 arXiv — NLP / Computation & Language research 5d ago ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch arXiv:2609.25081v1 Announce Type: new Abstract: We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at a total cost of about \$286 in… 38 arXiv — NLP / Computation & Language research 5d ago FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing arXiv:2609.25298v1 Announce Type: new Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally… 32 arXiv — NLP / Computation & Language research 5d ago Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference arXiv:2609.25537v1 Announce Type: new Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, energy consumption, and GPU memory demand as context length scales. Existing… 24 arXiv — NLP / Computation & Language research 5d ago MemoryAthena: Adaptive Routing over Latent and Generated Memories arXiv:2609.25853v1 Announce Type: new Abstract: Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated… 12 arXiv — NLP / Computation & Language research 5d ago BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval arXiv:2609.25859v1 Announce Type: new Abstract: Biomedical Entity Linking disambiguates mentions to entities in a knowledge base (KB), making it the cornerstone of information extraction pipelines. While embedding-based models are a popular approach for the task, they suffer… 18 arXiv — NLP / Computation & Language research 5d ago Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach arXiv:2609.26052v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to conventional Auto-Regressive (AR) Large Language Models (LLMs). By leveraging bidirectional attention and parallel decoding, dLLMs enable… 18 arXiv — NLP / Computation & Language research 5d ago CoVeR: Coverage-Based Routing of Verifier Calls in Agentic Retrieval arXiv:2609.26086v1 Announce Type: new Abstract: An agentic retrieval system issues a sequence of search queries and must decide, at each step, whether the evidence collected so far is enough to stop. Delegating that decision to an LLM verifier or a prompt judge makes stopping… 21 arXiv — NLP / Computation & Language research 5d ago Modality-Gated Deep Adapters: Adding a Modality to a Frozen Embedding Model with Exact Preservation arXiv:2609.26182v1 Announce Type: new Abstract: Multimodal embedding models are deployed at scale: retrieval indices, benchmark results, and behavioral audits all depend on the base model's exact outputs. Extending such a model to a new modality with existing parameter-efficient… 29 arXiv — NLP / Computation & Language research 5d ago HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing arXiv:2609.26368v1 Announce Type: new Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context… 38 arXiv — NLP / Computation & Language research 5d ago A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation arXiv:2609.26527v1 Announce Type: new Abstract: When two texts describe the same expression, standard metrics based on lexical overlap or whole-text similarity may fail to detect meaningful differences in how that expression is framed. We propose a framework to evaluate semiotic… 6 arXiv — NLP / Computation & Language research 5d ago Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models arXiv:2609.26610v1 Announce Type: new Abstract: Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural… 11 arXiv — NLP / Computation & Language research 5d ago Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives arXiv:2609.25007v1 Announce Type: cross Abstract: The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive… 18 arXiv — NLP / Computation & Language research 5d ago Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation arXiv:2609.25010v1 Announce Type: cross Abstract: Marketers increasingly use large language models (LLMs) as "synthetic personas" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples.… 16 arXiv — NLP / Computation & Language research 5d ago Efficient Iterative Retrieval with Heterogeneous Batching arXiv:2609.25405v1 Announce Type: cross Abstract: Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these… 23 OpenAI Python SDK releases dev-tools 5d ago v3.19.0 3.19.0 (2026-09-22) Features api: add GCP external storage support ( #3943 ) ( d12d60f ) api: add GPT-Rosalind research model ( #3940 ) ( d041d73 ) Bug Fixes _utils/_transform: propagate api_exclude in _async_transform_recursive ( #3324 ) ( 161ae65 ) api: handle omission markers… 7 r/MachineLearning community 5d ago LinearSolveBench: new benchmark for linear solvers [P] LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general numerical solvers for large sparse linear systems in C. The goal is to encourage algorithmic advances in numerical methods for solving linear systems of equations.… 27 llama.cpp releases dev-tools 6d ago b11099 ci : publish snapdragon builds in release workflow ( #29007 ) The snapdragon CI builds packages only to feed the QDC device tests, so Hexagon NPU binaries never reached the releases page. Build both targets in release.yml and attach them as release assets. Website:… 15 arXiv — Machine Learning research 6d ago LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling arXiv:2609.22117v1 Announce Type: new Abstract: Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility… 12 arXiv — Machine Learning research 6d ago From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning arXiv:2609.22155v1 Announce Type: new Abstract: Clinical decision support tools are most useful when accurate predictions are accompanied by understandable explanations. Rule-based models provide transparency, but rules derived directly from raw clinical measurements may miss… 12 arXiv — Machine Learning research 6d ago StepKV: Step-Aware KV Cache Compression for LLM Agents arXiv:2609.22158v1 Announce Type: new Abstract: Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this… 30 arXiv — Machine Learning research 6d ago WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction arXiv:2609.22191v1 Announce Type: new Abstract: Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines. Most models are evaluated using Average Precision (AP), which… 28 arXiv — Machine Learning research 6d ago List Counting Failures Are Not One Phenomenon arXiv:2609.22230v1 Announce Type: new Abstract: Counting the items in a bracketed list looks trivial, yet open-weight chat models often get it wrong. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different… 5 arXiv — Machine Learning research 6d ago Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing arXiv:2609.22244v1 Announce Type: new Abstract: This research proposes the Universal Observatory Graph (UOG), an AI-driven framework for distributed astronomical observation across the Solar System. The proposed architecture models autonomous observatories located at the Sun… 20 arXiv — Machine Learning research 6d ago CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents arXiv:2609.22247v1 Announce Type: new Abstract: Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of… 18 arXiv — Machine Learning research 6d ago EmbeddGAN: A Novel GAN Framework Using an Embedding Network and Gini Distance Correlation arXiv:2609.22508v1 Announce Type: new Abstract: Generative Adversarial Networks (GANs) have demonstrated strong performance in generating high-quality synthetic data. However, they are limited by no formal guarantees regarding convergence and the effectiveness of the learning… 12 arXiv — Machine Learning research 6d ago User-Level Handover Decision Making Based on Machine Learning Approaches arXiv:2609.22593v1 Announce Type: new Abstract: This letter covers a broad comparison of methods for classification and regression applications for a user-level handover decision making in scenarios with adverse propagation conditions involving buildings, coverage holes, and… 27 arXiv — Machine Learning research 6d ago D-IMPL: A Diffusion-based Solver for Parameterized BBOs arXiv:2609.22752v1 Announce Type: new Abstract: Diffusion models have demonstrated strong power in generative modeling tasks across multiple domains, exhibiting a remarkable capability of learning complex distributions from samples. In this paper, we leverage such capability to… 4 arXiv — Machine Learning research 6d ago Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting arXiv:2609.22820v1 Announce Type: new Abstract: Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress… 33 arXiv — Machine Learning research 6d ago Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories arXiv:2609.22867v1 Announce Type: new Abstract: Diffusion models generate a sample by traversing a denoising trajectory, a sequence of stochastic noise-reduction steps that transforms pure noise into a draw from a target distribution. At deployment time, additional computation… 25 arXiv — NLP / Computation & Language research 6d ago AdaMem: Adaptive Memory Token Allocation for Soft Compression in Retrieval-Augmented Generation arXiv:2609.22100v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves language models with retrieved evidence, but processing many long passages is costly and can introduce distracting information. Soft compression addresses this challenge by encoding… 27 arXiv — NLP / Computation & Language research 6d ago Balancing Reasoning and Hardware Constraints in RAG Pipelines for Ukrainian Multi-Domain Document Understanding arXiv:2609.22124v1 Announce Type: new Abstract: This paper describes the system submitted to the UNLP 2026 Shared Task on Multi-Domain Document Understanding. The challenge required extracting precise answers, document IDs, and page numbers from a diverse corpus of Ukrainian PDF… 18 arXiv — NLP / Computation & Language research 6d ago Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation arXiv:2609.22162v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves the factuality of large language models (LLMs) and vision-language models (VLMs) by grounding generation in external knowledge. However, most existing RAG frameworks assume a… 22 arXiv — NLP / Computation & Language research 6d ago A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification arXiv:2609.22212v1 Announce Type: new Abstract: Organizations in critical national infrastructure sectors must assess heterogeneous documents for sensitivity before routing or storage. Manual assessment is slow, inconsistent, and unscalable. Extending our prior… 5 arXiv — NLP / Computation & Language research 6d ago Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts arXiv:2609.22455v1 Announce Type: new Abstract: We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore… 38 arXiv — NLP / Computation & Language research 6d ago Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments arXiv:2609.22705v1 Announce Type: new Abstract: Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a… 24 arXiv — NLP / Computation & Language research 6d ago Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity arXiv:2609.23039v1 Announce Type: new Abstract: LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies… 19 arXiv — NLP / Computation & Language research 6d ago Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison arXiv:2609.23053v1 Announce Type: new Abstract: A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage… 24 arXiv — NLP / Computation & Language research 6d ago Bridging Static and Agentic RAG for Taiwanese Historical Question Answering arXiv:2609.23056v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed… 7 OpenAI Python SDK releases dev-tools 6d ago v3.17.0 3.17.0 (2026-09-22) Features api: add external storage configuration management ( #3909 ) ( 6332577 ) api: add safety case retrieval ( #3911 ) ( a87b938 ) api: add safety warning and deactivation webhook events ( #3908 ) ( 19f1f37 ) api: add session environment reset events (… 14 Vercel — AI dev-tools 6d ago Drives for Vercel Sandbox are now in public beta Drives for Vercel Sandbox are now available in public beta on Hobby, Pro, and Enterprise. A Drive is persistent storage that you mount as a directory in a Vercel Sandbox. It isn’t tied to a single sandbox, so you can reuse the same Drive across runs and different sandbox… 22 arXiv — Machine Learning research 7d ago FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift arXiv:2609.21057v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. Remedies built on classical Federated Averaging (FedAvg)… 29 arXiv — Machine Learning research 7d ago Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs arXiv:2609.21656v1 Announce Type: new Abstract: Recent Joint-Embedding Predictive Architectures (JEPAs) prevent representation collapse by constraining learned representations to follow a prescribed target distribution, such as an isotropic Gaussian or the uniform distribution… 9 Page 2 of 10 · 500 articles ← Newer Older →