News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs arXiv:2607.09121v1 Announce Type: new Abstract: In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on their reports as well as data and documents describing macroeconomic situation… 17 arXiv — NLP / Computation & Language research 1mo ago DKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured Data arXiv:2607.09348v1 Announce Type: new Abstract: Causal discovery from unstructured data is a challenging yet underexplored task in high-expertise domains such as healthcare, finance, and education. Existing methods typically leverage the general knowledge of large language… 26 arXiv — NLP / Computation & Language research 1mo ago VTaMo: Video-Text Alignment Model for Sign Language Translation arXiv:2607.09126v1 Announce Type: cross Abstract: Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation… 9 arXiv — NLP / Computation & Language research 1mo ago Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs arXiv:2508.14817v2 Announce Type: replace Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Methods: We defined three… 27 arXiv — NLP / Computation & Language research 1mo ago Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting arXiv:2511.00651v2 Announce Type: replace-cross Abstract: Telecom networks are rapidly growing in scale and complexity, making effective management, operation, and optimization increasingly challenging. Although Artificial Intelligence (AI) has been applied to many telecom… 25 r/LocalLLaMA community 1mo ago llama.cpp Agentic Workflows Ctx Checkpoints Fix b9978 Claude in one sentence what does this fix llama.cpp b9978 fixes a checkpoint bug that hit agentic workloads hardest: every agent turn created a new checkpoint (bypassing min-step spacing), collapsing the coverage window so that context rewinds — common in tool-calling… 35 Simon Willison community 1mo ago sqlite-utils 4.1.1 Release: sqlite-utils 4.1.1 Mainly a fix for an edge case that regular Claude chat spotted while experimenting with the 4.1 release to answer a question about ON DELETE. table.transform() now raises a TransactionError if called while a transaction is open with PRAGMA… 6 r/LocalLLaMA community 1mo ago Benchmark - 4x 5060 Ti (64GB VRAM) (P2P) - Qwen3.6 27B (INT8 /w bf16 kv cache) @ 8 concurrency with SGLang. SGLang seems to handle higher concurrency better with this setup I recently posted some posts with VLLM showing issues with TTFT and concurrency with 4x 5060 ti's. Wanted to share this benchmark to provide what worked for me so other people that are planning to go the 4x 5060 ti route aren't discouraged. Benchmark Results ============ Serving… 38 r/MachineLearning community 1mo ago Context and average best linear mappings [D] The context (in a border sense) viewpoint of neural networks is not thought about too much but it leads to a simple best average linear mapping viewpoint of a layer. https://archive.org/details/a-context-based-view-of-deep-neural-networks   submitted by  … 18 r/MachineLearning community 1mo ago VultronRetriever family of models released on HuggingFace![R] Thrilled to announce the VultronRetriever family of models, which were announced during Raise Summit Paris and demonstrated running Q&A and embedding documents on the iPhone, fully offline! 📱 Some highlights from the VultronRetriever model family: 🥇 Each model ranks #1 in its… 9 Hugging Face Daily Papers research 1mo ago LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models Abstract LongE2V enables high-quality video recovery from sparse event streams by leveraging pre-trained video diffusion priors and addressing temporal stability and frame interpolation challenges. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recovering high-quality video from… 12 arXiv — Machine Learning research 1mo ago Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE arXiv:2607.07740v1 Announce Type: new Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order… 37 arXiv — Machine Learning research 1mo ago NFTR: From Provable Mode-Averaging to Geodesic Subgoal Selection in Offline Goal-Conditioned RL arXiv:2607.07855v1 Announce Type: new Abstract: Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone. This rule has two coupled failure modes. Optimistic bias treats lucky stochastic outcomes as… 15 arXiv — Machine Learning research 1mo ago Distributed Sketching on Data Partitions for OLS Regression arXiv:2607.07888v1 Announce Type: new Abstract: This paper studies distributed sketching for ordinary least squares (OLS) regression, an approach that distributes small sketches of a large data set over multiple machines to separately construct OLS estimators and average them.… 27 arXiv — Machine Learning research 1mo ago Contrastive Order Learning: A General Framework for Ordinal Regression arXiv:2607.08109v1 Announce Type: new Abstract: We propose contrastive order learning (ConOrd), a contrastive learning framework for ordinal regression that integrates the strengths of contrastive learning and order learning. While contrastive learning effectively leverages all… 13 arXiv — Machine Learning research 1mo ago Eigenvalue Calibration for Semantic Embeddings of Large Language Models arXiv:2607.08377v1 Announce Type: new Abstract: Uncertainty quantification is central to the reliable deployment of large language models (LLMs), and eigenvalues of semantic embeddings have recently emerged as a key tool in state-of-the-art methods. However, conventional… 9 arXiv — Machine Learning research 1mo ago MatBind: A Shared Embedding Space for Multimodal Materials Characterization arXiv:2607.08470v1 Announce Type: new Abstract: Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a different facet… 31 arXiv — Machine Learning research 1mo ago BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression arXiv:2607.08643v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly constrained by memory capacity, weight bandwidth, and checkpoint storage during deployment. Existing low-bit compression methods mainly follow two directions. Scalar or group-wise… 21 arXiv — Machine Learning research 1mo ago Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph arXiv:2607.08746v1 Announce Type: new Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This graph… 11 arXiv — Machine Learning research 1mo ago Context Graphs for Proactive Enterprise Agents arXiv:2607.07721v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wait for a human query before acting. This paper argues that genuine enterprise… 33 arXiv — NLP / Computation & Language research 1mo ago A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding arXiv:2607.07974v1 Announce Type: new Abstract: Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there still exist challenges for detecting out-of-scope (OOS) intents. (i) The traditional methods… 35 arXiv — NLP / Computation & Language research 1mo ago COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation arXiv:2607.08117v1 Announce Type: new Abstract: Contextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA (Contextualized ASR Leveraging Biasing… 7 arXiv — NLP / Computation & Language research 1mo ago ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation arXiv:2502.15543v4 Announce Type: replace Abstract: Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence. However, they remain susceptible to unfaithful generation, where outputs… 9 arXiv — NLP / Computation & Language research 1mo ago Less Is More: Reducing Token Counts Without Compromising Performance arXiv:2506.15138v2 Announce Type: replace Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost. Although longer, multi-word tokens can reduce fertility, naively… 11 arXiv — NLP / Computation & Language research 1mo ago Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning arXiv:2602.23440v4 Announce Type: replace Abstract: Reinforcement learning has emerged as an effective paradigm for training large language models to interleave reasoning with search engine calls. However, existing approaches face a fundamental credit assignment problem: methods… 35 arXiv — NLP / Computation & Language research 1mo ago The Proxy Presumption: From Semantic Embeddings to Valid Social Measures arXiv:2605.07409v2 Announce Type: replace Abstract: Natural Language Processing is rapidly evolving into a primary instrument for Computational Social Science, with researchers increasingly using embeddings to measure latent constructs such as novelty, creativity, and bias.… 17 arXiv — NLP / Computation & Language research 1mo ago Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions arXiv:2604.12138v3 Announce Type: replace-cross Abstract: This position paper argues that Retrieval-Augmented Generation (RAG) systems exhibit a factual bias-optimizing for epistemic uncertainty reduction while ignoring the aleatoric uncertainty inherent in opinion-rich content.… 20 arXiv — NLP / Computation & Language research 1mo ago DeepTutor: Towards Agentic Personalized Tutoring arXiv:2604.26962v3 Announce Type: replace-cross Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-training knowledge and lack adaptation to individual learners, while existing RAG… 8 r/MachineLearning community 1mo ago Hyperparameter tuning approach question [R] I am doing some work with cell type classification, where I have 4.3 million cells and 512 features (condensed embeddings from the encoder of a transformer). The broader goal is to implement a contextual bandit for augmenting the training set of the dataset, as it is currently… 34 r/LocalLLaMA community 1mo ago If You Already Pay for an LLM Service, Running Local Embeddings and Rerankers Feels More Useful Than Running Local LLMs https://preview.redd.it/v0xtn3jdu9ch1.png?width=2047&format=png&auto=webp&s=628a6a541fe5f097d0f771ae0ba3b7f44126198f https://preview.redd.it/vjxiucsdu9ch1.png?width=2047&format=png&auto=webp&s=74f7a18a5a30276e206e2bfb5a0c529826ce86e4 This post was originally written in Korean,… 33 r/LocalLLaMA community 1mo ago Now brothers we know why we are so fucked up Samsung chip division's single-year profits beat its past 40 years of profits, combined, due to increased memory and storage prices — Samsung passes Nvidia to become most profitable company in the world, notches 19x quarterly increase in profit… 33 arXiv — Machine Learning research 1mo ago STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning arXiv:2607.06629v1 Announce Type: new Abstract: Brain age -- the age inferred from a physiological recording -- is an emerging biomarker whose deviation from chronological age tracks neurological and psychiatric burden, and EEG is an attractive substrate for it because it is… 10 arXiv — Machine Learning research 1mo ago At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics arXiv:2607.06639v1 Announce Type: new Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank at the grokking transition overstates the converged value by 3-5x on an MLP, and… 18 arXiv — Machine Learning research 1mo ago Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference arXiv:2607.07144v1 Announce Type: new Abstract: The key-value (KV) cache dominates the memory cost of long-context autoregressive inference, and a growing body of work compresses it through quantization, eviction, or offloading. We study a complementary question: once a… 36 arXiv — Machine Learning research 1mo ago ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation arXiv:2607.07640v1 Announce Type: new Abstract: Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within the corrupted input sequence. This reliance can be limiting in real-world… 5 arXiv — NLP / Computation & Language research 1mo ago Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering arXiv:2607.06641v1 Announce Type: new Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance. Retrieval-Augmented… 15 arXiv — NLP / Computation & Language research 1mo ago Riemannian Geometry for Pre-trained Language Model Embeddings arXiv:2607.07047v1 Announce Type: new Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token… 21 arXiv — NLP / Computation & Language research 1mo ago Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation arXiv:2607.07050v1 Announce Type: new Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool… 15 arXiv — NLP / Computation & Language research 1mo ago From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings arXiv:2607.07141v1 Announce Type: new Abstract: Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement… 28 arXiv — NLP / Computation & Language research 1mo ago Evaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations arXiv:2607.07302v1 Announce Type: new Abstract: This paper reports an empirical study evaluating the relevance of several RAG metrics. The experiment is based on a question-answering dataset created by human annotators from business data. The generated responses and retrieved… 4 arXiv — NLP / Computation & Language research 1mo ago Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning arXiv:2511.13726v2 Announce Type: replace Abstract: We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the final semantic representation by running multiple forward passes of the text embedding… 16 arXiv — NLP / Computation & Language research 1mo ago Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs arXiv:2601.21476v2 Announce Type: replace Abstract: On-policy reinforcement learning (RL) for language model post-training suffers from a fundamental tension: as training progresses, policy entropy collapses and sampling diversity diminishes, causing the model to ``forget'' its… 19 arXiv — NLP / Computation & Language research 1mo ago Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval arXiv:2604.18360v3 Announce Type: replace-cross Abstract: Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially… 6 Hugging Face Daily Papers research 1mo ago JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications Abstract A large-scale industrial platform leveraging LLMs and VLMs addresses key challenges in structured item knowledge production for e-commerce, achieving high precision and throughput across billions of products. Generated by Qwen/Qwen2.5-Coder-32B-Instruct JD.com, one of… 17 Latent.Space news-outlet 1mo ago Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO 2 years after our first coverage, we return with Modal's other cofounder to explore why Agent Experience is working now, and everything they have learned building the new agent cloud. 16 llama.cpp releases dev-tools 1mo ago b9931 opencl: ragged-tile MoE prefill FP16 GEMM optimization (skip padded expert tiles) ( #25433 ) opencl: ragged-tile MoE prefill GEMM (skip padded expert tiles) The MoE prefill GEMM groups tokens into TILESIZE_N=32 per-expert tiles; at low tokens-per-expert most tiles are mostly… 34 Hugging Face official-blog 1mo ago Data for Agents Back to Articles a]:hidden"> Data for Agents Enterprise + Article Published July 8, 2026 Upvote 2 Will Jennings WillJenningsDC nvidia Jane Polak Scowcroft jscowcroft nvidia Annie Surla asurla1998 nvidia Yev Meyer nv-3mei nvidia Rebecca Kao rebeccak-nv nvidia Leanna Chraghchian… 17 r/MachineLearning community 1mo ago First time ARR users - some questions [D] We submitted our first paper to ARR, intending to commit to IJCNLP-AACL. Area: Multilingualism and Cross-Lingual NLP Scores: (3,4) (2.5,3) (3,3) - average 2.83 for reviews, 3.33 for confidence 3 for soundness on all, 4 for reproducibility, and 2,3,3 for excitement. The reviewer… 5 Hugging Face Daily Papers research 1mo ago Attending to Multimodal Generation One Token at a Time Abstract Multimodal large language models exhibit distinct attention patterns during generation, with attention to visual and textual modalities shifting based on semantic requirements, and these patterns can be leveraged to improve task performance through targeted… 5 r/MachineLearning community 1mo ago DINOv2 way worse than SigLIP in k-NN. Is this expected? [R] Doing a bachelor thesis on fine-grained car classification (telling apart VW Golf generations from listing photos). Simple setup: frozen encoder → embeddings → weighted k-NN. On my small dataset (175 train / 132 test): SigLIP2 SO400M: ~92% CLIP ViT-L: ~59% DINOv2 Giant: ~41% I… 27 Page 10 of 10 · 500 articles ← Newer