News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow r/LocalLLaMA community 14d ago Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x From Raghu Sreeramaneni's memory tutorial at hot chips 2026. top line is normalised tflops for tpu v3 through r200, bottom line is hbm2e through hbm4, both log scale, so the distance between them is a lot wider than it looks. The three boxes down the right are the fixes people… 30 r/LocalLLaMA community 14d ago I have just moved from MacBook M5 pro 48 GB to RTX3090 Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s I think I can finally get rid of my Claude subscription, this is good enough for me. I am a software dev and I can get what I need from this set up… 10 r/LocalLLaMA community 15d ago internlm/Intern-S2 · Hugging Face from internlm: We introduce Intern-S2-397B , our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent… 6 r/LocalLLaMA community 15d ago This draft model is OP on 16 GB cards for Qwen 3.8 27b https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with the IQ3_XXS with 128k context and I saw it averaging about 60 tokens per second tg speed on the 16 GB RX 9070 XT. This… 16 r/LocalLLaMA community 16d ago Antirez Deepseek 4.1 flash gguf on HF Q2 is there and Q4 is uploading as I type. Has his github been updated yet? How do you run this? https://huggingface.co/antirez/deepseek-v4.1-flash-gguf/tree/main   submitted by   /u/Queasy_Asparagus69 [link]   [comments] 34 r/LocalLLaMA community 16d ago Learning/RSI through ngrams? Hey gang, im wondering if you in theory could use ngrams as seen with Qwen 3.8 Flash or DS4.1 in order to dynamically train the model? Normally the ngram embeddings behave similar to a lookup table of sorts. So instead of every token having to be represented only inside the main… 18 Hugging Face Daily Papers research 16d ago ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Abstract ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision. Generated by thinkingmachines/Inkling-Small As LLMs are increasingly used for… 18 OpenAI official-blog 17d ago Rapidly scaling online storage to serve over 1 billion ChatGPT users Learn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second. 4 r/LocalLLaMA community 17d ago What can you run on 8GB VRAM? Can you still do something with a 2050 or something like it? I mean for office work, loading embedding, reranking and chat models not at the same time but is anyone still using smaller models and have any good ones come out? I feel like small models are abandoned, I don’t care… 12 Hugging Face Daily Papers research 17d ago Generative Late-Interaction Embeddings For Visual Document Retrieval Abstract Generative Late-Interaction Embeddings compress visual document retrieval vectors by learning a small basis set that regenerates full embeddings on demand, improving accuracy under tight storage limits without retraining the encoder. Generated by… 28 Vercel — AI dev-tools 17d ago Vercel Sandbox now provides 64 GB of storage Every Vercel Sandbox now includes 64 GB of storage, up from 32 GB. This includes sandboxes created from a Vercel Managed Image or custom image, as well as those configured with the deprecated runtime property. The additional space provides more room for large repositories,… 8 arXiv — Machine Learning research 17d ago Conformal Calibration Transfer arXiv:2609.10737v1 Announce Type: new Abstract: Conformal prediction converts point predictions into set-valued predictions with coverage guarantees under exchangeability between calibration and deployment data. We study conformal calibration transfer, where this requirement… 27 arXiv — NLP / Computation & Language research 17d ago REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving arXiv:2609.11209v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and… 38 arXiv — Machine Learning research 17d ago DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis arXiv:2609.11504v1 Announce Type: new Abstract: A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration… 30 arXiv — Machine Learning research 17d ago Learnware and AI Model Management System arXiv:2609.11656v1 Announce Type: new Abstract: The transition from file storage to database management systems transformed stored data into managed resources. AI now faces an analogous transition from AI model storage to AI model management. Existing model pools essentially… 20 arXiv — NLP / Computation & Language research 17d ago When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text arXiv:2609.11067v1 Announce Type: new Abstract: Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise… 16 arXiv — NLP / Computation & Language research 17d ago A Fragility Spectrum for Recursive Language-Model Training arXiv:2609.11149v1 Announce Type: new Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols… 24 arXiv — NLP / Computation & Language research 17d ago Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech arXiv:2609.11545v1 Announce Type: new Abstract: Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker… 37 arXiv — NLP / Computation & Language research 17d ago A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings arXiv:2609.11620v1 Announce Type: new Abstract: High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces… 22 arXiv — NLP / Computation & Language research 17d ago Negative Self-Distillation: Learning to Reason by Avoiding Flaws arXiv:2609.11699v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions.… 6 arXiv — NLP / Computation & Language research 17d ago RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can… 25 arXiv — NLP / Computation & Language research 17d ago Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model arXiv:2609.11870v1 Announce Type: new Abstract: A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked… 28 arXiv — NLP / Computation & Language research 17d ago VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents arXiv:2609.11390v1 Announce Type: cross Abstract: State-of-the-art retrieval-augmented generation (RAG) methods exploit document structures to acquire sufficient evidence, but often incur substantial token costs. To reduce structural-context tokens without compromising high RAG… 10 arXiv — NLP / Computation & Language research 17d ago SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control arXiv:2609.11752v1 Announce Type: cross Abstract: For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation… 10 arXiv — NLP / Computation & Language research 17d ago CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models arXiv:2509.22360v2 Announce Type: replace Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however… 20 arXiv — NLP / Computation & Language research 17d ago Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection arXiv:2510.15685v2 Announce Type: replace Abstract: This paper investigates the use of an LLM to generate auxiliary background context for social media posts, and explores four methods to incorporate this context into the input of an SBERT-based Hate Speech Detection (HSD)… 28 r/LocalLLaMA community 17d ago Granular diff versioning for agent editing Had some ideas about version control: Attribute, down to individual words: word created manually or with ai. if ai, which sources cited (scoped to paragraphs). also: full agent trace that led to the edit (all tool calls + metadata). also: any multi-agent transactions at the… 10 The Information — AI news-outlet 17d ago Personal AI App Instinct Faces Compute Crunch That Could Lead to New Funding The year-old startup behind Instinct, a personal AI assistant that’s caught fire with Silicon Valley insiders, is seeking more computing power, leveraging its early buzz as it faces new competition from giant Meta Platforms. Over the last few weeks, Instinct, which launched to a… 20 Hugging Face Daily Papers research 17d ago From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution Abstract Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data. Generated by… 24 Hugging Face Daily Papers research 18d ago The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding Abstract A non-invasive brain decoding approach maps MEG responses to semantic embeddings to reconstruct sentence-level text without requiring word-level alignment. Generated by thinkingmachines/Inkling-Small Non-invasive speech decoding remains constrained by the low… 34 arXiv — Machine Learning research 18d ago When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic arXiv:2609.05508v1 Announce Type: new Abstract: Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. Its headline result is that performance improves as options are added. We explain that result, with theory and… 37 arXiv — Machine Learning research 18d ago Analysis of Respiratory Sinus Arrhythmia with Neural Networks arXiv:2609.05698v1 Announce Type: new Abstract: The paper introduces a neural network-based approach for analyzing ECG signals to estimate respiratory rate by leveraging the phe- nomenon of Respiratory Sinus Arrhythmia (RSA). Our method employs a deep learning model trained to… 5 arXiv — Machine Learning research 18d ago A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials arXiv:2609.05877v1 Announce Type: new Abstract: Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. The better choice can depend on how much data… 29 arXiv — Machine Learning research 18d ago CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning arXiv:2609.05884v1 Announce Type: new Abstract: Conventional federated learning relies on parameter averaging, which forces clients to be doubly homogeneous: all must run an identical architecture, and accuracy degrades when local data are non-IID. Decentralized federated… 38 arXiv — Machine Learning research 18d ago Memory in Deep Time-Series Models arXiv:2609.06006v1 Announce Type: new Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state-space models, retrieval-augmented predictors, foundation models, and tool-using… 9 arXiv — Machine Learning research 18d ago ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs arXiv:2609.06072v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow… 12 arXiv — Machine Learning research 18d ago All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs arXiv:2609.06161v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. Weight binarization is a promising solution, but existing binarization-based… 16 arXiv — Machine Learning research 18d ago Decision-Aware Suffix Prediction and Reasoning of Business Processes arXiv:2609.06169v1 Announce Type: new Abstract: Suffix prediction forecasts the remaining sequence of events of a running case until completion. Most approaches rely on neural networks trained on event logs, which, on average, perform well but struggle with short prefixes or… 6 arXiv — Machine Learning research 18d ago Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes arXiv:2609.06294v1 Announce Type: new Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from… 35 arXiv — Machine Learning research 18d ago Bi-HYCO: Bi-Objective Cooperative Learning for PDE Parameter Identification under Fragmented Observations arXiv:2609.06511v1 Announce Type: new Abstract: Physical and synthetic models may describe complementary aspects of the same PDE-governed system while receiving different, possibly fragmented, observations. We propose Bi-Objective HYCO (Bi-HYCO), a cooperative framework that… 15 arXiv — Machine Learning research 18d ago Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity arXiv:2609.06557v1 Announce Type: new Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. However, recent studies suggest that LLMs are… 30 arXiv — NLP / Computation & Language research 18d ago Osprey: Target-agnostic Pre-training Makes Stronger Drafters in Speculative Decoding arXiv:2609.09338v1 Announce Type: new Abstract: Speculative decoding is critical for accelerating LLM inference. However, the speedup is fragile: drafters are typically trained against a narrow distribution for a single target model, and their acceptance rate collapses under… 12 arXiv — NLP / Computation & Language research 18d ago Do LLMs Make More Mistakes If They Do Not Believe the Input Data? arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context… 36 arXiv — NLP / Computation & Language research 18d ago Reproducing Omitted Temporal Expressions in Japanese News for Retrieval-Augmented Applications arXiv:2609.09569v1 Announce Type: new Abstract: News articles often contain omitted temporal expressions, such as day-only or month-only mentions, which must be interpreted with reference to the publication date. When such articles are indexed or processed as standalone text in… 33 arXiv — NLP / Computation & Language research 18d ago Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services arXiv:2609.09889v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call… 37 arXiv — NLP / Computation & Language research 18d ago Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records arXiv:2609.09984v1 Announce Type: new Abstract: Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the… 37 arXiv — NLP / Computation & Language research 18d ago Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning arXiv:2609.10142v1 Announce Type: new Abstract: Large language models remain fragile against malicious fine-tuning, motivating training-time defenses against harmful persona drift. Preventative Steering injects undesirable-trait persona vectors during fine-tuning and removes… 36 arXiv — NLP / Computation & Language research 18d ago From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora arXiv:2609.10155v1 Announce Type: new Abstract: We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning.… 18 arXiv — NLP / Computation & Language research 18d ago The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs arXiv:2609.10237v1 Announce Type: new Abstract: A graph retrieval-augmented generation pipeline chooses which triples to put in the prompt, a syntax to write them in, an order to write them in, and a sentence telling the model what to do with them. We vary all four over six… 12 arXiv — NLP / Computation & Language research 18d ago KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints arXiv:2609.10266v1 Announce Type: new Abstract: LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different set of… 29 Page 5 of 10 · 500 articles ← Newer Older →