News / #rag Tag Rag 500 articles archived under #rag · RSS Sign in to follow arXiv — Machine Learning research 7d ago DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data arXiv:2608.05375v1 Announce Type: cross Abstract: Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines… 14 arXiv — NLP / Computation & Language research 7d ago Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability arXiv:2608.05153v1 Announce Type: new Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness analysis that holds the retrieval architecture fixed and varies three orthogonal… 27 arXiv — NLP / Computation & Language research 7d ago CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences arXiv:2608.05167v1 Announce Type: new Abstract: Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rely on contextual co-occurrence, degrading performance on rare and… 29 arXiv — NLP / Computation & Language research 7d ago Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation arXiv:2608.05576v1 Announce Type: new Abstract: When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however,… 30 arXiv — NLP / Computation & Language research 7d ago Sparse Mutual Information Graph Averaging for Improving Random Indexing Embeddings arXiv:2608.05724v1 Announce Type: new Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics. This paper studies Random Indexing (RI)… 8 arXiv — NLP / Computation & Language research 7d ago Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation arXiv:2608.05785v1 Announce Type: new Abstract: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow… 5 arXiv — NLP / Computation & Language research 7d ago Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing arXiv:2608.05857v1 Announce Type: new Abstract: Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geometric properties, complicating model migration and… 4 arXiv — NLP / Computation & Language research 7d ago Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers arXiv:2608.06111v1 Announce Type: new Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}. We introduce \textbf{S}yntax-\textbf{i}nformed \textbf{P}ositional \textbf{E}mbeddings… 15 arXiv — NLP / Computation & Language research 7d ago NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering arXiv:2608.06292v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning… 9 arXiv — NLP / Computation & Language research 7d ago Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services arXiv:2608.05159v1 Announce Type: cross Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial… 36 arXiv — NLP / Computation & Language research 7d ago Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations arXiv:2608.06305v1 Announce Type: cross Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents --… 5 arXiv — NLP / Computation & Language research 7d ago Building Open-Retrieval Conversational Question Answering Systems by Generating Synthetic Data and Decontextualizing User Questions arXiv:2507.04884v2 Announce Type: replace Abstract: We consider open-retrieval conversational question answering (OR-CONVQA), an extension of question answering where system responses need to be (i) aware of dialog history and (ii) grounded in documents (or document fragments)… 4 Hugging Face Daily Papers research 7d ago Lossless Tensor Compression as Program Synthesis Abstract Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-specific compressors rely on fixed… 4 Hugging Face Daily Papers research 8d ago GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Abstract Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not… 17 arXiv — Machine Learning research 8d ago SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors arXiv:2608.04060v1 Announce Type: new Abstract: Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps. We introduce SJEPA, a reconstruction-free JEPA… 38 arXiv — Machine Learning research 8d ago Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning arXiv:2608.04305v1 Announce Type: new Abstract: Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates,… 15 arXiv — Machine Learning research 8d ago A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction arXiv:2608.04710v1 Announce Type: new Abstract: Integrated Sensing and Communication (ISAC) combines sensing and communication to efficiently utilize wireless resources and is emerging as a key paradigm for next-generation wireless networks. By leveraging the wide bandwidth,… 5 arXiv — NLP / Computation & Language research 8d ago Leveraging Machine Learning to Gain Insights on Quantum Thermodynamic Entropy arXiv:2305.06177v1 Announce Type: cross Abstract: We present a thermodynamic analysis of a quantum engine that uses a single quantum particle as its working fluid, inspired by Szilard's classical single-particle engine. Our design is modeled after the classically-chaotic Szilard… 31 arXiv — NLP / Computation & Language research 8d ago Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks arXiv:2608.04286v1 Announce Type: new Abstract: Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG). However, these… 23 arXiv — NLP / Computation & Language research 8d ago Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings arXiv:2608.04311v1 Announce Type: new Abstract: Translating wordplay across languages has long challenged both professional translators and machine translation systems. We investigate three approaches to translating puns from English to French by combining large language models… 25 arXiv — NLP / Computation & Language research 8d ago D$^2$F-ReAG: Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation arXiv:2608.04444v1 Announce Type: new Abstract: Large language models (LLMs) often generate inaccurate answers due to their reliance on static internal knowledge. Retrieval-augmented generation (RAG) addresses this limitation by integrating external knowledge and excelling at… 29 arXiv — NLP / Computation & Language research 8d ago The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads arXiv:2608.04570v1 Announce Type: new Abstract: Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what… 16 arXiv — NLP / Computation & Language research 8d ago InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval arXiv:2608.04761v1 Announce Type: new Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight… 18 arXiv — NLP / Computation & Language research 8d ago DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models arXiv:2608.04477v1 Announce Type: cross Abstract: Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, SentinelLMs, TextObfuscator, and DPNR---mitigate this risk by transforming prompt… 25 arXiv — NLP / Computation & Language research 8d ago CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding arXiv:2608.04515v1 Announce Type: cross Abstract: Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture… 8 Hugging Face Daily Papers research 8d ago BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Abstract Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution… 4 Hugging Face Daily Papers research 8d ago The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads Abstract Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench,… 29 Hugging Face Daily Papers research 8d ago ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels Abstract Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this… 10 Hugging Face Daily Papers research 8d ago Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation Abstract MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is… 36 r/LocalLLaMA community 8d ago I updated my localy run benchmark with DeepSeek V4 Flash 0731 It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90 t/s gen (average). I tried different sampling params, you can check the detail. It's very efficient while scoring the best yet. Too bad… 12 Hugging Face Daily Papers research 9d ago When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills Abstract Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges… 29 arXiv — Machine Learning research 9d ago Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage arXiv:2608.02629v1 Announce Type: new Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develop a new multimodal auto-regressive transformer surrogate to model these operations under… 24 arXiv — Machine Learning research 9d ago GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks… 8 arXiv — NLP / Computation & Language research 9d ago ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. Aligned with this structure,… 19 arXiv — Machine Learning research 9d ago Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation arXiv:2608.03148v1 Announce Type: new Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging because retrieved context increases computation and memory. A direct way to reduce this… 26 arXiv — Machine Learning research 9d ago The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics arXiv:2608.03291v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process. Existing approaches that leverage verbalized CoTs to monitor reasoning… 6 arXiv — Machine Learning research 9d ago Tight Worst-Case Bounds for the Smallest Eigenvalue of ReLU NTK Gram Matrices arXiv:2608.03368v1 Announce Type: new Abstract: For $n$ unit vectors $x_1,\ldots,x_n \in \mathbb{R}^d$, we study the continuous ReLU derivative Gram matrix $H$, whose entries are obtained by averaging pairwise gated inner products over a standard Gaussian direction. Writing $… 10 arXiv — NLP / Computation & Language research 9d ago JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single… 36 arXiv — NLP / Computation & Language research 9d ago ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection… 21 arXiv — NLP / Computation & Language research 9d ago MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow… 15 arXiv — NLP / Computation & Language research 9d ago Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational consumption due to overly large vocabularies and embedding layers. Although… 20 arXiv — NLP / Computation & Language research 9d ago Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We… 11 arXiv — NLP / Computation & Language research 9d ago ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels arXiv:2608.03507v1 Announce Type: new Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve… 23 arXiv — NLP / Computation & Language research 9d ago Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual… 9 arXiv — NLP / Computation & Language research 9d ago SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG arXiv:2608.03860v1 Announce Type: new Abstract: We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three… 38 Hugging Face Daily Papers research 9d ago AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Abstract Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces… 32 r/LocalLLaMA community 9d ago Kimi K3 full model running on 16x GB10 cluster at 20+tps Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the… 13 MIT News — AI research 9d ago Solving the solvent problem By focusing on electrolytes, MIT scientists are making sodium-metal batteries a more practical energy storage option. 13 r/MachineLearning community 9d ago NeurIPS 2026 post-rebuttal score distribution poll [D] As the title suggests, because there's no data on Papercopilot yet, and people have been talking about the scores being lower in general than last year, I thought it could be interesting to survey the average score distribution after the rebuttal phase (not considering… 26 Hacker News — AI on Front Page community 9d ago AI-Generated Images Discourage Me from Reading Your Blog Article URL: https://nelson.cloud/ai-generated-images-discourage-me-from-reading-your-blog/ Comments URL: https://news.ycombinator.com/item?id=49167113 Points: 387 # Comments: 229 21 Page 3 of 10 · 500 articles ← Newer Older →