News / #paper Tag Research papers 500 articles archived under #paper · RSS Sign in to follow arXiv — NLP / Computation & Language research 2d ago Mapping and Measuring the Behavioral Evolution of Large Language Models arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using… 27 arXiv — NLP / Computation & Language research 2d ago ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization arXiv:2608.11045v1 Announce Type: cross Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization… 30 arXiv — Machine Learning research 2d ago Efficient Hypergradient Descent for Inverse Reinforcement Learning arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A natural approach is to formulate IRL as a bilevel optimization… 38 arXiv — Machine Learning research 2d ago Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study arXiv:2608.11054v1 Announce Type: new Abstract: Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in… 25 arXiv — Machine Learning research 2d ago Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training arXiv:2608.11061v1 Announce Type: new Abstract: Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory,… 38 arXiv — Machine Learning research 2d ago Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives arXiv:2608.11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and… 35 arXiv — Machine Learning research 2d ago Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting arXiv:2608.11114v1 Announce Type: new Abstract: Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existing approaches often face a fundamental trade-off between distributional flexibility and… 31 arXiv — Machine Learning research 2d ago A Recommendation System Approach for Interference-Robust Sensor Subset Selection arXiv:2608.11143v1 Announce Type: new Abstract: This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive… 38 arXiv — Machine Learning research 2d ago DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains arXiv:2608.11154v1 Announce Type: new Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth,… 35 arXiv — Machine Learning research 2d ago Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension arXiv:2608.11162v1 Announce Type: new Abstract: The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength… 14 arXiv — NLP / Computation & Language research 2d ago Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders arXiv:2608.11197v1 Announce Type: cross Abstract: Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We… 22 arXiv — Machine Learning research 2d ago Quantifying the noise sensitivity of the Wasserstein metric for images arXiv:2510.01015v3 Announce Type: cross Abstract: Wasserstein metrics are increasingly adopted as similarity scores for images. We consider the sensitivity of Wasserstein metrics with respect to pixel-wise additive noise when the images are treated as discrete measures on the… 38 arXiv — Machine Learning research 2d ago Optimized Sequential Testing for Binary Ensemble Classifiers arXiv:2606.15237v1 Announce Type: cross Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote. A classic example is random forests, which combine the predictions of decision trees. Ensembles that use more… 8 arXiv — NLP / Computation & Language research 2d ago Divergent Response Modes in Frontier Language Models Under Steering Pressure arXiv:2608.06578v1 Announce Type: cross Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study… 35 arXiv — Machine Learning research 2d ago HyperShape: Hyperelasticity Across Diverse Shapes arXiv:2608.09938v1 Announce Type: cross Abstract: Hyperelastic deformations are highly sensitive to domain geometry and boundary conditions, making generalization across both a critical capability for neural operators applied to these problems. However, existing benchmarks for… 34 arXiv — NLP / Computation & Language research 2d ago When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning arXiv:2608.09942v1 Announce Type: new Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of the H_dp bandwidth bound (Chen et al., 2024): although the formal bound binds… 21 arXiv — Machine Learning research 2d ago EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems arXiv:2608.09943v1 Announce Type: cross Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems (e.g., heat waves, parasitism, predator attacks). Animal… 35 arXiv — Machine Learning research 2d ago An adaptive and evolvable deep reinforcement learning framework for weather prediction arXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the forecasting problem as one of coordination. Here we present Feitian Adaptive… 28 arXiv — Machine Learning research 2d ago AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting arXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to physics-based NWP in forecasting tropical cyclone (TC) tracks, they dramatically… 24 arXiv — Machine Learning research 2d ago Projected climate memory and inherited warm-tail risk in accelerated European summer warming arXiv:2608.09966v1 Announce Type: cross Abstract: European summer warming reflects interactions among background change, persistent ocean--land--circulation states, and same-season variability. We develop an empirical reduced-dynamics framework that decomposes regional summer… 16 arXiv — Machine Learning research 2d ago SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning arXiv:2608.09967v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel… 35 arXiv — Machine Learning research 2d ago Do AI weather models miss extremes? arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European… 30 arXiv — Machine Learning research 2d ago MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis arXiv:2608.09986v1 Announce Type: cross Abstract: Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications frequently encounter incomplete or corrupted modalities, posing a critical challenge. Although… 29 arXiv — Machine Learning research 2d ago Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy arXiv:2608.09992v1 Announce Type: cross Abstract: Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the synthesis of samples with explicit constraints on semantic content, structural properties,… 6 arXiv — Machine Learning research 2d ago Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection arXiv:2608.09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. Deep learning (DL) models show strong potential for medical image analysis; however, as their… 25 arXiv — Machine Learning research 2d ago Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint arXiv:2608.09998v1 Announce Type: cross Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despite their benefits, growing attention has been directed toward their environmental… 8 arXiv — NLP / Computation & Language research 2d ago Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit… 36 arXiv — Machine Learning research 2d ago HIPNO: Symmetry-Aware Physics-Informed Neural Operators for Noninvasive Hemodynamic Inference arXiv:2608.10011v1 Announce Type: cross Abstract: Continuous hemodynamic monitoring guides treatment decisions in surgery and intensive care. However, gold-standard signals are only measured in severe cases due to risks associated with invasive measurement. In this work, we… 30 arXiv — Machine Learning research 2d ago Deep Learning-Based Statistical Downscaling of Sea Surface Temperature Using a Residual Corrective Neural Network arXiv:2608.10022v1 Announce Type: cross Abstract: The large-scale oceanic and atmospheric forecasts provided by global climate models typically lack sufficient resolution to accurately capture the response of the coastal ocean to atmospheric forcing and coastal circulation that… 37 arXiv — NLP / Computation & Language research 2d ago LLM Agents Factory: Retrieval of Domain-Specific LLM Agents arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the… 24 arXiv — NLP / Computation & Language research 2d ago Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025 arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement… 12 arXiv — NLP / Computation & Language research 2d ago Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural… 29 arXiv — NLP / Computation & Language research 2d ago The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a… 30 arXiv — NLP / Computation & Language research 2d ago Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitation by introducing absolute coordinates, relative distances, or… 11 arXiv — NLP / Computation & Language research 2d ago PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs,… 27 arXiv — NLP / Computation & Language research 2d ago The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding arXiv:2608.10137v1 Announce Type: new Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid masking distorts the model's underlying probability… 16 arXiv — NLP / Computation & Language research 2d ago Multimodal Item Parameter Estimation using Simulated Response Probabilitie arXiv:2608.10154v1 Announce Type: new Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to… 19 arXiv — NLP / Computation & Language research 2d ago Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems arXiv:2608.10216v1 Announce Type: new Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question:… 32 arXiv — NLP / Computation & Language research 2d ago Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usually treated as an obstacle to interpretation. We show it is functional. A… 38 arXiv — NLP / Computation & Language research 2d ago TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent arXiv:2608.10258v1 Announce Type: new Abstract: Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after… 29 arXiv — NLP / Computation & Language research 2d ago Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies arXiv:2608.10273v1 Announce Type: new Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic… 4 arXiv — NLP / Computation & Language research 2d ago Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a… 32 arXiv — NLP / Computation & Language research 2d ago Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic… 36 arXiv — NLP / Computation & Language research 2d ago Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility arXiv:2608.10315v1 Announce Type: new Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an… 19 arXiv — NLP / Computation & Language research 2d ago VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise… 14 arXiv — NLP / Computation & Language research 2d ago How Robust Are LLMs to Vietnamese Dialects? arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect… 22 arXiv — NLP / Computation & Language research 2d ago From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively… 21 arXiv — NLP / Computation & Language research 2d ago MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection arXiv:2608.10459v1 Announce Type: new Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models.… 25 arXiv — NLP / Computation & Language research 2d ago Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection (DCD) therefore aims to determine whether a given text is a… 28 arXiv — NLP / Computation & Language research 2d ago Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks.… 21 Page 10 of 10 · 500 articles ← Newer