News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — Machine Learning research 1mo ago Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System arXiv:2607.03125v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its "black-box" exploration risks violating strict hardware safety limits. Typically, these constraints are managed… 35 arXiv — Machine Learning research 1mo ago Unbiased Alignment for Large Language Models with Noisy Preferences arXiv:2607.03248v1 Announce Type: new Abstract: The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Preference Optimization. However, these methods are vulnerable to the significant noise… 37 arXiv — Machine Learning research 1mo ago Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning arXiv:2607.03453v1 Announce Type: new Abstract: Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to select high-quality responses generated by a reference LLM. However, the efficacy of these… 15 arXiv — NLP / Computation & Language research 1mo ago Improving LLMs via Validator-to-Generator Alignment arXiv:2607.02668v1 Announce Type: new Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs… 26 arXiv — NLP / Computation & Language research 1mo ago Alignment-Guided Largest Table Overlap Size Estimation arXiv:2607.03049v1 Announce Type: new Abstract: Fast estimation of the size of the largest overlap between tables enables blocking and query-by-table retrieval in large table repositories. The first and the state-of-the-art estimator Armadillo improves efficiency by embedding… 31 arXiv — NLP / Computation & Language research 1mo ago KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment arXiv:2607.03166v1 Announce Type: new Abstract: Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimization spreads supervision over mostly shared templates. We formalize this as the Resolution… 10 arXiv — NLP / Computation & Language research 1mo ago Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization arXiv:2607.03704v1 Announce Type: new Abstract: Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Models (LLMs) show promise, yet performance on clinically demanding tasks remains… 29 arXiv — NLP / Computation & Language research 1mo ago Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees arXiv:2607.04430v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses without reliable confidence estimates. Uncertainty quantification (UQ) offers a… 34 arXiv — NLP / Computation & Language research 1mo ago Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 arXiv:2607.04510v1 Announce Type: new Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights.… 7 arXiv — NLP / Computation & Language research 1mo ago Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations arXiv:2607.04645v1 Announce Type: new Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently… 35 arXiv — NLP / Computation & Language research 1mo ago FormalRx: Rectify and eXamine Semantic Failures in Autoformalization arXiv:2607.04655v1 Announce Type: new Abstract: The veracious semantic alignment in autoformalization is significant for formal mathematical reasoning. However, existing evaluations provide only opaque binary verdicts or scalar scores, offering no interpretable insight into… 34 arXiv — NLP / Computation & Language research 1mo ago Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment arXiv:2607.04728v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. To resolve this, Importance sampling (IS) is… 36 Hugging Face Daily Papers research 1mo ago Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification Abstract Automated safety testing framework Vera uses a three-stage pipeline to identify and test safety risks in LLM agents through structured risk taxonomies, combinatorial case generation, and adaptive sandbox execution with evidence-based verification. Generated by… 7 r/LocalLLaMA community 1mo ago ThinkingCap-Qwen3.6-27B: same accuracy as base Qwen3.6 with ~50% fewer thinking We rigorously evaluate the resulting checkpoint across general reasoning, non-reasoning multiple-choice question answering, everyday multi-turn conversations, system prompt adherence, safety, math, code and agentic use cases. Due to the high variability of reasoning quality at… 12 Hugging Face Daily Papers research 1mo ago Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment Abstract Geo-Anchored Cloud Removal framework addresses semantic drift in cloud removal by combining physically grounded residual inversion with semantic manifold constraints from vision foundation models. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Cloud removal (CR) is… 29 Hugging Face Daily Papers research 1mo ago Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming Abstract AI-Infra-Guard is an open-source framework that addresses AI infrastructure security through layered detection paradigms spanning infrastructure, protocol, agent behavior, and model layers. Generated by Qwen/Qwen2.5-Coder-32B-Instruct The fast growth of open-source AI… 4 r/MachineLearning community 1mo ago Best models for generating red-team attacks? Also looking for public datasets [R] Hi everyone, I'm currently working on a framework to evaluate the security of LLM applications and AI agents, and I've been stuck on one part for a while. Most red-teaming frameworks rely on an LLM to generate adversarial prompts. My question is more about which model to use .… 29 r/MachineLearning community 1mo ago What does "Safe AI" look like? [D] ​ For open-weight LLMs, how practical is it to study defenses against post-release fine-tuning that weakens refusal or safety behavior? I've been seeing “uncensored” or “heretic” variants of new models appear very quickly after release, which raises a question I’m curious… 28 Hugging Face Daily Papers research 1mo ago Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Abstract Transfer-Aware Curriculum (TAC) improves multi-domain reinforcement learning by prioritizing domains that provide broad benefits to other domains, using gradient-geometry alignment to estimate cross-domain transferability. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 25 arXiv — Machine Learning research 1mo ago IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electrochemical diagnostics, second-life analytics, and battery safety research.… 36 arXiv — Machine Learning research 1mo ago Multi-modal Rail Crossing Safety Analysis arXiv:2607.01365v1 Announce Type: new Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our ability to do so by introducing structured data (such as official accident reports)… 9 arXiv — Machine Learning research 1mo ago CALM: Interpretable Cross-Modal Alignment for Biomarker Discovery from Unpaired Data arXiv:2607.01656v1 Announce Type: new Abstract: The interaction between brain structure and genetic influences is key to understanding neuropsychiatric disorders. However, most large-scale datasets are unimodal, providing either neuroimaging or genetics data. We propose CALM, a… 15 arXiv — NLP / Computation & Language research 1mo ago Safeguarding LLM Agents from Misalignment through Provenance Analysis arXiv:2607.01236v1 Announce Type: new Abstract: As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent's proposed tool invocation deviates from the user's intent -- a phenomenon… 17 arXiv — NLP / Computation & Language research 1mo ago Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment arXiv:2607.01239v1 Announce Type: new Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural mechanism: BPE tokenization fragments safety-critical words into sub-word… 9 arXiv — NLP / Computation & Language research 1mo ago Multi-Objective Exploration and Preference Optimization via Mutual Information arXiv:2607.01392v1 Announce Type: new Abstract: Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicting preference dimensions. Current methods achieve this trade-off by training… 33 arXiv — NLP / Computation & Language research 1mo ago MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering arXiv:2607.01420v1 Announce Type: new Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal attributions have been explored in depth, the… 10 arXiv — NLP / Computation & Language research 1mo ago HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art performance on English and multilingual prompt-safety benchmarks at roughly one-tenth… 25 arXiv — NLP / Computation & Language research 1mo ago Structuring the Space of Sociotechnical Alignment arXiv:2607.01250v1 Announce Type: cross Abstract: Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP research increasingly addresses its technical aspects, it often leaves underspecified what… 23 arXiv — NLP / Computation & Language research 1mo ago Safety Targeted Embedding Exploit via Refinement arXiv:2607.01859v1 Announce Type: cross Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this… 24 arXiv — NLP / Computation & Language research 1mo ago Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters arXiv:2607.01893v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix. Block (DLM-style) drafters predict the whole… 4 arXiv — NLP / Computation & Language research 1mo ago Online Safety Monitoring for LLMs arXiv:2607.02510v1 Announce Type: cross Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time… 26 Hugging Face Daily Papers research 1mo ago From SRA to Self-Flow: Data Augmentation or Self-Supervision? Abstract Research investigates the mechanisms behind self-alignment methods in diffusion transformers, finding that performance improvements stem primarily from data augmentation along the noise dimension rather than token interactions between noise levels. Generated by… 11 Hugging Face Daily Papers research 1mo ago Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation Abstract Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Vision-language… 10 Hugging Face Daily Papers research 1mo ago SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Abstract Scientific image generation faces challenges in semantic alignment and logical reasoning, prompting the creation of SciIR-82k dataset and SciIR-Bench evaluation framework to improve scientific reasoning capabilities in text-to-image models. Generated by… 23 MIT Technology Review — AI news-outlet 1mo ago Teaching AI to run with the turbines Artificial intelligence may have captured the public imagination through chatbots and image generators, but some of its most consequential use cases are unfolding far from consumer-facing tools. In industries where physical infrastructure, operational continuity, and safety are… 29 MIT Technology Review — AI news-outlet 1mo ago Building the foundation for an autonomous enterprise Artificial intelligence may have captured the public imagination through chatbots and image generators, but some of its most consequential use cases are unfolding far from consumer-facing tools. In industries where physical infrastructure, operational continuity, and safety are… 22 arXiv — NLP / Computation & Language research 1mo ago MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, many… 22 arXiv — Machine Learning research 1mo ago PAPA: Online Personalized Active Preference Alignment arXiv:2607.00486v1 Announce Type: new Abstract: Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like personalized recommender systems, the objective often shifts to modeling specific regions of the… 11 arXiv — Machine Learning research 1mo ago Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment arXiv:2607.00603v1 Announce Type: new Abstract: We give a descent-free, alignment-free measurement of singular structure on trained networks. At a single frozen checkpoint the read recovers the order $k$ of each dead direction from the directional-Fisher rate, the master… 34 arXiv — Machine Learning research 1mo ago Beyond Activation Alignment:The Alignment-Diversity Tradeoff in Task-Aware LLM Quantization arXiv:2607.00908v1 Announce Type: new Abstract: Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints. We first identify a phenomenon that we term the Perplexity Illusion: layers ranked as… 7 arXiv — Machine Learning research 1mo ago Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling arXiv:2607.01022v1 Announce Type: new Abstract: Spatiotemporal point processes (STPPs) model event data in continuous time and space, with applications in mobility, epidemiology, and public safety. Recent neural STPPs span expressive intensity models, conditional density models,… 14 arXiv — Machine Learning research 1mo ago Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search arXiv:2607.01144v1 Announce Type: new Abstract: While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow regions of the underlying distribution. These approaches struggle when preferences are unknown… 16 arXiv — NLP / Computation & Language research 1mo ago A Mechanistic View of Authority Hierarchy in LLM Sycophancy arXiv:2607.00415v1 Announce Type: new Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than… 17 arXiv — NLP / Computation & Language research 1mo ago Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image implicit toxicity (MIIT), where each image appears benign in isolation, but… 15 arXiv — NLP / Computation & Language research 1mo ago MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark arXiv:2607.00724v1 Announce Type: new Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption… 8 arXiv — NLP / Computation & Language research 1mo ago Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity arXiv:2607.01153v1 Announce Type: new Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded… 14 Hugging Face Daily Papers research 1mo ago Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts Abstract Vision-Language-Action models can be efficiently adapted to new environments using a single demonstration through weight vector arithmetic that isolates domain-specific information via subspace alignment. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 17 Hugging Face Daily Papers research 1mo ago ABot-M0.5: Unified Mobility-and-Manipulation World Action Model Abstract ABot-M0.5 is a World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency in autoregressive prediction. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Mobile… 16 Hugging Face Daily Papers research 1mo ago AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation Abstract GraphRAG extends RAG by incorporating graph-structured data for LLMs, addressing latent feature misalignment through Adaptive-masking for Graph Embedding (AGE) that uses Transformer-based self-supervised learning with learnable node sampling. Generated by… 16 r/MachineLearning community 1mo ago Making Optimization Work When Labels Are Scarce [R] https://www.gnosyslabs.com/case-studies/safety-classifier-sparse-labels Gnosys is an autonomous model engineer: it improves prompts and classifiers when ground truth is too sparse for conventional optimization. On ToxicChat, a public safety benchmark, under realistic label… 23 Page 8 of 10 · 500 articles ← Newer Older →