News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding arXiv:2607.11131v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by letting a lightweight drafter propose multiple tokens that are verified by a larger target model. Although effective for text-only LLMs, speculative decoding yields… 24 arXiv — NLP / Computation & Language research 1mo ago Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference Alignment arXiv:2607.11434v1 Announce Type: new Abstract: Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF… 11 arXiv — NLP / Computation & Language research 1mo ago Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection arXiv:2607.11597v1 Announce Type: new Abstract: The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced… 10 arXiv — NLP / Computation & Language research 1mo ago Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises arXiv:2606.01592v2 Announce Type: cross Abstract: This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we… 12 arXiv — Machine Learning research 1mo ago Reward Transport: Property Control in Flow Matching via Noise-Space Alignment arXiv:2607.08781v1 Announce Type: new Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data… 4 arXiv — Machine Learning research 1mo ago Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal arXiv:2607.08883v1 Announce Type: new Abstract: Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directions in activation space. This raises questions… 13 arXiv — Machine Learning research 1mo ago Forget Narrowly, Retain Broadly: Unlearning as an Asymmetric Generalization Problem arXiv:2607.09236v1 Announce Type: new Abstract: Machine unlearning in LLMs is the targeted removal of specific knowledge while preserving all other capabilities, critical for privacy and safety. Yet existing benchmarks measure it unreliably. They miss knowledge that resurfaces… 27 arXiv — Machine Learning research 1mo ago Interval Certifications for Multilayered Perceptrons via Lattice Traversal arXiv:2607.08773v1 Announce Type: cross Abstract: In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness. In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal… 34 arXiv — NLP / Computation & Language research 1mo ago An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? arXiv:2607.09053v1 Announce Type: new Abstract: Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, alongside evidence that this behavior can be… 14 arXiv — NLP / Computation & Language research 1mo ago VTaMo: Video-Text Alignment Model for Sign Language Translation arXiv:2607.09126v1 Announce Type: cross Abstract: Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation… 9 arXiv — NLP / Computation & Language research 1mo ago Entity Alignment Method of Science and Technology Patent based on Graph Convolution Network and Information Fusion arXiv:2311.00300v2 Announce Type: replace Abstract: The entity alignment of science and technology patents aims to link the equivalent entities in the knowledge graph of different science and technology patent data sources. Most entity alignment methods only use graph neural… 19 Don't Worry About the Vase community 1mo ago AI #176 Part 2: Plan B This is part 2 of the weekly, broadly covering speculation, rhetoric and policy, along with alignment research. 16 arXiv — NLP / Computation & Language research 1mo ago Efficient Safety Alignment of Language Models via Latent Personality Traits arXiv:2607.07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can… 15 arXiv — Machine Learning research 1mo ago Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA arXiv:2607.08054v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic… 15 arXiv — Machine Learning research 1mo ago CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal Consistency arXiv:2607.08555v1 Announce Type: new Abstract: The operational integrity of complex industrial systems relies on precise anomaly detection and diagnosis. The vast majority of existing methods narrowly focus on capturing temporal similarities of representations, often… 12 arXiv — Machine Learning research 1mo ago Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks arXiv:2607.08561v1 Announce Type: new Abstract: A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between… 22 arXiv — NLP / Computation & Language research 1mo ago PLURAL: A Global Dataset for Value Alignment arXiv:2607.08034v1 Announce Type: new Abstract: Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset… 14 arXiv — NLP / Computation & Language research 1mo ago Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment arXiv:2607.08256v1 Announce Type: new Abstract: Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a… 18 arXiv — NLP / Computation & Language research 1mo ago Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition arXiv:2607.08374v1 Announce Type: new Abstract: Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psychological taxonomies rather than uncovering shared underlying behavioral structure. This… 19 Hacker News — AI on Front Page community 1mo ago GPT-5.6 https://deploymentsafety.openai.com/gpt-5-6/gpt-5-6.pdf https://developers.openai.com/api/docs/guides/latest-model https://x.com/levie/status/2075287443411222628 , https://xcancel.com/levie/status/2075287443411222628 Comments URL: https://news.ycombinator.com/item?id=48849066… 27 Hugging Face Daily Papers research 1mo ago Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs Abstract Splash is a mask-isolated tactile alignment learning framework that enables multimodal LLMs to acquire tactile sensing capabilities without sacrificing vision-language reasoning through selective parameter updating that prevents catastrophic forgetting. Generated by… 22 arXiv — Machine Learning research 1mo ago Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia arXiv:2607.06626v1 Announce Type: new Abstract: Recent Vision-Language Models capture increasingly complex aspects of human cognition. Here we ask whether this alignment extends to reward valuation, which we assess in a mechanistic framework built on clinical tests that were… 32 arXiv — Machine Learning research 1mo ago When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models arXiv:2607.06630v1 Announce Type: new Abstract: Formal robustness certificates for embedded neural-interface models can pass while task accuracy collapses: at perturbation budget e=0.25, EEGNet classification accuracy drops by 25.7% under projected-gradient attack while the… 21 arXiv — Machine Learning research 1mo ago Online Data Selection Is Implicit Alignment arXiv:2607.07023v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separation is incomplete: when examples are scored and… 33 arXiv — Machine Learning research 1mo ago A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving… 6 arXiv — Machine Learning research 1mo ago Predicting LLM Safety Before Release by Simulating Deployment arXiv:2607.07184v1 Announce Type: new Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model behavior will occur in deployment: they generally have… 31 arXiv — Machine Learning research 1mo ago Avoiding unsafe sets when training with Langevin Dynamics arXiv:2607.07538v1 Announce Type: new Abstract: Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics on the loss landscape, and a natural safety question is to bound the probability $\nu_t(\mathcal{A}_H) = \mathbb{P}(Q_t \in… 16 arXiv — NLP / Computation & Language research 1mo ago Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs arXiv:2607.06831v1 Announce Type: new Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an… 14 arXiv — NLP / Computation & Language research 1mo ago Riemannian Geometry for Pre-trained Language Model Embeddings arXiv:2607.07047v1 Announce Type: new Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token… 21 arXiv — NLP / Computation & Language research 1mo ago R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement arXiv:2607.07318v1 Announce Type: new Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasible, particularly for video advertisements. However, existing safety-driven… 28 arXiv — NLP / Computation & Language research 1mo ago Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs,… 20 arXiv — NLP / Computation & Language research 1mo ago $C$-$\Delta\Theta$: Circuit-Restricted Weight Arithmetic for Selective Refusal arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compute cost and serving complexity. Activation steering is widely used, but it… 38 arXiv — NLP / Computation & Language research 1mo ago Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection arXiv:2604.07831v2 Announce Type: replace-cross Abstract: Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by… 37 r/MachineLearning community 1mo ago Agentic safety triggers aren't textual safety triggers — MCP attacks that beat SOTA guardrails more than half the time (code + dataset) [R] Most safety alignment work treats "detect the attack" as a text classification problem — does the prompt contain language the model's safety guardrails should catch. That assumption breaks down for LLM agents with real tool access. Here's a concrete case: take a known, public… 14 OpenAI official-blog 1mo ago Our approach to government and national security partnerships Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountability, and public safety. 8 arXiv — Machine Learning research 1mo ago TILDE: TILt-based Distributional Erasure for Concept Unlearning arXiv:2607.06432v1 Announce Type: new Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and safety regulations, deployed systems must be able to… 22 arXiv — NLP / Computation & Language research 1mo ago Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language… 7 arXiv — NLP / Computation & Language research 1mo ago PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails arXiv:2607.05910v1 Announce Type: cross Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product,… 27 arXiv — NLP / Computation & Language research 1mo ago Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing arXiv:2505.10356v3 Announce Type: replace Abstract: Decoding language from the human brain remains a grand challenge for Brain-Computer Interfaces (BCIs). Current approaches typically rely on unimodal brain representations, neglecting the brain's inherently multimodal… 31 arXiv — NLP / Computation & Language research 1mo ago Quantifying Retriever-Generator Alignment in RAG with Local Explanations arXiv:2601.21803v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems combine dense retrievers and language models to ground their outputs in external documents. However, the interaction between these components remains opaque, creating challenges for… 8 arXiv — NLP / Computation & Language research 1mo ago Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis arXiv:2602.00846v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar scores… 13 arXiv — NLP / Computation & Language research 1mo ago Geometric Stability: The Missing Axis of Representations arXiv:2601.09173v5 Announce Type: replace-cross Abstract: Representational similarity analysis and related methods compare the internal geometries of neural networks, but they measure only alignment between spaces, leaving a blind spot -- whether a representation's structure is… 15 Hugging Face Daily Papers research 1mo ago Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Abstract A supervised contrastive alignment framework maps WavLM embeddings from English and Mandarin into a shared clinical space for depression detection, addressing cross-lingual generalization challenges and revealing performance artifacts caused by speaker identity leakage.… 38 Hugging Face Daily Papers research 1mo ago Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Abstract Light-Omni is a multimodal agent framework that enables efficient video understanding through dual contextual states, achieving faster and more accurate video processing by eliminating iterative reasoning while maintaining semantic alignment. Generated by… 14 r/MachineLearning community 1mo ago Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D] late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or… 23 Hugging Face Daily Papers research 1mo ago PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Abstract PixWorld presents a unified pixel-space diffusion approach for 3D reconstruction and generation that overcomes limitations of latent-space methods through direct image-level supervision and geometry-aware feature alignment. Generated by Qwen/Qwen2.5-Coder-32B-Instruct… 19 arXiv — Machine Learning research 1mo ago Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data arXiv:2607.02636v1 Announce Type: new Abstract: Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, operational security environments, infrastructure monitoring and defense… 31 arXiv — Machine Learning research 1mo ago Safe Inference-Time Alignment via Lagrangian Reward Augmentation arXiv:2607.02781v1 Announce Type: new Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-time alignment methods typically optimize a single… 28 arXiv — Machine Learning research 1mo ago Bootstrap Flow-Map Tree Sampling Enables Online Feedback Driven Search arXiv:2607.02915v1 Announce Type: new Abstract: In many scientific and engineering domains, maximizing discovery within a limited sampling budget demands strategic, observation-guided exploration. While generative models have enabled training-free reward alignment, current… 27 arXiv — Machine Learning research 1mo ago Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification arXiv:2607.03075v1 Announce Type: new Abstract: Safety-critical applications require classifiers that are both robust and reliable. Adversarial training is a widely adopted defense for improving robustness in deep neural networks; however, its effect on the reliability of… 34 Page 7 of 10 · 500 articles ← Newer Older →