News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Smol AI News news-outlet 10d ago not much happened today **Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning, while **Mistral AI** released **Shieldstral**, a 3B parameter open-weights safety model for… 17 arXiv — Machine Learning research 10d ago Inference-Time Policy Alignment for Fair Reinforcement Learning arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For… 22 arXiv — Machine Learning research 10d ago Neural operator learning for collision-aware trajectory planning of spacecraft swarms arXiv:2608.00320v1 Announce Type: new Abstract: Autonomous spacecraft swarms must plan fuel-efficient, collision-free maneuvers in increasingly congested orbits, yet classical trajectory optimization scales poorly as pairwise safety constraints multiply with swarm size, and… 37 arXiv — NLP / Computation & Language research 10d ago A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard) arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts.… 26 arXiv — NLP / Computation & Language research 10d ago OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that is… 14 arXiv — NLP / Computation & Language research 10d ago Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable to malicious jailbreak prompts. Transfer-based attacks construct adversarial… 6 arXiv — NLP / Computation & Language research 10d ago ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, Egyptian, Algerian, Palestinian, and Moroccan. The dataset is annotated with… 19 arXiv — NLP / Computation & Language research 10d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 arXiv — NLP / Computation & Language research 10d ago Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese arXiv:2608.01629v1 Announce Type: new Abstract: Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we… 30 Hugging Face Daily Papers research 11d ago In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing Abstract Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex and lacks well-established… 18 Hugging Face Daily Papers research 11d ago SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift Abstract RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture benchmarking and insufficient… 32 Hugging Face Daily Papers research 11d ago Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Abstract Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and… 8 arXiv — Machine Learning research 11d ago LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment arXiv:2607.28669v1 Announce Type: new Abstract: We present LARA (Lightweight Additive Residual Adaptation), a method for efficient adaptation that operates in the residual stream of a frozen model rather than in its weights. Where LoRA adds an update of low rank to weight… 36 arXiv — Machine Learning research 11d ago Beyond Feature and Structure Alignment: Learning Transferable Propagation Knowledge for Graph Foundation Models arXiv:2607.28980v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for enabling knowledge transfer across diverse domains. Unlike traditional graph learning methods that are typically designed for in-domain settings, GFMs… 29 arXiv — NLP / Computation & Language research 11d ago Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers. Existing retrieval-augmented methods… 26 OpenAI official-blog 13d ago Advancing responsible AI across Europe OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances. 17 Hugging Face Daily Papers research 14d ago Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing Abstract In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive,… 6 arXiv — Machine Learning research 14d ago Regularizing modality contribution drift in multimodal continual learning arXiv:2607.27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic… 36 arXiv — Machine Learning research 14d ago The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models arXiv:2607.27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step… 23 arXiv — Machine Learning research 14d ago Context-Informed Ship Trajectory Prediction via Conditional Attention arXiv:2607.27418v1 Announce Type: new Abstract: Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architectures have improved forecasting horizons, they predominantly rely on historical… 10 arXiv — Machine Learning research 14d ago When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment arXiv:2607.27530v1 Announce Type: new Abstract: Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a class label or molecular property. Multiple heads can separate these aspects, but… 23 arXiv — Machine Learning research 14d ago Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for… 26 arXiv — Machine Learning research 14d ago Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning arXiv:2607.27626v1 Announce Type: new Abstract: Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor's peak Age of Information (peak AoI, also abbreviated PAoI) to stay below a hard per-slot deadline,… 29 arXiv — Machine Learning research 14d ago DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement arXiv:2607.27761v1 Announce Type: new Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view… 14 arXiv — NLP / Computation & Language research 14d ago Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups arXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we… 28 arXiv — NLP / Computation & Language research 14d ago BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences arXiv:2607.27366v1 Announce Type: new Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than… 22 arXiv — NLP / Computation & Language research 14d ago Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a… 32 arXiv — NLP / Computation & Language research 14d ago Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and… 26 arXiv — NLP / Computation & Language research 14d ago Inducing language models to assert their own consciousness restores human beliefs and values arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety… 18 arXiv — NLP / Computation & Language research 14d ago Measuring Alignment With Reader Highlights Net of Position and Length arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting… 38 Ars Technica — AI news-outlet 14d ago Google reveals Gemini Robotics 2.0, promising improved dexterity and safety Gemini Robotics 2 includes three models, but only one is publicly available right now. 22 Hugging Face Daily Papers research 15d ago GPT-Red: Automated Red Teaming via Self-Play at Scale Abstract We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially… 27 arXiv — Machine Learning research 15d ago Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation arXiv:2607.26164v1 Announce Type: new Abstract: Automated molecular structure elucidation from infrared (IR) spectroscopy data has seen significant advancements in recent years, but its broad applicability is limited by a reliance on pre-determined chemical formulas provided as… 19 arXiv — Machine Learning research 15d ago Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models arXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a… 5 arXiv — Machine Learning research 15d ago Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon… 4 arXiv — Machine Learning research 15d ago Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal… 15 arXiv — Machine Learning research 15d ago Shape-Based Inductive Bias for Glioma Grading from Tumor Contours arXiv:2607.26090v1 Announce Type: cross Abstract: Glioma grading from tumor contours is often treated as a pixel problem even when the signal of interest is shape. We align closed contours with a functional shape-alignment framework, separate global deformation from residual… 11 arXiv — NLP / Computation & Language research 15d ago GPT-Red: Automated Red Teaming via Self-Play at Scale arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production… 11 arXiv — Machine Learning research 15d ago A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment arXiv:2607.26170v1 Announce Type: cross Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background… 21 arXiv — NLP / Computation & Language research 15d ago Steering Instruction Hierarchies at Inference Time arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often… 33 arXiv — NLP / Computation & Language research 15d ago Misalignment Has a Personality: A Big Five Account of Emergent Misalignment arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in… 13 arXiv — NLP / Computation & Language research 15d ago Constitutional Midtraining: Content Presence Drives Alignment Gains arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional… 31 arXiv — NLP / Computation & Language research 15d ago DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models arXiv:2607.26891v1 Announce Type: new Abstract: Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alignment and low inference efficiency. To address these issues, we propose DIRECT,… 20 arXiv — NLP / Computation & Language research 15d ago OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate… 35 arXiv — NLP / Computation & Language research 15d ago Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in… 29 arXiv — NLP / Computation & Language research 15d ago On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that… 5 TechCrunch — AI news-outlet 15d ago Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI Weng previously served as the VP of AI Safety Research at OpenAI. 27 Hugging Face Daily Papers research 15d ago Projection Pursuit CPCANet for Domain Generalization Abstract Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract domain-invariant structures through batch-wise Common Principal Component Analysis (CPCA). However, CPCANet suffers from… 7 arXiv — NLP / Computation & Language research 16d ago MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To… 33 arXiv — NLP / Computation & Language research 16d ago IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment arXiv:2607.25579v1 Announce Type: new Abstract: Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods mainly exploit explicit graph structures and textual fields, which often provide insufficient… 7 Page 3 of 10 · 500 articles ← Newer Older →