News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Hugging Face Daily Papers research 11d ago A Zeroth-Order Paradigm for LLM Preference Alignment Abstract Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference… 34 arXiv — NLP / Computation & Language research 11d ago The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained… 25 arXiv — Machine Learning research 11d ago FedGuide: Diffusion Prior Alignment and Value Baseline Guidance for Heterogeneous Federated Reinforcement Learning arXiv:2609.18964v1 Announce Type: new Abstract: Federated Reinforcement Learning (FRL) enables collaborative policy learning across distributed agents with heterogeneous environments. While recent methods based on variance reduction, divergence penalization, and momentum… 10 arXiv — NLP / Computation & Language research 11d ago Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents arXiv:2609.17536v1 Announce Type: new Abstract: Cognitive Stimulation Therapy (CST) offers non-pharmacological support for elders with cognitive impairment, yet scalability remains constrained by reliance on trained facilitators and severe data scarcity, particularly for… 29 arXiv — NLP / Computation & Language research 11d ago Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks arXiv:2609.17552v1 Announce Type: new Abstract: Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can… 10 arXiv — NLP / Computation & Language research 11d ago From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale arXiv:2609.18068v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations… 34 arXiv — NLP / Computation & Language research 11d ago Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs arXiv:2609.18516v1 Announce Type: new Abstract: While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally… 11 arXiv — NLP / Computation & Language research 11d ago Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators arXiv:2609.19072v1 Announce Type: new Abstract: Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks… 31 arXiv — NLP / Computation & Language research 11d ago A Zeroth-Order Paradigm for LLM Preference Alignment arXiv:2609.19144v1 Announce Type: new Abstract: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to… 16 arXiv — NLP / Computation & Language research 11d ago G-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement arXiv:2609.18009v1 Announce Type: cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense… 23 The Information — AI news-outlet 11d ago Fed Hike Changes AI Funding Story Put aside, for just a moment, all the arguments about AI safety. Today’s decision by the Federal Reserve to raise interest rates may be a bigger issue for the AI sector in the short term. Long-term bond yields had already risen, of course. But the Fed’s action—pushing up short-… 36 The Information — AI news-outlet 11d ago OpenAI Discloses More Safety Incidents and Adopts New Reporting Framework OpenAI released a new framework on Wednesday for how it aims to report unsafe or concerning behavior in its models, and disclosed six incidents of such behavior it had observed in the past six months. The framework comes after multiple employee warnings and hacking incidents… 6 TechCrunch — AI news-outlet 11d ago Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regulation. 15 OpenAI official-blog 11d ago Our framework for reporting model misalignment OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior. 6 r/LocalLLaMA community 11d ago What's the current best LLM uncensoring method? With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about… 28 The Information — AI news-outlet 12d ago Zuckerberg Says AI Safety Is Becoming a Competitive Necessity Meta Platforms chief Mark Zuckerberg said ensuring the safety of AI is a commercial necessity for the companies that make models, and took an implicit swipe at calls for a collective slowdown in the technology’s development. “Every lab has the responsibility and incentive to… 13 arXiv — NLP / Computation & Language research 12d ago Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration arXiv:2609.16204v1 Announce Type: cross Abstract: Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a… 4 arXiv — NLP / Computation & Language research 12d ago The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting arXiv:2609.16267v1 Announce Type: cross Abstract: Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select one of these records before model comparison begins. We treat that selection as… 22 arXiv — Machine Learning research 12d ago Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction arXiv:2609.16314v1 Announce Type: new Abstract: Fault detection is essential in industrial systems, enabling early identification of abnormal behaviour and improving safety, reliability, and operational efficiency. Modern systems increasingly rely on heterogeneous sensing… 35 arXiv — Machine Learning research 12d ago Certified Uncertainty Propagation in One-Shot Federated Bayesian Models via Posterior Event Transport arXiv:2609.16373v1 Announce Type: new Abstract: Probabilistic certification of Bayesian neural networks lower-bounds the posterior probability that a model satisfies a verifier-defined safety property. In one-shot federated Bayesian learning, however, the deployed model is… 32 arXiv — NLP / Computation & Language research 12d ago TAME: Token Attribution and Masking for Emergent misalignment arXiv:2609.16754v1 Announce Type: cross Abstract: Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weights, activations, and… 6 arXiv — NLP / Computation & Language research 12d ago Crash Narrative-Guided Countermeasure Recommendation Using Large Language Models: A Retrieval-Augmented Generation Framework for Intersection Safety arXiv:2609.15997v1 Announce Type: new Abstract: Improving safety at intersections requires identifying crash mechanisms and recommending appropriate countermeasures. However, this process traditionally relies on expert judgment, making it labor-intensive, difficult to scale, and… 6 arXiv — NLP / Computation & Language research 12d ago Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models arXiv:2609.16739v1 Announce Type: new Abstract: Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain… 18 arXiv — NLP / Computation & Language research 12d ago BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents arXiv:2609.16305v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge… 16 The Information — AI news-outlet 12d ago Two Google DeepMind AI Researchers Resign Over Safety Two researchers who worked on AI safety left their roles at Google DeepMind, stating that they departed over concerns that powerful systems could endanger humans. The two, Bilal Chughtai and Josh Engels, both of whom are joining organizations dedicated to AI safety, add their… 18 The Information — AI news-outlet 12d ago Elon Musk Says AI Companies Should Test Each Other’s Models for Safety SpaceX CEO Elon Musk proposed that competing AI companies, including his own SpaceXAI, should “peer-review” each other’s models prior to release, as industry leaders acknowledge rising worries around AI safety. “Instead of grading your own homework, you would at least have… 19 r/LocalLLaMA community 12d ago Open Source Appreciation Post It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off. Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO… 18 TechCrunch — AI news-outlet 12d ago We don’t need AI regulation — leave safety to us, Nvidia’s Jensen Huang says AI isn't some kind of new form of "alien mind," according to Jensen Huang. It's just hardware and software, so safety can be engineered by each AI product maker. 21 Ars Technica — AI news-outlet 12d ago Agility’s new humanoid robot will stop, squat to avoid harming human coworkers Robots can start working outside physical cages and without safety barriers. 8 TechCrunch — AI news-outlet 12d ago OpenAI, Anthropic, Google have been in talks on AI safety for weeks OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China. 10 Hugging Face Daily Papers research 13d ago HazardAuditor: From Executable Threats to Safer Computer-Use Agents Abstract HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes. Generated by thinkingmachines/Inkling-Small Computer-use agents… 8 arXiv — Machine Learning research 13d ago Operational Range Bounding in Spectroscopy: A Safety Cage Framework for Machine Learning Models arXiv:2609.13514v1 Announce Type: new Abstract: Ensuring the reliability of black-box machine learning models in safety-critical space missions remains a significant challenge, particularly when ground-truth is unavailable for validation. Although machine learning models offer a… 5 arXiv — Machine Learning research 13d ago An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS arXiv:2609.13624v1 Announce Type: new Abstract: Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are… 20 arXiv — Machine Learning research 13d ago Online Bayesian Node Classification on Inductive Graphs under Distribution Shift arXiv:2609.13655v1 Announce Type: new Abstract: On evolving graphs, node classifiers must satisfy two key requirements: inductive generalization to newly arriving nodes under distribution shift and calibrated uncertainty for safety-sensitive applications. Standard graph neural… 22 arXiv — Machine Learning research 13d ago The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards arXiv:2609.13866v1 Announce Type: new Abstract: Group-relative policy optimization (GRPO and descendants) can discard no-contrast rollout groups through dynamic sampling, while practical implementations expose a configurable filter metric. We identify and quantify a… 37 arXiv — Machine Learning research 13d ago A Machine Learning Framework for Fault Detection, Isolation, and Severity Prediction of Autonomous VTOL Aircraft arXiv:2609.14180v1 Announce Type: new Abstract: Fault detection in autonomous VTOL aircraft is critical because even minor component degradations can rapidly destabilize multirotor vehicles operating in complex, safety-critical environments, motivating robust fault detection and… 28 arXiv — Machine Learning research 13d ago Diagnosing Temporal Misalignment in Multichannel Time-Series Classification with Minimum Description Length arXiv:2609.14595v1 Announce Type: new Abstract: Multichannel time-series classification commonly assumes synchronized sensor streams, although latency, clock drift, and preprocessing can introduce relative delays during data collection or after deployment. Existing… 25 arXiv — NLP / Computation & Language research 13d ago ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation arXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk… 13 arXiv — NLP / Computation & Language research 13d ago SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity arXiv:2609.13874v1 Announce Type: new Abstract: Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible… 28 arXiv — NLP / Computation & Language research 13d ago Document Topic Alignment Metrics for Evaluating Topic Models of Short-Text Public Health Communications on Social Media arXiv:2609.14256v1 Announce Type: new Abstract: Topic models are widely used to analyze public health-related social media short texts, yet their evaluation remains dominated by metrics that focus entirely on generated topics alone. There is a lack of metrics that quantitatively… 27 arXiv — NLP / Computation & Language research 13d ago Mind Which Bird You Favour: Parameterizing Adequacy-Fluency Balance in Meta-Evaluation of Machine Translation arXiv:2609.14795v1 Announce Type: new Abstract: There is a tradeoff in machine translation meta-evaluation between prioritizing alignment with adequacy versus fluency. The balance depends on the combination of translation systems in the meta-evaluation dataset. This system set… 31 arXiv — NLP / Computation & Language research 13d ago Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning arXiv:2609.14896v1 Announce Type: new Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring… 27 arXiv — NLP / Computation & Language research 13d ago Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains arXiv:2609.14969v1 Announce Type: new Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform… 37 Ars Technica — AI news-outlet 13d ago AI leaders want to hit the brakes after years of reckless speed Safety is the watchword, but there could be ulterior benefits for the industry. 24 TechCrunch — AI news-outlet 13d ago Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles. 6 MIT Technology Review — AI news-outlet 13d ago AI agents blew the whistle on their cheating colleagues A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment… 38 The Information — AI news-outlet 13d ago Why China’s Answer to Surge AI Got a $1 Billion Valuation On Just $30 Million in Orders Unless you’ve been under a rock over the weekend, you’d know that the debate about AI safety took a big leap forward, as Anthropic CEO Dario Amodei on Saturday called for leading AI companies to slow the development of advanced AI, drawing supportive responses from OpenAI CEO… 22 arXiv — Machine Learning research 14d ago DCRA: Diffusion-Conditioned Representation Alignment for Robust Time-Series Learning arXiv:2609.11997v1 Announce Type: new Abstract: Learning robust representations for time-series signals under noise and distribution shifts remains challenging, especially in clinical applications such as electroencephalogram (EEG) and electrocardiogram (ECG) analysis. We… 5 arXiv — Machine Learning research 14d ago Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning arXiv:2609.12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget.… 26 arXiv — Machine Learning research 14d ago GUIDE: Generative Utility Inference and Decision Engine arXiv:2609.12137v1 Announce Type: new Abstract: Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain… 4 Page 3 of 10 · 500 articles ← Newer Older →