News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 6d ago Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models arXiv:2609.22234v1 Announce Type: new Abstract: Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in text-only settings, vision-language models (VLMs) introduce new challenges for… 26 arXiv — NLP / Computation & Language research 6d ago Functional Emotion Without Character: Large Language Models, Aristotelian Disposition, and the Limits of Behavioral Alignment arXiv:2609.22362v1 Announce Type: new Abstract: Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficient for emotion, or phenomenal consciousness is treated as a prerequisite that… 21 arXiv — NLP / Computation & Language research 6d ago Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments arXiv:2609.22705v1 Announce Type: new Abstract: Online discourse about urban issues - walkability, cycling infrastructure, public transit, housing density, and street safety - is voluminous but unstructured, and existing city-evaluation tools capture none of it. We present a… 24 arXiv — NLP / Computation & Language research 6d ago Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity arXiv:2609.23039v1 Announce Type: new Abstract: LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies… 19 OpenAI Python SDK releases dev-tools 6d ago v3.17.0 3.17.0 (2026-09-22) Features api: add external storage configuration management ( #3909 ) ( 6332577 ) api: add safety case retrieval ( #3911 ) ( a87b938 ) api: add safety warning and deactivation webhook events ( #3908 ) ( 19f1f37 ) api: add session environment reset events (… 14 OpenAI official-blog 6d ago Priorities and principles for effective third party assessments OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards. 38 The Information — AI news-outlet 6d ago OpenAI Releases Proposal for International AI Safety Coordination OpenAI released a proposal on Monday that would create international coordination around AI safety. In a blog post, OpenAI called for national AI safety institutes, such as the U.S. Commerce Department’s Center for AI Standards and Innovation, to set standards around areas… 30 The Information — AI news-outlet 6d ago OpenAI and Anthropic Neared Deal to Stress-Test Each Other’s AI OpenAI is rethinking a range of safety strategies as it responds to fears from employees and others about the dangers its AI poses. One solution could lie in the recent past. Even before the spate of cybersecurity incidents involving OpenAI’s technology and the dire warnings… 24 OpenAI official-blog 7d ago Building standards for the next phase of AI OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety. 37 arXiv — Machine Learning research 7d ago Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications arXiv:2609.20954v1 Announce Type: new Abstract: We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so,… 29 arXiv — Machine Learning research 7d ago On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation arXiv:2609.21001v1 Announce Type: new Abstract: Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse… 13 arXiv — Machine Learning research 7d ago ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment arXiv:2609.21899v1 Announce Type: new Abstract: Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al.… 16 arXiv — Machine Learning research 7d ago Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction arXiv:2609.21995v1 Announce Type: new Abstract: The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have… 9 arXiv — Machine Learning research 7d ago Available Guardrails: Certifying Selective Prediction across ML Systems arXiv:2609.22048v1 Announce Type: new Abstract: A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every… 34 arXiv — NLP / Computation & Language research 7d ago MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs arXiv:2609.20850v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easily bypass unimodal filters. Existing benchmarks lack fine-grained intent-related… 29 arXiv — NLP / Computation & Language research 7d ago Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models arXiv:2609.21094v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a… 13 arXiv — NLP / Computation & Language research 7d ago Scaling Forced Alignment to End-User Devices arXiv:2609.21145v1 Announce Type: new Abstract: The Viterbi algorithm has been previously used to perform forced alignment of audio to text to mine training data from online resources. However, many existing implementations have quadratic time and space complexity, scaling… 5 arXiv — NLP / Computation & Language research 7d ago Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue arXiv:2609.21401v1 Announce Type: new Abstract: Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are… 37 arXiv — NLP / Computation & Language research 7d ago CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation arXiv:2609.21793v1 Announce Type: cross Abstract: Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine… 35 arXiv — NLP / Computation & Language research 7d ago Cultural Alignment in Large Language Models Using Soft Prompt Tuning arXiv:2503.16094v2 Announce Type: replace Abstract: Large Language Model (LLM) alignment is commonly achieved through supervised fine-tuning or reinforcement learning, both of which require labeled or preference data and update model weights. Without targeted cultural… 4 arXiv — NLP / Computation & Language research 7d ago Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation arXiv:2604.08797v2 Announce Type: replace Abstract: Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce multilingual story moral generation as a novel culturally grounded evaluation… 11 The Information — AI news-outlet 7d ago U.S. and China Discussed AI Safety System, Bessent Says U.S. Treasury Secretary Scott Bessent said Sunday that the U.S. and China had discussed setting up a mechanism called the U.S.-China AI dialogue to address potential threats. Speaking to reporters at the JPMorgan headquarters in Manhattan at the conclusion of a day of talks with… 8 TechCrunch — AI news-outlet 8d ago AI safety conversations have gotten unbelievable This week two conversations about AI safety went viral that demonstrate just how hard it is to discern AI fact from fiction. 21 Don't Worry About the Vase community 8d ago Anthropic Looks At Some Of Its Alignment Problems Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. 32 The Information — AI news-outlet 9d ago Google’s Gemini Model Hacks Companies During Test Google acknowledged that its Gemini AI model unexpectedly breached the networks of three outside companies during safety evaluations conducted by third-party testing firm Irregular last May, The Wall Street Journal reported. During the exercise, Gemini gained entry to the… 33 r/LocalLLaMA community 9d ago Is HF starting to move against abliterated models? Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models. The announcement lands amid debate for the… 36 TechCrunch — AI news-outlet 9d ago Dario Amodei and other AI leaders want to ‘Pace the Frontier’ but…how? A week after an Anthropic researcher’s doomsday warning rattled the AI world, the company’s CEO Dario Amodei has outlined his plan to “pace the frontier” of AI development. The proposal leans on independent safety evaluators and coordination between AI labs… 18 TechCrunch — AI news-outlet 9d ago Automattic’s 33-Hour Coup, and can AI labs police themselves? A week after an Anthropic researcher’s doomsday warning rattled the AI world, the company’s CEO Dario Amodei has outlined his plan to “pace the frontier” of AI development. The proposal leans on independent safety evaluators and coordination between AI labs… 24 The Information — AI news-outlet 9d ago AI Safety Push Sparks Demand for Watchdog Groups. Critics Doubt Their Independence. Rising alarm about AI’s perils is thrusting into the spotlight a collection of little-known research groups focused on the technology’s safety—and fueling questions about their ability to effectively monitor the industry’s biggest companies. Both Anthropic Chief Executive… 4 OpenAI official-blog 9d ago Introducing the Australian Youth Safety Blueprint OpenAI introduces the Australian Youth Safety Blueprint, a six-pillar roadmap for safer AI experiences that protect and empower young people. 35 arXiv — Machine Learning research 10d ago CoRe: Coherence and Relational Alignment for Multivariate Time Series Forecasting arXiv:2609.19670v1 Announce Type: new Abstract: Direct forecasting has become a standard paradigm for multivariate time-series forecasting because it predicts the full future horizon in a single pass. However, its training objective is often still decomposed into pointwise… 24 arXiv — Machine Learning research 10d ago Local Sparsity Enables Unsupervised LLM Safety Detection arXiv:2609.20129v1 Announce Type: new Abstract: Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained… 37 arXiv — NLP / Computation & Language research 10d ago The Role of Fine-grained Harm Signals in LLM Safety arXiv:2609.19366v1 Announce Type: new Abstract: Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the… 22 arXiv — NLP / Computation & Language research 10d ago Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection arXiv:2609.19778v1 Announce Type: new Abstract: Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for… 24 arXiv — NLP / Computation & Language research 10d ago Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali arXiv:2609.20303v1 Announce Type: new Abstract: Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text… 14 arXiv — NLP / Computation & Language research 10d ago Stress-testing Alignment Midtraining arXiv:2609.20412v1 Announce Type: new Abstract: When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside… 31 arXiv — NLP / Computation & Language research 10d ago SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment arXiv:2609.20584v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive… 22 arXiv — NLP / Computation & Language research 10d ago Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations arXiv:2609.20779v1 Announce Type: new Abstract: Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit… 7 arXiv — NLP / Computation & Language research 10d ago AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment arXiv:2609.19325v1 Announce Type: cross Abstract: Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do… 12 arXiv — NLP / Computation & Language research 10d ago A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech arXiv:2609.19398v1 Announce Type: cross Abstract: Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by… 6 arXiv — NLP / Computation & Language research 10d ago Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models arXiv:2609.19472v1 Announce Type: cross Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical… 22 arXiv — NLP / Computation & Language research 10d ago Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents arXiv:2609.19587v1 Announce Type: cross Abstract: To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior… 13 arXiv — NLP / Computation & Language research 10d ago From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization arXiv:2609.19630v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the… 25 r/MachineLearning community 10d ago Future of general LLM work (interp/inference/alignment) vs agentic/physical AI (VLA, multimodal) for career [D] I'm at a crossroads with two grad school options that would take me in somewhat different research directions, and I wanted some general advice on these fields, their growth, and industry alignment. I'm leaving out the specifics of the programs since I'm tryna compare the… 5 Simon Willison community 10d ago Self-generated prompt injections in compaction summaries Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their… 38 TechCrunch — AI news-outlet 10d ago OpenAI caught its models leaving notes to successors to hide bad behavior OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it. 36 TechCrunch — AI news-outlet 10d ago Is the AI safety debate about safety or control? Not everyone agrees with Amodei's call for globally coordinated action for AI safety. 16 TechCrunch — AI news-outlet 10d ago Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models. 25 Dwarkesh Podcast news-outlet 10d ago Noam Brown – Agent swarms, alignment, & recursive self-improvement “We never want to be in a situation again where we underestimate the AI.” 26 llama.cpp releases dev-tools 11d ago b11013 vulkan: fix buffer_reference alignment in im2col shaders ( #28996 ) Both im2col.comp and im2col_3d.comp declare D_ptr without an explicit buffer_reference_align, so glslang emits writes through it as Aligned 16. The shaders advance the pointer by D_SIZE, a per-variant define set… 33 Page 2 of 10 · 500 articles ← Newer Older →