News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow Hacker News — AI on Front Page community 24d ago Flock Credibility Lost as It Repeatedly Lies to City Councils, Police, & Public Article URL: https://www.aclu.org/news/privacy-technology/tracking-alpr-cameras/flock-safety-credibility-lost-as-it-repeatedly-lies-to-city-councils-police-departments-and-public-across-the-country Comments URL: https://news.ycombinator.com/item?id=48986731 Points: 237 #… 14 r/LocalLLaMA community 24d ago Head of US AI safety agency resigns   submitted by   /u/fallingdowndizzyvr [link]   [comments] 11 OpenAI official-blog 25d ago Safety and alignment in an era of long-horizon models OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment. 36 arXiv — Machine Learning research 25d ago Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation arXiv:2607.15562v1 Announce Type: new Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits. Existing packing assistants are template-driven and… 17 arXiv — Machine Learning research 25d ago Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework arXiv:2607.15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer… 6 arXiv — Machine Learning research 25d ago CoG-Guided Weight Correction for Fault-Tolerant Deep Neural Networks arXiv:2607.15753v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) used in safety-critical applications are vulnerable to hardware and memory faults that corrupt network weights and degrade reliability. In this paper, we propose a Center of Gravity (CoG) guided weight… 8 arXiv — Machine Learning research 25d ago QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides arXiv:2607.15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4… 26 arXiv — Machine Learning research 25d ago Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment arXiv:2607.15928v1 Announce Type: new Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce. Existing… 38 arXiv — Machine Learning research 25d ago Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes arXiv:2607.16087v1 Announce Type: new Abstract: AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to… 33 arXiv — Machine Learning research 25d ago PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment arXiv:2607.16156v1 Announce Type: new Abstract: Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands… 7 arXiv — NLP / Computation & Language research 25d ago Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior arXiv:2607.15286v1 Announce Type: cross Abstract: We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a… 35 arXiv — Machine Learning research 25d ago Unsupervised Keypoints for Real-Time Fall Detection: Comparative Analysis Under Real-world Conditions with Predictive Bandwidth Reduction arXiv:2607.15400v1 Announce Type: cross Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth.… 11 arXiv — NLP / Computation & Language research 25d ago Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair arXiv:2607.15282v1 Announce Type: cross Abstract: Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state. This synchronic framing has shaped artificial systems, where empathic behavior is defined as affect recognition and… 8 arXiv — NLP / Computation & Language research 25d ago Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs arXiv:2508.10029v3 Announce Type: replace Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ), which works by pairing a harmful query with a… 31 arXiv — NLP / Computation & Language research 25d ago Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We… 10 arXiv — NLP / Computation & Language research 25d ago Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing arXiv:2606.07636v2 Announce Type: replace-cross Abstract: Long-form video editing over heterogeneous footage requires agents to coordinate source selection, multimodal analysis, timeline construction, narration and subtitle alignment, rendering, and revision while exposing… 37 r/MachineLearning community 26d ago AAAI 27 AI Alignment track [D] How to submit to AI alignment track? I can only see these at openReview: AAAI 2027 AAAI 2027 Artificial Intelligence for Social Impact Track AAAI 2027 Conference AAAI 2027 Innovative Applications of AI   submitted by   /u/Silencer_Wasd [link]   [comments] 8 Don't Worry About the Vase community 28d ago AI #177 Part 2: Wish You Were Here As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. 35 Hugging Face Daily Papers research 28d ago SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment Abstract CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation models to match image regions to… 22 arXiv — Machine Learning research 28d ago LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration arXiv:2607.14410v1 Announce Type: new Abstract: Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines. We present LATTICE (Latent Alignment of Tissue-level and… 8 arXiv — NLP / Computation & Language research 28d ago Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak arXiv:2607.14147v1 Announce Type: new Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a… 24 arXiv — NLP / Computation & Language research 28d ago Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text arXiv:2607.14103v1 Announce Type: new Abstract: Multi-agent systems (MAS) are utilized in many contexts and many professions. Those MAS rely on inter-agent communication, usually implemented by clear-text message passing. We hypothesize that Large Language Models may have a… 9 arXiv — NLP / Computation & Language research 28d ago Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning arXiv:2607.14117v1 Announce Type: new Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous… 17 arXiv — NLP / Computation & Language research 28d ago Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv:2607.14682v1 Announce Type: cross Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge. Current approaches bifurcate into Supervised Fine-Tuning… 14 arXiv — NLP / Computation & Language research 28d ago MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection arXiv:2607.15166v1 Announce Type: cross Abstract: Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary failed? We present a clinician-built synthetic benchmark and failure atlas that labels… 27 arXiv — NLP / Computation & Language research 28d ago Decoupled Alignment for Robust Plug-and-Play Adaptation arXiv:2406.01514v4 Announce Type: replace Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust… 13 arXiv — NLP / Computation & Language research 28d ago Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity arXiv:2510.01171v4 Announce Type: replace Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level… 10 VentureBeat — AI news-outlet 28d ago The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty… 26 r/LocalLLaMA community 29d ago Filings: Dario Amodei gave $1M in May to Public First, a super PAC advocating for AI safety regulations, seemingly his first seven-figure political donation   submitted by   /u/pscoutou [link]   [comments] 26 arXiv — Machine Learning research 29d ago CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion arXiv:2607.13120v1 Announce Type: new Abstract: Inferring gene regulatory networks (GRNs) from single-cell transcriptomic data is crucial for biological discovery, yet existing approaches suffer from a fundamental misalignment with real-world needs. Researchers typically seek a… 35 arXiv — Machine Learning research 29d ago SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy arXiv:2607.13175v1 Announce Type: new Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate,… 8 arXiv — Machine Learning research 29d ago Distributionally Robust and Safe Imitation Learning arXiv:2607.13436v1 Announce Type: new Abstract: Imitation learning (IL) has achieved remarkable success in complex decision-making tasks. However, its performance is highly sensitive to distribution shifts, which can pose significant safety risks. We propose a distributionally… 28 arXiv — Machine Learning research 29d ago Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as… 38 Hugging Face Daily Papers research 29d ago PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Abstract Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed… 25 OpenAI official-blog 1mo ago The US is advancing AI safety through state and federal action OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI. 35 OpenAI official-blog 1mo ago GPT-Red: Unlocking Self-Improvement for Robustness Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. 14 arXiv — Machine Learning research 1mo ago Scalable Optimal Transport Algorithm for Network Alignment arXiv:2607.11952v1 Announce Type: new Abstract: Network alignment identifies node correspondences across different networks and is a fundamental primitive in many data science applications, including social network analysis, fraud detection, and knowledge graph integration.… 20 arXiv — Machine Learning research 1mo ago Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financial risk management, yet conventional approaches rely on application-specific… 13 arXiv — Machine Learning research 1mo ago Predictive Modeling of High-Altitude Clear Air Turbulence in the United States: A Machine Learning Approach arXiv:2607.11899v1 Announce Type: cross Abstract: High-altitude Clear Air Turbulence (CAT) poses significant risks to aviation safety due to its unpredictability and challenges in detection. This study leverages machine learning models to improve CAT prediction within U.S.… 24 arXiv — NLP / Computation & Language research 1mo ago CANDI: Contextual Alignment for Niche Domains Question Answering arXiv:2607.11891v1 Announce Type: new Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabilities beyond general knowledge. Traditional question-answering benchmarks often… 33 arXiv — NLP / Computation & Language research 1mo ago Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings arXiv:2607.12071v1 Announce Type: new Abstract: Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces and neural coding patterns, which severely impedes cross-modal alignment… 34 arXiv — NLP / Computation & Language research 1mo ago Optimization Is Not All You Need arXiv:2607.11977v1 Announce Type: cross Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering… 4 arXiv — NLP / Computation & Language research 1mo ago Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety… 30 arXiv — NLP / Computation & Language research 1mo ago From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models arXiv:2604.26052v4 Announce Type: replace Abstract: Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide how risk changes between prompt and response.… 37 Marcus on AI community 1mo ago Breaking: Demis Hassabis endorses preflight safety testing for AI Good news, for once. 19 Hugging Face Daily Papers research 1mo ago Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Abstract Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive… 36 arXiv — Machine Learning research 1mo ago Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs arXiv:2607.09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility. Model fine-tuning achieves robust safety but compromises general utility. Input-side safety guardrails… 16 arXiv — Machine Learning research 1mo ago SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification arXiv:2607.09936v1 Announce Type: new Abstract: Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those threats first appear. Generalized zero-shot learning offers a natural solution by enabling… 28 arXiv — NLP / Computation & Language research 1mo ago Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement arXiv:2607.10590v1 Announce Type: new Abstract: We investigate how annotator demographic attributes, supplied as prompt cues, shape the alignment between large language model (LLM) predictions and human annotations across five tasks. Using five open-source LLMs, we… 26 arXiv — NLP / Computation & Language research 1mo ago MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment arXiv:2607.11070v1 Announce Type: new Abstract: Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn… 4 Page 6 of 10 · 500 articles ← Newer Older →