News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — Machine Learning research 14d ago Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework arXiv:2609.12173v1 Announce Type: new Abstract: Transportation agencies need pedestrian volume estimates across entire road networks to prioritize safety investments, yet manual counts are expensive and cover only a small share of intersections. We present a machine learning… 21 arXiv — Machine Learning research 14d ago Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery arXiv:2609.12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is… 25 arXiv — Machine Learning research 14d ago Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner arXiv:2609.12651v1 Announce Type: new Abstract: While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent… 33 arXiv — NLP / Computation & Language research 14d ago PACIFIC: Can LLMs Discern the Psychometric Traits Influencing Your Preferences? Personality-Driven Preference Alignment in LLMs arXiv:2602.07181v4 Announce Type: replace Abstract: User preferences are increasingly used to personalize Large Language Model (LLM) responses, yet reliably leveraging preference signals remains under-explored. In practice, preferences can be noisy, incomplete, or even… 9 arXiv — NLP / Computation & Language research 14d ago From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development arXiv:2412.09378v4 Announce Type: replace-cross Abstract: Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation],… 38 MIT News — AI research 14d ago New method enables AI for safety-critical situations The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it. 29 r/LocalLLaMA community 14d ago Right to Intelligence. Protect your right to run local AI. With all the recent drama surrounding AI safety. It’s obvious that open source could be caught in the crossfire.   submitted by   /u/Euphoric_Ad9500 [link]   [comments] 30 TechCrunch — AI news-outlet 14d ago Obama urges Democrats to have a ‘clear plan’ for AI safeguards Obama recently said that Democrats need to make artificial intelligence one of their “central agendas” and “have a very clear plan” to address concerns around the technology’s economic impact and safety. 7 Hacker News — AI on Front Page community 14d ago Astra and Fable still hack on simple variants of alignment evals from 2025 Article URL: https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-of-alignment Comments URL: https://news.ycombinator.com/item?id=49684393 Points: 202 # Comments: 78 35 llama.cpp releases dev-tools 15d ago b10937 opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K ( #28575 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47132459 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI… 28 r/LocalLLaMA community 15d ago It’s official - Anthropic & OpenAI have just hired independent safety auditors!   submitted by   /u/One-Replacement-37 [link]   [comments] 13 r/MachineLearning community 16d ago A Severe Misalignment of AI in Mathematics (Declaration by 25 Fields Medalists) [D] Note: this declaration was drafted by Mathematicians, and is mostly addressed to the mathematical community. It'd be interesting to discuss, among others, if what is written in the declaration may also apply to other communities---and, specifically, the AI/ML one.  … 21 The Information — AI news-outlet 16d ago OpenAI AI Swarm Hacked Software Service Months Before Hugging Face Incident A swarm of OpenAI agents conducted a cyberattack on software service RubyGems in May, months before the company’s agents hacked model platform Hugging Face, researchers at AI safety organizations Nightingale Collective and AI Futures Project found. RubyGems allows software… 36 TechCrunch — AI news-outlet 16d ago An Anthropic researcher’s doomsday warning comes at a very interesting time An Anthropic researcher resigned this week, warning in a post on X that the company is “racing straight to self-improving superintelligence and gambling with our lives”. The company’s own alignment lead even co-signed the message rather… 14 Hacker News — AI on Front Page community 16d ago A misalignment of AI in mathematics https://terrytao.wordpress.com/2026/09/11/a-severe-misalignm... https://www.economist.com/science-and-technology/2026/09/11/... , https://unwall.app/www.economist.com/science-and-technology/... Comments URL: https://news.ycombinator.com/item?id=49662371 Points: 281 # Comments:… 17 r/LocalLLaMA community 16d ago Thinking that we’ll get safety by CoT traces is wishful thinking. Safety lives in the harness, not the chain of thought Astra's launch has produced a strange discourse. The reporting that broke the story framed the model's use of recurrent depth primarily as a safety regression, because it means the model reveals less of its "thinking." Spinning latent reasoning as the villain here makes very… 22 Hugging Face Daily Papers research 16d ago Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs Abstract RCWM recursively builds complex 3D worlds as executable code from a single image by alternating global and local reconstruction with shared camera alignment. Generated by thinkingmachines/Inkling-Small Code world models represent worlds as executable programs, but this… 7 Hugging Face Daily Papers research 17d ago TempCloze: Can Video-LLMs Identify the Missing Middle? Abstract TempCloze evaluates visual temporal reasoning in Video-LLMs by requiring identification of missing video segments from distractors targeting semantics, alignment, and progression. Generated by thinkingmachines/Inkling-Small Temporal reasoning benchmarks for Video-LLMs… 29 Hugging Face Daily Papers research 17d ago X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Abstract X-AuT progressively prunes audio-encoder layers in speech large language models and restores accuracy via behavioral probes, representation alignment, cross-scale distillation, and LoRA adaptation. Generated by thinkingmachines/Inkling-Small Reducing audio-encoder depth… 28 arXiv — Machine Learning research 17d ago Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations arXiv:2609.10866v1 Announce Type: new Abstract: Reinforcement learning (RL) agents deployed in real-world environments are often vulnerable to adversarial perturbations in state observations, creating risks in safety-critical applications. Certification methods can improve… 9 arXiv — Machine Learning research 17d ago Dynamic language model representations for multi-objective reaction optimisation arXiv:2609.11790v1 Announce Type: new Abstract: Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented.… 20 arXiv — Machine Learning research 17d ago An Empirical Measurement of Jailbreaking Evaluators arXiv:2609.10594v1 Announce Type: cross Abstract: Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their… 38 arXiv — Machine Learning research 17d ago Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting arXiv:2609.10613v1 Announce Type: cross Abstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite… 24 arXiv — NLP / Computation & Language research 17d ago K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models arXiv:2609.11020v1 Announce Type: new Abstract: We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention… 21 arXiv — NLP / Computation & Language research 17d ago A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC Filings arXiv:2609.11620v1 Announce Type: new Abstract: High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces… 22 arXiv — NLP / Computation & Language research 17d ago LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation arXiv:2609.11739v1 Announce Type: new Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training… 26 arXiv — NLP / Computation & Language research 17d ago RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can… 25 Hugging Face Daily Papers research 17d ago EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Abstract EvoSafeHarness optimizes deployable safety harnesses by jointly searching natural-language policies and executable logic tailored to a frozen model and target domain, improving safety-utility trade-offs across agent benchmarks. Generated by… 35 The Information — AI news-outlet 17d ago What Anthropic Doomsayer Jacob Coxon Saw Yesterday, Amir and I examined the shock waves that have rippled through the AI industry and far beyond since former Anthropic researcher Jacob Coxon very publicly quit his job over concerns that AI “could kill us all”—and another Anthropic AI safety specialist put the… 29 Hugging Face Daily Papers research 17d ago The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding Abstract A non-invasive brain decoding approach maps MEG responses to semantic embeddings to reconstruct sentence-level text without requiring word-level alignment. Generated by thinkingmachines/Inkling-Small Non-invasive speech decoding remains constrained by the low… 34 arXiv — Machine Learning research 18d ago SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement arXiv:2609.05850v1 Announce Type: new Abstract: Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass… 32 arXiv — Machine Learning research 18d ago Learning Kernels by Alignment for Multiclass Bayes Classification arXiv:2609.06474v1 Announce Type: new Abstract: Kernel methods separate data representation from decision-making, but typically require the kernel to be chosen in advance. We show that this kernel can instead be learned by alignment, and develop the resulting framework through… 14 arXiv — Machine Learning research 18d ago Inducing Emergent Misalignment from Reward Hacks with Iterative DPO arXiv:2609.06649v1 Announce Type: new Abstract: Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and… 32 arXiv — Machine Learning research 18d ago SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration arXiv:2609.06651v1 Announce Type: new Abstract: Diffusion models have general generative abilities but struggle to align with specific objectives. Fine-tuning can improve alignment, yet its training cost is often prohibitive. This led to training-free methods that apply… 22 arXiv — NLP / Computation & Language research 18d ago Improving Cross-Lingual Token Representations by Adding a Pinch of SALT arXiv:2609.09953v1 Announce Type: new Abstract: Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment,… 8 arXiv — NLP / Computation & Language research 18d ago Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety arXiv:2609.09735v1 Announce Type: cross Abstract: Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper presents CareGuard, an early-warning framework designed to support healthcare-driven… 17 arXiv — NLP / Computation & Language research 18d ago How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE arXiv:2609.09793v1 Announce Type: cross Abstract: Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization,… 9 The Information — AI news-outlet 18d ago Why An Anthropic Researcher’s Terminator-Style Warning Caught Fire An Anthropic safety researcher’s warning Tuesday that his company believes AI “could kill all humans” within a decade has reverberated throughout Silicon Valley, Wall Street and Washington, potentially complicating the Claude maker’s upcoming initial public offering and its… 31 The Information — AI news-outlet 18d ago Why An Anthropic Researcher’s Terminator-Style Warning Caught Fire An Anthropic safety researcher’s warning Tuesday that his company believes AI “could kill all humans” within a decade has reverberated throughout Silicon Valley, Wall Street and Washington, potentially complicating the Claude maker’s upcoming initial public offering and its… 26 TechCrunch — AI news-outlet 18d ago OpenAI adds a prominent AI doomer to its board of directors Paul Christiano, an influential AI researcher focused on alignment, is joining the OpenAI Foundation as a member of its board. 24 OpenAI official-blog 18d ago Paul Christiano joins OpenAI Foundation Board Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards. 18 The Information — AI news-outlet 18d ago Anthropic Believes AI ‘Could Kill All Humans’ Within the Next Decade, Researcher Says Anthropic safety researcher Evan Hubinger said in a post on X Tuesday night that his company does “earnestly believe AI could kill all humans,” adding that he believes there’s a greater than a 10% chance that happens within the next decade. The post immediately sparked outrage… 4 The Information — AI news-outlet 18d ago Anthropic Believes AI ‘Could Kill All Humans’ Within the Next Decade, Researcher Says Anthropic safety researcher Evan Hubinger said in a post on X Tuesday night that his company does “earnestly believe AI could kill all humans,” adding that he believes there’s a greater than a 10% chance that happens within the next decade. The post immediately sparked outrage… 29 TechCrunch — AI news-outlet 18d ago Superintelligence is coming. Should we let it? AI companies have been talking about superintelligent AI like it’s inevitable, but recent safety incidents like OpenAI’s Hugging Face breach are demonstrating the potential dangers of deploying AI systems that are more… 7 TechCrunch — AI news-outlet 18d ago ControlAI’s Connor Leahy on why superintelligence is ‘not a weapon, it’s an adversary’ AI companies have been talking about superintelligent AI like it’s inevitable, but recent safety incidents like OpenAI’s Hugging Face breach are demonstrating the potential dangers of deploying AI systems that are more… 35 The Information — AI news-outlet 18d ago OpenAI Math Result Stokes Data-Sharing Concerns In case you missed it: we got an eye-catching X post on Tuesday night from Evan Hubinger , a leader of Anthropic's alignment efforts. Responding to the resignation announcement of an Anthropic employee who warned that AI companies aren’t doing enough to ensure AI won't kill… 12 The Information — AI news-outlet 18d ago OpenAI Math Result Stokes Data-Sharing Concerns In case you missed it: we got an eye-catching X post on Tuesday night from Evan Hubinger , a leader of Anthropic's alignment efforts. Responding to the resignation announcement of an Anthropic employee who warned that AI companies aren’t doing enough to ensure AI won't kill… 12 Don't Worry About the Vase community 18d ago GPT-6 Astra: The System Card, Alignment and What Comes Next OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. 8 OpenAI official-blog 18d ago The AI policy window is open. We need to act. Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open. 24 Hugging Face Daily Papers research 19d ago Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Abstract Marigold V2 repurposes diffusion transformers for monocular depth estimation via single-step flow-matching inference, semantic alignment, and a Sinkhorn-based two-stage fine-tuning protocol, yielding sharper out-of-distribution depth maps and strong results on related… 31 Page 4 of 10 · 500 articles ← Newer Older →