News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 3d ago On the use of foundation models in cognitive science arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive… 34 arXiv — NLP / Computation & Language research 3d ago SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages.… 21 arXiv — NLP / Computation & Language research 3d ago Safety Cost of Steering Vectors Is Separable and Reducible arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests,… 34 arXiv — NLP / Computation & Language research 3d ago Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks arXiv:2608.09624v1 Announce Type: new Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the… 6 arXiv — Machine Learning research 4d ago Online Conformal Prediction Beyond Feedback arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box… 37 arXiv — Machine Learning research 4d ago Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize… 11 arXiv — Machine Learning research 4d ago Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits arXiv:2608.07430v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as… 36 arXiv — Machine Learning research 4d ago Game-Theoretic Inverse Reinforcement Learning for Modeling Competitive Human Driving: A Cut-in Prediction Study arXiv:2608.06445v1 Announce Type: cross Abstract: Capturing the strategic decision-making inherent in competitive human driving is critical for autonomous vehicle safety and traffic simulation. This study demonstrates that game-theoretic Inverse Reinforcement Learning (IRL)… 25 arXiv — NLP / Computation & Language research 4d ago Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models arXiv:2608.06409v1 Announce Type: new Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a… 4 arXiv — NLP / Computation & Language research 4d ago StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a… 5 Interconnects (Nathan Lambert) research 4d ago Lessons from the hacks Musings on model alignment, what determines safety, and where we go from here. 36 TechCrunch — AI news-outlet 4d ago The AI safety test is becoming a safety risk AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry standards and regulation can keep pace with increasingly powerful models. 12 Dwarkesh Podcast news-outlet 6d ago 8 Predictions for the Era of Continual Learning Locking in AI safety regulation now is a mistake. 18 Ars Technica — AI news-outlet 6d ago AI chatbots have failed people in crisis. Can that be fixed? Clinicians and researchers say AI companies need to open up their safety data. 16 TechCrunch — AI news-outlet 6d ago New Mexico court orders Meta to pay additional $567M in child safety case Meta's total fine has raked up to $942 million in this case 33 arXiv — Machine Learning research 7d ago Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language arXiv:2608.05238v1 Announce Type: new Abstract: Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is… 31 arXiv — Machine Learning research 7d ago Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG arXiv:2608.05315v1 Announce Type: new Abstract: Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and sessions. While Riemannian alignment methods like the Riemannian Centering… 37 arXiv — Machine Learning research 7d ago Align-RAG: Alignment Is All You Need for TSFM In-Context Learning arXiv:2608.05571v1 Announce Type: new Abstract: Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge… 11 arXiv — Machine Learning research 7d ago CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits arXiv:2608.05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions… 14 arXiv — Machine Learning research 7d ago A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric… 33 arXiv — Machine Learning research 7d ago SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models arXiv:2608.06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains… 17 arXiv — Machine Learning research 7d ago A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance arXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning,… 15 arXiv — Machine Learning research 7d ago From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction arXiv:2608.05203v1 Announce Type: cross Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning. Motivated… 23 arXiv — NLP / Computation & Language research 7d ago Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025)… 20 arXiv — NLP / Computation & Language research 7d ago From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a… 33 arXiv — NLP / Computation & Language research 7d ago ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment arXiv:2608.06110v1 Announce Type: cross Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared… 28 arXiv — NLP / Computation & Language research 7d ago Explanations of Large Language Models Explain Language Representations in the Brain arXiv:2502.14671v4 Announce Type: replace Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment. We test whether explainable AI (XAI) can help answer this: using… 24 Hugging Face Daily Papers research 8d ago Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Abstract Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming… 25 arXiv — Machine Learning research 8d ago Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that… 35 arXiv — Machine Learning research 8d ago Local Violation Certification for Linear Predict-Then-Optimize Pipelines arXiv:2608.04474v1 Announce Type: new Abstract: Data-driven decision pipelines combining predictive machine learning models with downstream optimization software are increasingly used to make high-stakes operational decisions. Certifying the safety, fairness, and reliability of… 30 arXiv — Machine Learning research 8d ago Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think arXiv:2608.04613v1 Announce Type: new Abstract: Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and industrial monitoring. Despite the large number of proposed anomaly detection… 30 arXiv — NLP / Computation & Language research 8d ago DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific fine-tuning can also weaken the safety guardrails of aligned LLMs. A widely… 27 arXiv — NLP / Computation & Language research 8d ago Social Pressure Breaks Majority Voting in LLM Safety Panels arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the… 36 arXiv — NLP / Computation & Language research 8d ago EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment arXiv:2608.04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic… 18 Simon Willison community 8d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 22 Simon Willison community 8d ago Third-party cyber evaluations involving OpenAI models Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular :… 19 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 37 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 14 arXiv — Machine Learning research 9d ago Learning Molecular Representations from Cellular Phenotypes with Structure Preservation arXiv:2608.02688v1 Announce Type: new Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment… 11 arXiv — NLP / Computation & Language research 9d ago Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety arXiv:2608.02617v1 Announce Type: new Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a… 15 arXiv — NLP / Computation & Language research 9d ago Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound… 7 arXiv — NLP / Computation & Language research 9d ago Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation arXiv:2608.03044v1 Announce Type: new Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak… 38 arXiv — NLP / Computation & Language research 9d ago HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition? arXiv:2608.03105v1 Announce Type: new Abstract: Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, and free word order. Prior ensemble work has emphasized architectural diversity… 9 arXiv — NLP / Computation & Language research 9d ago Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach arXiv:2608.03204v1 Announce Type: new Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the requirements of visual reasoning tasks. However, existing RL-based alignment methods are… 18 arXiv — NLP / Computation & Language research 9d ago ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently emerged as a promising attack paradigm. They bypass… 22 arXiv — NLP / Computation & Language research 9d ago Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough? arXiv:2608.03446v1 Announce Type: new Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual… 37 arXiv — NLP / Computation & Language research 9d ago Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by… 5 arXiv — NLP / Computation & Language research 9d ago Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving… 38 r/LocalLLaMA community 9d ago China’s Open-Weight Models Will Be Spared US Safety Tests   submitted by   /u/fallingdowndizzyvr [link]   [comments] 36 TechCrunch — AI news-outlet 9d ago Open-weight AI models are catching up to the frontier. The safety gap remains. A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards. 11 Page 2 of 10 · 500 articles ← Newer Older →