News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — Machine Learning research 22d ago Variance-reduced Domain Adaptation using Paired Sampling arXiv:2607.20367v1 Announce Type: new Abstract: Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their… 11 arXiv — Machine Learning research 22d ago Online Variance Reduction for Domain Adaptation on Streaming Data arXiv:2607.20374v1 Announce Type: new Abstract: This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been… 18 arXiv — NLP / Computation & Language research 22d ago Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment arXiv:2607.19371v1 Announce Type: cross Abstract: Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaffolding collapse: under sustained student pressure, a tutor gradually abandons guided… 9 arXiv — NLP / Computation & Language research 22d ago Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework arXiv:2607.19361v1 Announce Type: new Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk… 30 arXiv — NLP / Computation & Language research 22d ago The Two-Process Theory of Machine Self-Report arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc… 11 arXiv — NLP / Computation & Language research 22d ago OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety… 12 arXiv — NLP / Computation & Language research 22d ago The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models arXiv:2607.20265v1 Announce Type: new Abstract: Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives… 34 arXiv — NLP / Computation & Language research 22d ago Sound Probabilistic Safety Bounds for Large Language Models arXiv:2607.20286v1 Announce Type: new Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to… 16 arXiv — NLP / Computation & Language research 22d ago LKValues: Aligning Large Language Models with Sri Lankan Societal Values arXiv:2607.20410v1 Announce Type: new Abstract: Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique… 38 arXiv — NLP / Computation & Language research 22d ago JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an… 37 arXiv — NLP / Computation & Language research 22d ago Rewarding Better Thinking for LLM Preference Alignment arXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often… 9 arXiv — NLP / Computation & Language research 22d ago JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety arXiv:2607.19913v1 Announce Type: cross Abstract: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate… 26 arXiv — NLP / Computation & Language research 22d ago LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization arXiv:2407.00740v2 Announce Type: replace Abstract: As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, such as non-toxicity and logical consistency, as well as task- and… 14 arXiv — NLP / Computation & Language research 22d ago Abstraction Induces the Brain Alignment of Language and Speech Models arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very little is known about the… 13 arXiv — NLP / Computation & Language research 22d ago Meta-Learning Preferences for Multilingual LLM Alignment arXiv:2607.13315v2 Announce Type: replace Abstract: Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language… 29 arXiv — NLP / Computation & Language research 22d ago Prompt Programming for Cultural Bias and Alignment of Large Language Models arXiv:2603.16827v2 Announce Type: replace-cross Abstract: Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural biases that misalign with target populations. As LLMs are increasingly used for… 11 Hugging Face Daily Papers research 22d ago Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Abstract Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained representations that support… 31 r/LocalLLaMA community 22d ago China’s Kimi K3 fuels fears safety curbs are holding back US AI interesting to see the reverse of the American frontier model makers' stance coming from the Chinese side via South China Morning Post   submitted by   /u/zxyzyxz [link]   [comments] 9 r/MachineLearning community 23d ago Institution Prestige VS Research Alignment When Choosing University For Masters [D] When choosing a university for a masters in ML/DL, what is more important if someone wants to go into research and an eventual PhD. Is it the ranking/prestige factor of the university or the strength of the research groups in the university? Should an admission decision be made… 9 Stratechery (Ben Thompson) community 23d ago OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips OpenAI accidentally hacked Hugging Face, but the takeaways are more encouraging than people realize. 10 arXiv — Machine Learning research 23d ago Dual-domain fused LSTM modeling for efficient time-dependent reliability analysis arXiv:2607.18291v1 Announce Type: new Abstract: Time-dependent reliability analysis is crucial for ensuring the long-term safety and performance of engineering systems under uncertainties. However, traditional surrogate model methods often struggle to incorporate… 33 arXiv — Machine Learning research 23d ago On the Limits of Support-Preserving Alignment and Bounded Filtering arXiv:2607.18295v1 Announce Type: new Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models. Recent research… 11 arXiv — Machine Learning research 23d ago TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue arXiv:2607.18304v1 Announce Type: new Abstract: The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often… 21 arXiv — Machine Learning research 23d ago Conditioned Direct Feedback Alignment via Activity and Error Geometry arXiv:2607.18574v1 Announce Type: new Abstract: Direct feedback alignment (DFA) trains hidden layers with fixed random projections of the output error, avoiding the transposed-weight backward pass of backpropagation (BP). We study a failure mode of DFA training that is distinct… 9 arXiv — NLP / Computation & Language research 23d ago Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs arXiv:2607.18639v1 Announce Type: cross Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain… 34 arXiv — Machine Learning research 23d ago Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals arXiv:2607.18860v1 Announce Type: new Abstract: Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning methods depend mainly on temperature and may therefore miss mechanical precursors that emerge… 29 arXiv — Machine Learning research 23d ago KALE: Kernel Alignment with Loss Equilibration for Stable CLIP-DINOv2 Alignment at Web Scale arXiv:2607.18885v1 Announce Type: new Abstract: Kernel-based alignment of CLIP toward a vision centric teacher such as DINOv2 (KUEA) improves CLIP's visual representations while preserving text-encoder compatibility, using a fixed trade-off weight tuned on curated ImageNet-1K.… 35 arXiv — NLP / Computation & Language research 23d ago CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness arXiv:2607.18820v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used to improve both the performance and interpretability of large language models (LLMs), yet the generated reasoning may not faithfully support the final answer. We study this problem… 22 arXiv — NLP / Computation & Language research 23d ago Operational Hallucination and Safety Drift in AI Agents arXiv:2607.18366v1 Announce Type: cross Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal… 26 arXiv — NLP / Computation & Language research 23d ago MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications arXiv:2409.07314v3 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical… 17 arXiv — NLP / Computation & Language research 23d ago Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT arXiv:2509.23381v2 Announce Type: replace Abstract: We introduce Guard Vector, a safety task vector computed as the parameter difference between a guardrail model (Guard Model) and a same-architecture pretrained language model. Composing this vector with a target language model… 28 Don't Worry About the Vase community 23d ago OpenAI Shares Some Alignment Problems Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. 13 r/LocalLLaMA community 23d ago New Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size) https://huggingface.co/Nanbeige/Nanbeige4.2-3B Nanbeige4.2-3B is a compact agentic model built on Nanbeige4.2-3B-Base , designed to combine strong agentic behavior with broad reasoning and alignment capabilities. Its Looped Transformer architecture reuses the transformer layers… 33 r/LocalLLaMA community 24d ago Be Careful when Purchasing CMP 170HX on Alibaba! Just a heads up. Shops in China are running like chickens without a head after the news the Falcon Exploit working to jailbreak some of the functions of these cards. Is not just happening on Alibaba but also Ebay. Usually from Chinese sellers. I spent 2 days contacting lost of… 28 arXiv — Machine Learning research 24d ago LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats arXiv:2607.16227v1 Announce Type: new Abstract: LLMs are increasingly deployed in security-critical systems across healthcare, finance, education, and decision support, yet their inability to forget creates serious cybersecurity, privacy, and safety risks. Sensitive personal… 9 arXiv — Machine Learning research 24d ago Normalized Rewards for Preference Optimization arXiv:2607.16240v1 Announce Type: new Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood… 20 arXiv — Machine Learning research 24d ago TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment arXiv:2607.16242v1 Announce Type: new Abstract: Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, service providers need to recover models' safety… 32 arXiv — Machine Learning research 24d ago Self-Evolving Just-In-Time Memory for Proactive Embodied Safety arXiv:2607.16247v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during closed-loop interactions. Existing safety approaches… 12 arXiv — Machine Learning research 24d ago Bridging battery design and health assessment through virtual sensing and physics-informed learning arXiv:2607.16864v1 Announce Type: new Abstract: Supercharging of lithium-ion batteries (LiBs) requires robust health monitoring to ensure durability, safety, and user confidence, particularly for emerging vehicle-to-grid applications with bidirectional energy flows. Yet battery… 38 arXiv — Machine Learning research 24d ago When Can Safe Controllers Adapt? Information before Commitment arXiv:2607.16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently across environments as data arrive. Only its… 18 arXiv — Machine Learning research 24d ago Distilled Reinforcement Learning for LLM Post-training arXiv:2607.17247v1 Announce Type: new Abstract: Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL… 21 arXiv — NLP / Computation & Language research 24d ago Group Entropy-Controlled Policy Optimization arXiv:2607.16850v1 Announce Type: new Abstract: Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on… 31 arXiv — NLP / Computation & Language research 24d ago Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models arXiv:2607.17270v1 Announce Type: new Abstract: Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small… 35 arXiv — NLP / Computation & Language research 24d ago Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila arXiv:2607.18066v1 Announce Type: new Abstract: The value alignment of large language models (LLMs) is crucial for ensuring responses align with human intention and value preferences. However, most evaluations of value alignment focus on Western or universal values, while… 18 arXiv — NLP / Computation & Language research 24d ago How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? arXiv:2607.18114v1 Announce Type: new Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We… 32 arXiv — NLP / Computation & Language research 24d ago How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv:2607.17152v1 Announce Type: cross Abstract: Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a… 22 arXiv — NLP / Computation & Language research 24d ago L1 Augmented Attention as an Improved Vector Similarity Metric arXiv:2607.18027v1 Announce Type: cross Abstract: Scaled dot product attention conflates directional alignment and vector magnitude, limiting its effectiveness as a similarity metric in Transformer models. We introduce L1 augmented attention, a simple and computationally… 18 Hugging Face Daily Papers research 24d ago Distilled Reinforcement Learning for LLM Post-training Abstract Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome… 30 Hugging Face Daily Papers research 24d ago Group Entropy-Controlled Policy Optimization Abstract Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of heterogeneous tasks, which induce… 36 Hugging Face Daily Papers research 24d ago DiFA: Inference-Time Forward-Process Alignment for Diffusion Models Abstract The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of the denoising process. In this… 5 Page 5 of 10 · 500 articles ← Newer Older →