News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow arXiv — NLP / Computation & Language research 16d ago Evaluation of forced alignment of code-mixed speech: the case of Hindi-English arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We… 35 arXiv — NLP / Computation & Language research 16d ago Shieldstral arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety… 19 arXiv — NLP / Computation & Language research 16d ago LLM Scheming Inversely Scales with Pretraining Language Coverage arXiv:2607.24769v1 Announce Type: cross Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned… 36 arXiv — NLP / Computation & Language research 16d ago Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging. Pure parametric Large Language models (LLMs) do not contain… 10 arXiv — NLP / Computation & Language research 16d ago Towards Robust Reinforcement Learning for Small-Scale Language Model Agents arXiv:2607.25091v1 Announce Type: cross Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the… 37 arXiv — NLP / Computation & Language research 16d ago MemSFT: Mitigating Alignment Tax with an External Parametric Memory arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We… 37 arXiv — NLP / Computation & Language research 16d ago Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently… 14 Hugging Face Daily Papers research 16d ago Shieldstral Abstract We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content… 14 Hugging Face Daily Papers research 16d ago Towards Robust Reinforcement Learning for Small-Scale Language Model Agents Abstract The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen… 34 arXiv — NLP / Computation & Language research 17d ago Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail… 21 arXiv — Machine Learning research 17d ago Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation arXiv:2607.22766v1 Announce Type: new Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks,… 25 arXiv — Machine Learning research 17d ago Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis arXiv:2607.22797v1 Announce Type: new Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers… 38 arXiv — Machine Learning research 17d ago Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign… 30 arXiv — Machine Learning research 17d ago Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing falsification approaches rely on conditional sampling strategies that factor the… 19 arXiv — Machine Learning research 17d ago Context-Aware Concept Distillation for Trustworthy Flood Prediction arXiv:2607.23237v1 Announce Type: new Abstract: Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates a barrier to trust and accountability in high-stakes public safety decisions. While existing… 21 arXiv — Machine Learning research 17d ago Directional Influence Function: Estimating Training Data Influence in Constrained Learning arXiv:2607.23388v1 Announce Type: new Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training… 25 arXiv — NLP / Computation & Language research 17d ago Not All LLM Reasoning is Visible in the Chain-of-Thought arXiv:2607.22925v1 Announce Type: new Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically… 16 arXiv — NLP / Computation & Language research 17d ago Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining arXiv:2607.23175v1 Announce Type: new Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language… 35 arXiv — NLP / Computation & Language research 17d ago SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding arXiv:2607.23991v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which… 24 arXiv — NLP / Computation & Language research 17d ago From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages arXiv:2607.24542v1 Announce Type: new Abstract: Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and… 5 arXiv — NLP / Computation & Language research 17d ago STAIF: A Stage-wise Optimization for Complex Instruction Following arXiv:2607.22649v1 Announce Type: cross Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment methods, such as DPO, optimize holistic reward signals that often… 17 arXiv — NLP / Computation & Language research 17d ago How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift arXiv:2607.22676v1 Announce Type: cross Abstract: Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader… 11 Hugging Face Daily Papers research 17d ago OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Abstract Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging… 12 TechCrunch — AI news-outlet 17d ago OpenAI’s Hugging Face breach has reignited the debate over alignment and control OpenAI's Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better contained, or both. 14 Hugging Face Daily Papers research 18d ago LAMAR: An Open Language-Aware Multilingual Alignment Reranker Abstract In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document… 25 arXiv — Machine Learning research 18d ago Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning arXiv:2607.21646v1 Announce Type: new Abstract: Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement… 24 arXiv — Machine Learning research 18d ago Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs arXiv:2607.22314v1 Announce Type: new Abstract: Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random… 8 arXiv — NLP / Computation & Language research 18d ago Adversarial Prompts for Acceptance Collapse in Speculative Decoding arXiv:2607.21804v1 Announce Type: cross Abstract: Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence… 10 arXiv — NLP / Computation & Language research 18d ago Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimization arXiv:2607.21619v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved impressive performance, but their safety alignment remains vulnerable to jailbreak attacks. Existing content-based jailbreaks are often inconsistent and show unsatisfying… 23 arXiv — NLP / Computation & Language research 18d ago Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from… 37 arXiv — NLP / Computation & Language research 18d ago When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a critical concern. While prior work has examined LLMs' moral judgment and… 24 r/LocalLLaMA community 20d ago Any idea about these jailbreaks? Do you guys have any idea what these jailbreaks are? I searched, but I couldn't find any.   submitted by   /u/Suhan_XD [link]   [comments] 29 Simon Willison community 20d ago Quoting Boris Cherny More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. — Boris Cherny… 10 arXiv — Machine Learning research 21d ago End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier functions (CBFs). Incorporating CBFs into end-to-end policy training requires embedding… 5 arXiv — Machine Learning research 21d ago TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging arXiv:2607.20887v1 Announce Type: new Abstract: Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment. We formulate merging as a finite descent problem in which checkpoints are local objects,… 23 arXiv — Machine Learning research 21d ago Emergent Misalignment Recruits a Pre-existing Persona Subspace arXiv:2607.21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment. We ask why the narrow lesson generalizes… 5 arXiv — NLP / Computation & Language research 21d ago Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models arXiv:2607.20436v1 Announce Type: new Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same… 35 arXiv — NLP / Computation & Language research 21d ago The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs arXiv:2607.20449v1 Announce Type: new Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed… 12 arXiv — NLP / Computation & Language research 21d ago Rushes: A Human Preference Dataset for Pluralistic Alignment arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collected through a game interface where users interact with AI-generated branching… 15 arXiv — NLP / Computation & Language research 21d ago QuantiBias: Benchmarking Quantization-Induced Bias in LLMs arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side… 20 arXiv — NLP / Computation & Language research 21d ago Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin arXiv:2607.21332v1 Announce Type: new Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners… 20 arXiv — NLP / Computation & Language research 21d ago Expectation Alignment of Language Models for Real-World User Expectations arXiv:2607.20485v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model… 34 OpenAI Python SDK releases dev-tools 21d ago v2.48.0 2.48.0 (2026-07-23) Full Changelog: v2.47.0...v2.48.0 Features api: accept None for prompt_cache_key/safety_identifier ( 36820e6 ) api: add support for spend_limit admin apis ( 1ff13af ) 9 Don't Worry About the Vase community 21d ago AI #178: A Fire Alarm For General Intelligence The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to… 9 Hugging Face Daily Papers research 22d ago SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Abstract Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to… 37 arXiv — Machine Learning research 22d ago Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning arXiv:2607.19394v1 Announce Type: new Abstract: Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and neural signal patterns vary substantially across individuals. To investigate such… 38 arXiv — Machine Learning research 22d ago Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely… 11 arXiv — Machine Learning research 22d ago OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization arXiv:2607.19806v1 Announce Type: new Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors… 4 arXiv — Machine Learning research 22d ago Test Case Prioritization for DNNs via Neural Collapse Instability arXiv:2607.20046v1 Announce Type: new Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization… 15 arXiv — Machine Learning research 22d ago Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library arXiv:2607.20277v1 Announce Type: new Abstract: Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires interpretability. While fuzzy rule-based systems offer transparent, linguistically… 23 Page 4 of 10 · 500 articles ← Newer Older →