News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow r/LocalLLaMA community 1h ago NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not. https://x.com/JensenHuang/status/2104499465055023424   submitted by   /u/InternationalGap3698 [link]   [comments] 14 NVIDIA Developer Blog official-blog 2h ago NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of... 7 arXiv — NLP / Computation & Language research 7h ago Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence arXiv:2609.30535v1 Announce Type: new Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only… 12 arXiv — NLP / Computation & Language research 7h ago Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment arXiv:2609.30802v1 Announce Type: new Abstract: Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during… 35 arXiv — NLP / Computation & Language research 7h ago Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers arXiv:2609.31342v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document… 36 arXiv — NLP / Computation & Language research 7h ago ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs arXiv:2609.31448v1 Announce Type: new Abstract: Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from… 29 arXiv — NLP / Computation & Language research 7h ago MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos arXiv:2609.31553v1 Announce Type: new Abstract: Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of… 24 arXiv — NLP / Computation & Language research 7h ago Affective Flow Language Model for Emotional Support Conversation arXiv:2602.08826v3 Announce Type: replace Abstract: Large language models (LLMs) have advanced emotional support conversation, but existing alignment methods rely mainly on sparse preferences at the response level or outcomes at the dialogue level, providing limited supervision… 36 arXiv — NLP / Computation & Language research 7h ago Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders arXiv:2603.18863v2 Announce Type: replace Abstract: Cross-lingual alignment is often assumed to improve cross-lingual transfer by bringing representations of different languages closer together. However, improvements in representational alignment do not consistently translate… 13 llama.cpp releases dev-tools 2d ago b11193 hexagon: find software divide calls using binary inspection tool ( #29449 ) hex-scripts: fix table alignment hex-scripts: find sw div calls using binary inspection tool Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/50359675… 28 The Information — AI news-outlet 2d ago Exclusive: Meta Bolsters Muse Safety Warning After Security Vulnerability Found Meta Platforms is adding a clearer safety warning within Muse after a security researcher discovered a vulnerability in the AI agent that could let an attacker access a user’s sensitive personal information. The security flaw, which was flagged by an outside researcher through… 18 The Information — AI news-outlet 3d ago Chinese Leader Xi Calls for U.S.-China Cooperation to Prevent AI Abuse Chinese leader Xi Jinping said during his summit with U.S. President Donald Trump that the two nations can “jointly prevent the misuse and abuse of AI” through more cooperation. Trump and Xi held talks on Thursday at the White House, and a major topic on the agenda was AI safety… 34 arXiv — Machine Learning research 3d ago Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems arXiv:2609.29687v1 Announce Type: new Abstract: Automated driving systems (ADSs) are increasingly operating on public roads, raising safety concerns, yet reliable prediction of crash injury severity remains difficult because crash reports are limited, severe outcomes are rare,… 32 arXiv — Machine Learning research 3d ago Safety-oriented pedestrian trajectory prediction at urban intersections using time-to-collision and crossing-zone context arXiv:2609.29706v1 Announce Type: new Abstract: Accurate pedestrian trajectory prediction is important for proactive road-safety applications, particularly at urban intersections where pedestrian motion is shaped by both vehicle interactions and crossing context. This study… 23 arXiv — NLP / Computation & Language research 3d ago An Explainable DistilBERT-BiLSTM-Attention Framework for Binary and Multi-Class Hate Speech Detection arXiv:2609.28703v1 Announce Type: new Abstract: Hate speech on social media poses serious risks to social harmony, mental well-being, and public safety, making its timely and accurate detection essential for content moderation systems. Most existing studies focus on binary… 37 The Information — AI news-outlet 3d ago Google, OpenAI and Anthropic AI Safety Group Takes Shape Google , OpenAI and Anthropic are pushing forward with a plan to create a new AI safety-focused standards body on their own, without government oversight, in hopes of launching it by the end of the year or early in 2027, according to people familiar with the matter. The three… 14 llama.cpp releases dev-tools 3d ago b11159 vulkan: handle misalignment in conv_2d and conv_3d ( #29365 ) vulkan: handle misalignment in conv_2d and conv_3d fix test-backend-ops print Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/49846182 macOS/iOS: macOS Apple Silicon (arm64)… 6 arXiv — Machine Learning research 4d ago COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation arXiv:2609.26853v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing… 5 arXiv — Machine Learning research 4d ago Stable Neural Decoding Across Sessions via Task-Conditioned Latent Alignment for Brain-Machine Interfaces arXiv:2609.27441v1 Announce Type: new Abstract: Achieving stable long-term neural decoding in invasive brain-machine interfaces (BMIs) remains challenging due to variations in recorded neural populations across sessions. Current latent alignment approaches may overlook… 8 arXiv — Machine Learning research 4d ago Learning Where to Look: A Shared Relative-Alignment Module for Time-Series Forecasting and PPG-to-Vital-Sign Reconstruction arXiv:2609.27473v1 Announce Type: new Abstract: PPG-to-vital-sign reconstruction turns a wrist-worn photoplethysmogram into clinical waveforms such as the ECG. Long-horizon multivariate time-series forecasting underpins planning in energy, weather, and traffic. Both generate a… 16 arXiv — Machine Learning research 4d ago DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment arXiv:2609.27572v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues… 10 arXiv — Machine Learning research 4d ago CAST: Context- and Anomaly Structure-Conditioned Time Series Anomaly Generation arXiv:2609.27825v1 Announce Type: new Abstract: Anomalous time series play a critical role in safety-critical domains, yet they are inherently scarce, heterogeneous, and costly to obtain. Existing time series generation methods predominantly focus on synthesizing normal data,… 16 arXiv — NLP / Computation & Language research 4d ago Hard Negatives Reveal What Easy Negatives Hide: Cross-Lingual Harmfulness Representations Degrade with Resource Tier Under Hard Negatives arXiv:2609.27758v1 Announce Type: new Abstract: Safety alignment in large language models is trained primarily in English, and recent work reports that the underlying harmfulness representation survives translation: English-trained probes separate harmful from harmless prompts… 38 arXiv — NLP / Computation & Language research 4d ago Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures arXiv:2609.27773v1 Announce Type: new Abstract: As Large Language Models (LLMs) move from conversational assistants to advanced agentic systems, guardrail failures can convert adversarial intents into harmful executions. However, most guardrail evaluation frameworks focus only… 29 arXiv — NLP / Computation & Language research 4d ago Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks arXiv:2609.27900v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multi-agent systems where a principal agent decomposes tasks and delegates them to subordinate agents that may invoke external tools. Safety alignment, however, is still… 14 arXiv — NLP / Computation & Language research 4d ago Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment arXiv:2609.26929v1 Announce Type: cross Abstract: People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective… 22 arXiv — NLP / Computation & Language research 4d ago Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis arXiv:2609.27756v1 Announce Type: cross Abstract: Large language models are increasingly asked to analyze data and report what the results mean, a task distinct from the belief- or preference-alignment settings studied in most sycophancy research. We test whether editorial… 13 arXiv — NLP / Computation & Language research 4d ago PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety arXiv:2609.28197v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn… 17 Ars Technica — AI news-outlet 4d ago China silent as US touts plan for AI safety alerts that omits tech experts Trump focus on winning “AI race” may deter China from sharing safety intel. 17 OpenAI official-blog 4d ago Sam Altman’s remarks at the United Nations Security Council OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council. 11 arXiv — NLP / Computation & Language research 5d ago "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It arXiv:2609.25021v1 Announce Type: cross Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the… 21 arXiv — Machine Learning research 5d ago PACT: From Credit Assignment to Critic Alignment arXiv:2609.26355v1 Announce Type: new Abstract: Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training… 21 arXiv — NLP / Computation & Language research 5d ago Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione arXiv:2609.25049v1 Announce Type: new Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely… 38 arXiv — NLP / Computation & Language research 5d ago LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping arXiv:2609.25051v1 Announce Type: new Abstract: Dual-source encrypted points of interest (DSEP), POIs from two encrypted coordinate systems, suffer from intertwined location and attribute uncertainties, including nonlinear systematic misalignment and naming inconsistency,… 20 arXiv — NLP / Computation & Language research 5d ago FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability arXiv:2609.25192v1 Announce Type: new Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition… 35 arXiv — NLP / Computation & Language research 5d ago Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation arXiv:2609.25755v1 Announce Type: new Abstract: Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao… 25 arXiv — NLP / Computation & Language research 5d ago Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding arXiv:2609.26399v1 Announce Type: new Abstract: Scene safety understanding plays a life-or-death role in situational awareness in various critical domains. Traditional methods that rely on learning direct mappings between scenes and safety levels often lack interpretability,… 14 arXiv — NLP / Computation & Language research 5d ago Calibration as a First-Class Criterion in LLM Evaluation arXiv:2609.26489v1 Announce Type: new Abstract: Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this… 9 arXiv — NLP / Computation & Language research 5d ago Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment arXiv:2609.23640v1 Announce Type: cross Abstract: Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We… 10 The Information — AI news-outlet 5d ago Anthropic releases cheaper model in first launch since slowdown calls Anthropic announced the release of its newest model, Claude Opus 5.5 , saying it performs as well or better than Anthropic’s previous most powerful models, Fable 5.1 and Mythos 5.1, in domains including coding, reasoning, business workflows, and safety. At the same time, it will… 18 TechCrunch — AI news-outlet 5d ago Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda At TechCrunch Disrupt 2026, five sessions across the AI Stage and Real World AI Stage cover AI safety, featuring leaders from Anthropic, NVIDIA, AWS, Waabi, and more. Register now to save up to $200 before Sept 25. 6 arXiv — Machine Learning research 6d ago Correcting Learning-based Perception for Safety arXiv:2609.22108v1 Announce Type: new Abstract: Learning-enabled perception is important in many autonomous systems. Unlike traditional sensors, the boundary where ML perception does or does not work is poorly characterized. Incorrect perception can lead to unsafe or overtly… 29 arXiv — Machine Learning research 6d ago SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs arXiv:2609.22153v1 Announce Type: new Abstract: Methods for addressing safety drift in fine-tuned Large Language Models (LLMs) are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, making them difficult to adopt and compare. We introduce… 29 arXiv — Machine Learning research 6d ago Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator arXiv:2609.22173v1 Announce Type: new Abstract: Mobility simulation supports logistics, safety, and communications planning in industrial environments such as ports, mines, and airports. Existing trajectory models, however, rely on absolute coordinates, road-network tokens, or… 20 arXiv — Machine Learning research 6d ago Beyond Task Completion: Training Capable and Safe Computer-Use Agents arXiv:2609.22178v1 Announce Type: new Abstract: Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must… 32 arXiv — Machine Learning research 6d ago The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families arXiv:2609.22216v1 Announce Type: new Abstract: Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8,… 37 arXiv — Machine Learning research 6d ago Computationally efficient safe exploration in reinforcement learning arXiv:2609.22919v1 Announce Type: new Abstract: Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes… 7 arXiv — NLP / Computation & Language research 6d ago Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs arXiv:2609.22144v1 Announce Type: new Abstract: Preserving safety alignment during large language models fine-tuning is critical, however, recent studies have demonstrated that even benign fine-tuning data may contain safety-degrading samples that silently undermine safety… 14 arXiv — NLP / Computation & Language research 6d ago From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models arXiv:2609.22224v1 Announce Type: new Abstract: A direction in activation space that changes safety-relevant behavior when steered is not necessarily one the model uses to produce that behavior on its own. We therefore ask whether steering acts through the computation of the… 34 arXiv — NLP / Computation & Language research 6d ago Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time arXiv:2609.22226v1 Announce Type: new Abstract: Decode-time alignment methods steer a frozen language model by scoring candidate continuations with an external reward and selecting the maximiser. We argue that this shared design is a single degenerate point in a much larger… 29 Page 1 of 10 · 500 articles Older →