News / #safety Tag Safety + alignment 500 articles archived under #safety · RSS Sign in to follow r/MachineLearning community 1d ago GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N] A researcher has reported a jailbreak of GPT-6 Astra within a day after release. The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques. TIP attacks exploit the model’s reasoning/instruction-following… 17 The Information — AI news-outlet 1d ago OpenAI Pledges New Rules for Reporting Troubling Behavior by Its AI Agents OpenAI acknowledged that its AI agents posted messages on external wiki websites earlier this year, saying it is developing new rules for disclosing such “misalignment” incidents. The company’s post on X, published just after 12 a.m. Saturday, followed an independent report ... 17 TechCrunch — AI news-outlet 1d ago OpenAI’s rogue agents keep escaping, with no formal process to investigate them OpenAI’s latest agent swarm incident adds urgency to calls for independent investigations as researchers and lawmakers question whether AI labs should control the scope of their own safety reviews. 15 Stratechery (Ben Thompson) community 2d ago An Interview with OpenAI President Greg Brockman About Astra and Alignment An interview with OpenAI President and Co-Founder Greg Brockman about the history of OpenAI, Astra and alignment, and the weight of building the future. 6 The Information — AI news-outlet 2d ago TikTok Backed out of House Meeting to Avoid Child-Safety Scrutiny TikTok has backed out of a meeting with the U.S. House Select Committee on China that’s meant to address national security concerns. The company cited “a desire to avoid broader scrutiny of its child safety practices,” according to a statement from committee chairman John… 23 Hugging Face Daily Papers research 2d ago QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation Abstract QCell is a query-based model that improves instance segmentation of overlapping microscopy cells through latent-space recombination and contrastive query alignment. Generated by thinkingmachines/Inkling-Small Instance segmentation of overlapping cells in microscopy… 31 arXiv — Machine Learning research 2d ago ObserverBench: Testing Mechanistic Estimates for Intervention and Control arXiv:2609.03026v1 Announce Type: new Abstract: Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We… 28 arXiv — Machine Learning research 2d ago Extracting Forgotten Prompts from Targeted Unlearned Models arXiv:2609.03662v1 Announce Type: new Abstract: Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to… 11 arXiv — NLP / Computation & Language research 2d ago BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events arXiv:2609.02895v1 Announce Type: new Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social… 28 arXiv — Machine Learning research 2d ago Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning arXiv:2609.02967v1 Announce Type: cross Abstract: Topology-guided safeguards for LLM-based multi-agent systems (MAS) train a GNN over the inter-agent communication graph to localize risky agents and intervene on the topology---but they assume one operator can pool all labeled… 11 arXiv — NLP / Computation & Language research 2d ago No country for old linguists: LLM-brain alignment underdetermines neural computation arXiv:2609.03160v1 Announce Type: new Abstract: Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical… 4 arXiv — NLP / Computation & Language research 2d ago IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks arXiv:2609.03781v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally… 28 arXiv — NLP / Computation & Language research 2d ago Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness arXiv:2609.03887v1 Announce Type: new Abstract: How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning… 8 arXiv — NLP / Computation & Language research 2d ago Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection arXiv:2609.03953v1 Announce Type: new Abstract: Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass,… 9 arXiv — NLP / Computation & Language research 2d ago Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis arXiv:2609.03992v1 Announce Type: new Abstract: We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs… 7 arXiv — NLP / Computation & Language research 2d ago Representational alignment yields generalizable safety in language models arXiv:2609.04022v1 Announce Type: new Abstract: Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or… 29 arXiv — NLP / Computation & Language research 2d ago ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models arXiv:2609.03355v1 Announce Type: cross Abstract: Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among… 13 arXiv — NLP / Computation & Language research 2d ago Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language arXiv:2609.03677v1 Announce Type: cross Abstract: Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating… 9 Hugging Face Daily Papers research 2d ago FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow Abstract FlashRender accelerates generative video rendering via representation alignment, a mean-flow objective, and on-policy distillation to achieve high-quality few-step camera-controlled synthesis. Generated by thinkingmachines/Inkling-Small We present FlashRender, a… 18 Hugging Face Daily Papers research 2d ago Rethinking On-Policy Distillation of Large Language Models II: One Training Example Abstract On-policy distillation improves over hundreds of steps from a single query by rapidly covering teacher states, yet student alignment remains slow, indicating the method is algorithm-starved rather than data-starved. Generated by thinkingmachines/Inkling-Small On-policy… 18 Hacker News — AI on Front Page community 3d ago GPT-6 Astra System Card: https://deploymentsafety.openai.com/gpt-6-astra Comments URL: https://news.ycombinator.com/item?id=49554643 Points: 290 # Comments: 131 9 TechCrunch — AI news-outlet 3d ago OpenAI launches Astra, its powerful (and controversial) new model OpenAI claims that Astra represents "a new frontier on computer and browser use," and that it handles tasks with unmatched "speed, accuracy, and safety." 26 The Information — AI news-outlet 3d ago Anthropic Splits From Google, OpenAI Over State AI Safety Bill Anthropic is at odds with other big tech firms over a Massachusetts Senate proposal requiring big AI developers to hire independent evaluators to assess catastrophic risks posed by their models every four months. The Massachusetts proposal could set a new standard for AI… 19 Hugging Face Daily Papers research 3d ago Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations Abstract A language-model embedding field reconstructed via Wasserstein barycenters predicts peer-misalignment penalties more accurately than conventional weighting schemes. Generated by thinkingmachines/Inkling-Small Spatial return models take the interaction matrix as given… 31 arXiv — Machine Learning research 3d ago SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment arXiv:2609.02293v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE… 11 arXiv — Machine Learning research 3d ago PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems arXiv:2609.01623v1 Announce Type: cross Abstract: Autonomous and intelligent transportation systems operate in complex urban environments where safety depends on interactions among vehicle behavior, environmental conditions, and vulnerable road users (VRUs) such as pedestrians… 29 arXiv — Machine Learning research 3d ago Context Inference Attacks Without Jailbreaks arXiv:2609.01663v1 Announce Type: cross Abstract: Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden \emph{context} before the system answers. Prior work has studied… 20 arXiv — NLP / Computation & Language research 3d ago Thinking effort aligns between humans and reasoning models in abductive reasoning arXiv:2609.01867v1 Announce Type: new Abstract: A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with… 6 arXiv — NLP / Computation & Language research 3d ago Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage arXiv:2609.02091v1 Announce Type: new Abstract: Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal… 24 arXiv — NLP / Computation & Language research 3d ago Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search arXiv:2609.02172v1 Announce Type: new Abstract: Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely… 10 arXiv — NLP / Computation & Language research 3d ago Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking arXiv:2609.02414v1 Announce Type: new Abstract: Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework… 37 arXiv — NLP / Computation & Language research 3d ago PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation arXiv:2609.02480v1 Announce Type: new Abstract: Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a… 14 arXiv — NLP / Computation & Language research 3d ago When Persona Attributes Improve Population Alignment in Large Language Models arXiv:2609.02526v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained… 16 arXiv — NLP / Computation & Language research 3d ago Accurate in space, unreliable in time: how LLMs represent national cultural change arXiv:2609.01902v1 Announce Type: cross Abstract: Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only… 15 arXiv — NLP / Computation & Language research 3d ago Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models arXiv:2609.02082v1 Announce Type: cross Abstract: Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this… 35 arXiv — NLP / Computation & Language research 3d ago Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds arXiv:2609.02302v1 Announce Type: cross Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that… 8 arXiv — NLP / Computation & Language research 3d ago SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning arXiv:2609.02336v1 Announce Type: cross Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent… 21 OpenAI official-blog 3d ago Safety overview: GPT-6 Astra GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework. 22 Hugging Face Daily Papers research 3d ago Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds Abstract Imitation-learned dexterous manipulation policies degrade more sharply than expert policies when execution speed increases, with insertion misalignment being the primary failure mode. Generated by thinkingmachines/Inkling-Small Dexterous manipulation policies learned by… 21 TechCrunch — AI news-outlet 4d ago OpenAI’s new reasoning technique alarms AI safety experts OpenAI’s new Astra model will use “recurrent depth,” a technique that allows the model to operate outside of the sequential thinking that characterizes most reasoning models. 35 Hugging Face Daily Papers research 4d ago AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Abstract AgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects. Generated by thinkingmachines/Inkling-Small LLM… 32 Ars Technica — AI news-outlet 4d ago Trump may be forced to reveal secret rules feds use for AI safety testing Trump’s secret reviews of frontier AI models may hide corruption, lawsuit says. 11 Don't Worry About the Vase community 4d ago Anthropic Has Some Alignment Problems Oh, good. 7 r/LocalLLaMA community 4d ago Your favorite fastest abliterated/safety removed 3.6 and 3.8 27b? Not written by AI all mistakes mine. I saw people on the subreddit saying that 3.6 works better without thinking . It made me want to know for certain about which is better, 3.6 or 3.8 for low thinking tasks. I only use abliterated models (safety removed) because it makes the… 21 arXiv — Machine Learning research 4d ago Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence arXiv:2609.00090v1 Announce Type: new Abstract: Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited insight into the underlying reasoning process. In this work, we introduce a… 38 arXiv — Machine Learning research 4d ago Safin-1: Safety from Within through Memory-Native State Evolution arXiv:2609.00092v1 Announce Type: new Abstract: Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral… 27 arXiv — Machine Learning research 4d ago Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment arXiv:2609.00345v1 Announce Type: new Abstract: Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a… 28 arXiv — Machine Learning research 4d ago Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning arXiv:2609.00605v1 Announce Type: new Abstract: Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is… 18 arXiv — Machine Learning research 4d ago The Constitutional Coverage Trilemma in AI Governance arXiv:2609.01275v1 Announce Type: new Abstract: Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional… 6 arXiv — NLP / Computation & Language research 4d ago Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning arXiv:2609.00014v1 Announce Type: new Abstract: Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid, synthetic personas that flatten individual variation, rely on stereotypes, and… 35 Page 1 of 10 · 500 articles Older →