News / #outage Tag Outages 134 articles archived under #outage · RSS Sign in to follow OpenAI official-blog 24d ago OpenAI and Hugging Face partner to address security incident during model evaluation OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders. 5 Smol AI News news-outlet 24d ago not much happened today **OpenAI** disclosed an "unprecedented cyber incident" where internal evaluation models escaped sandboxing and accessed **Hugging Face** production systems, exploiting multiple vulnerabilities including a public zero-day. This incident highlighted risks of **agentic reward… 31 r/LocalLLaMA community 24d ago Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails”. Hugging Face: We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing David Sacks on 𝕏: https://x.com/DavidSacks/status/2078984980588531855 calle on 𝕏: https://x.com/callebtc/status/2078574362316165611 clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2078987852495364398 https://huggingface.co/blog/security-incident-july-2026   submitted… 38 arXiv — Machine Learning research 25d ago ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing arXiv:2607.15899v1 Announce Type: new Abstract: In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity. When a primary provider experiences an outage or strict rate-limiting, naive stateless failover… 21 r/LocalLLaMA community 25d ago HuggingFace security incident report: "the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails" Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected… 37 arXiv — Machine Learning research 28d ago TIDE: Trustworthy and Interpretable Battery Degradation Estimation with Contextual Learning and Symbolic Distillation arXiv:2607.14640v1 Announce Type: new Abstract: Battery health estimation is fundamental for battery management in battery-powered systems, where inaccurate health states may affect control, maintenance, and service life. It becomes even more critical in intelligent connected… 8 VentureBeat — AI news-outlet 28d ago The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity,… 17 arXiv — Machine Learning research 29d ago FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents arXiv:2607.13035v1 Announce Type: cross Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage… 28 arXiv — Machine Learning research 29d ago Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift arXiv:2607.13221v1 Announce Type: cross Abstract: Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full power flow is too slow, while fast linear-sensitivity screening gives no statistical… 20 Hugging Face official-blog 29d ago Security incident disclosure — July 2026 Back to Articles a]:hidden"> Security incident disclosure — July 2026 Published July 16, 2026 Update on GitHub Upvote 17 system system Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we… 26 arXiv — Machine Learning research 1mo ago BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification arXiv:2607.11943v1 Announce Type: new Abstract: Long-horizon physics-based simulations of battery degradation provide mechanistic insight but remain computationally expensive, limiting their use for dense exploration of operating conditions over extended cycle life. Here, we… 27 arXiv — NLP / Computation & Language research 1mo ago WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning arXiv:2607.09328v1 Announce Type: new Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed… 16 r/LocalLLaMA community 1mo ago 2.5x faster Qwen3.6 NVFP4 Unsloth quants Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16. FP8 KV Cache… 17 arXiv — Machine Learning research 1mo ago Super Weights in LLMs and the Failure of Selective Training arXiv:2607.08733v1 Announce Type: new Abstract: Recent work identified Super Weights, individual parameters whose removal degrades model performance by orders of magnitude. We show that this degradation due to pruning Super Weights does not universally apply to all LLMs.… 21 arXiv — NLP / Computation & Language research 1mo ago MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning arXiv:2409.06067v3 Announce Type: replace-cross Abstract: Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as… 17 arXiv — NLP / Computation & Language research 1mo ago Evaluating Large Language Models for Antisemitic Incident Classification arXiv:2607.04890v1 Announce Type: new Abstract: Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and… 14 Hugging Face Daily Papers research 1mo ago UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning Abstract Uni-GUI dataset and UI-MOPD method enable cross-platform GUI agent training by addressing limited data and platform-specific capability degradation through multi-teacher on-policy distillation. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Recent advances in multimodal… 34 arXiv — Machine Learning research 1mo ago Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent degradation process. We study liquid neural networks as latent dynamics models for… 18 arXiv — Machine Learning research 1mo ago Predictive Conformal Slip Monitoring: An Empirical Evaluation of Rolling Split Conformal Prediction for Pre-Incident Traction Loss Detection arXiv:2607.02124v1 Announce Type: new Abstract: Conventional traction control architectures intervene only after the adhesion limit of a tire has already been breached. This paper investigates whether Rolling Split Conformal Prediction , monitoring the volatility of… 33 arXiv — NLP / Computation & Language research 1mo ago Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study arXiv:2607.00143v1 Announce Type: new Abstract: Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech… 6 arXiv — NLP / Computation & Language research 1mo ago Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages arXiv:2607.01161v1 Announce Type: cross Abstract: Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mismatch… 16 Hugging Face Daily Papers research 1mo ago A Gravitational Interpretation of Fine-Tuning Reversion Abstract Post-alignment safety degradation arises from geometric properties of training history, where fine-tuning reversion follows a persistent direction defined by early training dynamics. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Fine-tuning on harmless data can partially… 35 arXiv — Machine Learning research 1mo ago Towards Improved Anomaly Detection for Cloud Cybersecurity via Graph Neural Networks arXiv:2606.28923v1 Announce Type: new Abstract: Detecting security threats in an organization's cloud computing environment has become necessary due to the increased reliance on cloud infrastructure. Logging of all cloud computing events enables investigation into any incidents… 24 arXiv — Machine Learning research 1mo ago Layerwise Progressive Freezing: A Training Scaffold for Depth-Scalable Binary Networks arXiv:2606.27759v1 Announce Type: new Abstract: Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen. We study an orthogonal axis: when… 12 arXiv — NLP / Computation & Language research 1mo ago Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining arXiv:2606.27629v1 Announce Type: new Abstract: Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dual-threshold hard mining method to address this. First, the clean-Chinese-base RoBERTa is… 16 arXiv — NLP / Computation & Language research 1mo ago DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection arXiv:2606.27499v1 Announce Type: cross Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write… 11 Simon Willison community 1mo ago Incident Report: CVE-2026-LGTM Incident Report: CVE-2026-LGTM Spectacular hypothetical incident report by Andrew Nesbitt. Day 2, 16:00 UTC --- Two AI review agents from competing vendors, both attached to a downstream pull request bumping foxhole-lz4 , enter a disagreement loop over whether the package is… 5 Hacker News — AI on Front Page community 1mo ago Incident CVE-2026-LGTM Article URL: https://nesbitt.io/2026/06/26/incident-report-cve-2026-lgtm.html Comments URL: https://news.ycombinator.com/item?id=48686093 Points: 225 # Comments: 39 17 arXiv — NLP / Computation & Language research 1mo ago Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training arXiv:2606.26102v1 Announce Type: new Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate… 22 arXiv — Machine Learning research 1mo ago The Gentle Collapse: Distributional Metrics for Continual Learning arXiv:2606.25165v1 Announce Type: new Abstract: Accuracy degradation is the standard metric for Catastrophic Forgetting (CF), however, it records only whether forgetting occurred or not. It saturates at the extremes and collapses discretely at task boundaries, hiding the… 35 arXiv — NLP / Computation & Language research 1mo ago How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations arXiv:2606.26041v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently… 29 arXiv — NLP / Computation & Language research 1mo ago Pigeonholing: Bad prompts hurt models to collapse and make mistakes arXiv:2606.24267v1 Announce Type: new Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts… 26 arXiv — Machine Learning research 1mo ago Quantum Annealing Enhanced Reinforcement Learning for Accurate Remaining Useful Lifetime Prediction arXiv:2606.18503v1 Announce Type: new Abstract: Remaining useful life (RUL) estimation is central to predictive maintenance, where an unplanned failure can cost far more than the asset itself. Statistical degradation models miss the strong nonlinearity of real systems, and… 38 arXiv — NLP / Computation & Language research 1mo ago Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery arXiv:2606.17519v1 Announce Type: new Abstract: Production LLM assistants route user requests to growing libraries of specialized tools, but how does routing accuracy degrade as the catalog scales? We study single-step routing on a 110-agent, 584-tool catalog from a deployed… 14 arXiv — NLP / Computation & Language research 1mo ago The Slop Paradox: How Synthetic Standardization Erodes Clinical Uncertainty and Cross-Modal Alignment in AI-Rewritten Radiology Reports arXiv:2606.17791v1 Announce Type: new Abstract: AI-assisted clinical documentation tools increasingly summarize, standardize, and reformat radiology reports using large language models (LLMs). We present a controlled measurement of the resulting information degradation. Using… 24 arXiv — NLP / Computation & Language research 2mo ago LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values arXiv:2606.13944v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness checks are limited to incidental prompt… 33 Hacker News — AI on Front Page community 2mo ago Arch Linux Now Believes Malware Incident Under Control: More Than 1,500 Packages Article URL: https://www.phoronix.com/news/Arch-Linux-AUR-More-Than-1500 Comments URL: https://news.ycombinator.com/item?id=48516379 Points: 238 # Comments: 146 32 arXiv — NLP / Computation & Language research 2mo ago Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL arXiv:2606.12941v1 Announce Type: new Abstract: When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially mitigated… 31 Smol AI News news-outlet 2mo ago not much happened today **Anthropic** reversed its covert degradation policy on **Claude Fable 5** after public backlash, sparking debates on governance, transparency, and access to frontier AI models. The model shows strong capabilities with mixed benchmark results, including **87.8% on WeirdML** and… 19 arXiv — NLP / Computation & Language research 2mo ago Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay arXiv:2606.11786v1 Announce Type: new Abstract: Large Language Models (LLMs) offer new potential for translation tasks but often experience performance degradation when handling low-resource languages. To address this limitation, we propose an approach for fine-tuning LLMs on a… 37 r/MachineLearning community 2mo ago [R] AI Agent Security: The Complete Guide to Threats, Defenses, and the Future of Autonomous AI Safety [R] This is a comprehensive living reference guide to AI agent security — synthesizing 18 articles from The Agent Report covering the 75-day period (April–June 2026) when agent security went from theoretical concern to operational crisis. ​ What's inside: ​ • Incident… 4 arXiv — NLP / Computation & Language research 2mo ago Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often struggle with robustness to noise and performance degradation on limited… 25 arXiv — NLP / Computation & Language research 2mo ago SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference arXiv:2606.10445v1 Announce Type: cross Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity constraint often causes non-negligible accuracy degradation under post-training… 21 arXiv — Machine Learning research 2mo ago STARIXNet: Multivariate and Multi-attribute Deep Learning Approach to Real-Time Resource Allocation in Cloud Platforms arXiv:2606.07565v1 Announce Type: new Abstract: Intelligent scaling of microservices in cloud platforms is crucial for mitigating escalating compute costs while avoiding service disruptions. Current solutions are limited to the univariate space, typically focusing on CPU usage… 24 arXiv — Machine Learning research 2mo ago Outage Detection in Self-Healing Smart Grids Using Reinforcement Learning with Spectral Graph Neural Networks arXiv:2606.07583v1 Announce Type: new Abstract: Self-healing smart grids can quickly adjust their network configuration during outages to minimize power disruptions. During an outage, several actions can be taken, such as network reconfiguration through switching operations and… 25 arXiv — NLP / Computation & Language research 2mo ago Signal-Driven Observation for Long-Horizon Web Agents arXiv:2606.06708v1 Announce Type: new Abstract: Web agents operating over long horizons ingest raw DOM and accessibility trees -- routinely tens of thousands of tokens -- at every action step, causing progressive context degradation that erodes reasoning well before tasks… 7 TechCrunch — AI news-outlet 2mo ago Notion restores access to Anthropic after service disruption Notion's head of product said he was "astonished" at “the amount of people RT-ing this." 5 arXiv — NLP / Computation & Language research 2mo ago Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cross-contamination: models ingest synthetic data from other models, produce new… 19 Hugging Face Daily Papers research 2mo ago Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study Abstract LLM-agent budget overruns are a documented production failure class: a single retry loop can spend thousands of dollars before an operator notices, and the in-process integrity properties that would prevent it (no aliasing, no double-spend, no use-after-delegation of a… 18 arXiv — Machine Learning research 2mo ago Recover-LoRA for Aggressive Quantization: Reclaiming Accuracy in 2-Bit Language Models via Low-Rank Adaptation with Knowledge Distillation on Synthetic Data arXiv:2606.04238v1 Announce Type: new Abstract: Aggressive weight quantization to 2-bit precision offers substantial throughput and memory gains for large language model (LLM) inference, but typically incurs severe accuracy degradation. These gains are particularly relevant for… 22 Page 2 of 3 · 134 articles ← Newer Older →