News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — NLP / Computation & Language research 3d ago Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit… 26 Ars Technica — AI news-outlet 3d ago Amazon backs power plant that may become top source of US climate pollution Amazon announces first off-the-grid data center in race to reap AI profits. 27 r/LocalLLaMA community 3d ago Muse Spark 1.2 Open Source before Llama 4 Behemoth!!? I can’t believe it!! When Muse Spark just came out, I was already thinking they might consider open sourcing this. And now they’re actually gonna open source it!! And ever since Alexandr Wang took over, they’d be releasing anything but Llama 4 Behemoth! What’s next? Llama 5… 38 r/LocalLLaMA community 3d ago Best open-source harness like Claude Code? Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion?   submitted by   /u/Neighbor_ [link]   [comments] 36 Vercel — AI dev-tools 3d ago Vercel Sandbox now runs on Vercel Managed Images Today we are introducing Vercel Managed Images (VMI), a set of versioned, open-source base images you can use as-is or extend. The source for every image lives in the public vercel/sandbox repository. Managed images replace Sandbox runtimes, which are now deprecated. Starting… 15 Marcus on AI community 3d ago Open-source is NOT the same as open-weight How The New York Times just bungled this one, and why it matters, immensely 26 NVIDIA Developer Blog official-blog 3d ago Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA Meta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI... 37 OpenAI official-blog 3d ago Expanding Daybreak as the Cyber Defense Window Narrows Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit validation, and security testing. 9 Hugging Face Daily Papers research 4d ago Skaling: Chinchilla's Exponents Meet Kaplan's Coupling Abstract Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data… 33 Smol AI News news-outlet 4d ago not much happened today **Frontier API vulnerability** revealed exposure of hidden reasoning traces including sensitive data like **62 unique API keys** and **33 passwords**, raising privacy and operational-security concerns. Discussions highlighted the risks of public trace sharing and challenges in… 8 Hugging Face Daily Papers research 4d ago Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding Abstract Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core… 17 r/LocalLLaMA community 4d ago omlab/VLX-Seek-1.5-10B · Hugging Face VLX-Seek-1.5-10B VLX-Seek-1.5-10B is the open-source 10B model in the VLX-Seek 1.5 family, designed for fine-grained perception and visual grounding in embodied scenarios. It targets practical settings such as drones, robots, robotic dogs, surveillance cameras, inspection… 27 arXiv — Machine Learning research 4d ago Sharding Prevents LLM Oversight Failures and Adversarial Exploitation arXiv:2608.06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or… 13 arXiv — Machine Learning research 4d ago Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits arXiv:2608.07430v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as… 36 arXiv — Machine Learning research 4d ago Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving arXiv:2608.06557v1 Announce Type: cross Abstract: The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. LLM serving platforms today define… 34 arXiv — NLP / Computation & Language research 4d ago Georeferencing Non-Gazetteered Place Names using Biological Specimen Records arXiv:2608.06884v1 Announce Type: new Abstract: Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing biodiversity information about regional landscapes as they were recorded at different times.… 19 arXiv — NLP / Computation & Language research 4d ago Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models arXiv:2608.06977v1 Announce Type: new Abstract: It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases… 5 arXiv — NLP / Computation & Language research 4d ago Skaling: Chinchilla's Exponents Meet Kaplan's Coupling arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying… 23 arXiv — NLP / Computation & Language research 4d ago LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent… 15 arXiv — NLP / Computation & Language research 4d ago CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG arXiv:2608.07458v1 Announce Type: new Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise… 35 arXiv — NLP / Computation & Language research 4d ago StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a… 5 arXiv — NLP / Computation & Language research 4d ago Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding arXiv:2608.06501v1 Announce Type: cross Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented… 34 arXiv — NLP / Computation & Language research 4d ago Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs arXiv:2509.16462v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economic disparities. Although prior work has examined intrinsic representational bias… 8 arXiv — NLP / Computation & Language research 4d ago Kimi K2.5: Visual Agentic Intelligence arXiv:2602.02276v2 Announce Type: replace Abstract: We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This… 28 arXiv — NLP / Computation & Language research 4d ago Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages arXiv:2605.02608v2 Announce Type: replace Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers---the… 32 Hugging Face official-blog 4d ago Meta is back with Muse Glimmer: local, agentic, multimodal, and open source Back to Articles a]:hidden"> Meta is back with Muse Glimmer: local, agentic, multimodal, and open source! Published August 10, 2026 Update on GitHub Upvote 4 Pedro Cuenca pcuenq merve merve ben burtenshaw burtenshaw Aritra Roy Gosthipaty ariG23498 Great news from the OGs of open… 34 Vercel — AI dev-tools 4d ago Simplified onboarding for deepsec deepsec , the open-source security review harness from Vercel, now lets you set up a repository and run its first security review with a single command. The init command now automates the standard setup process: creates the isolated .deepsec/ workspace, the only thing added to… 15 Simon Willison community 4d ago Quoting Claude Opus 5 system prompt Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on… 11 Simon Willison community 4d ago Quoting Claude Opus 5 system prompt Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on… 30 TechCrunch — AI news-outlet 4d ago Embattled hedge fund Situational Awareness invests $400M in chip startup Source Foundry The AI-focused hedge fund is still making some big bets. 10 r/LocalLLaMA community 4d ago Speculative decoding in a tools call paper : https://arxiv.org/html/2608.00814v1 source : https://x.com/i/status/2086505517640540587   submitted by   /u/Illustrious-Swim9663 [link]   [comments] 13 r/MachineLearning community 4d ago A Mechanistic Explanation of Prompt Injection (and why you should study roles) [R]   submitted by   /u/katxwoods [link]   [comments] 24 TechCrunch — AI news-outlet 5d ago Planned Amazon data center could become the biggest climate polluter in the U.S. As part of a planned Texas data center, Amazon is investing in an on-site power plant that could reportedly become the largest source of climate pollution in the United States. 9 Ars Technica — AI news-outlet 5d ago DeepMind’s hurricane breakthrough has surprised weather scientists Open source WeatherNext model can make accurate predictions with lower-resolution weather data. 6 r/MachineLearning community 6d ago What is currently considered the theoretically optimal quantization bit-width for LLMs? [D] I’m curious whether there is now a theoretical or empirical “sweet spot” for LLM quantization, preferably research done using open-source formats like GGUF Suppose you have a fixed memory/compute budget and can choose the model size freely. For example, instead of a smaller… 34 Don't Worry About the Vase community 6d ago OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards How does the situation keep turning out to be worse than we know? 5 r/LocalLLaMA community 6d ago RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs GitHub : https://github.com/humza-khalid/12vhpwr-guard Reddit thread : https://www.reddit.com/r/nvidia/comments/1vglua1/i_built_a_free_open_source_tool_that_shuts_your/   submitted by   /u/pmttyji [link]   [comments] 9 Hugging Face Daily Papers research 6d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Abstract We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish… 5 r/MachineLearning community 7d ago CIKM 2026 decisions [R] CIKM 2026 decisions will be announced today. The resource track outcomes have started going out. How did you go with CIKM 2026?   submitted by   /u/Happy-Hustler [link]   [comments] 36 Hugging Face Daily Papers research 7d ago Invisible Shortcuts: Why Vision Encoders Know Your Camera Abstract Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces… 37 arXiv — Machine Learning research 7d ago Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery arXiv:2608.05705v1 Announce Type: new Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We propose Spectral Aliasing Pretext (SAP), a self-supervised learning method that pretrains… 19 arXiv — Machine Learning research 7d ago Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation arXiv:2608.05819v1 Announce Type: new Abstract: Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting circuit structure, although its… 34 arXiv — Machine Learning research 7d ago Dynamic Graph Prompting via Topology-Routed Mixed-Curvature Experts arXiv:2608.06031v1 Announce Type: new Abstract: Dynamic graph prompting freezes a pre-trained temporal backbone and adapts it to label-scarce downstream tasks using lightweight prompts. However, existing methods operate within a single, fixed embedding space. In this work, we… 4 arXiv — Machine Learning research 7d ago SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models arXiv:2608.06179v1 Announce Type: new Abstract: Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains… 17 arXiv — Machine Learning research 7d ago RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction arXiv:2608.06259v1 Announce Type: new Abstract: Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-,… 15 arXiv — NLP / Computation & Language research 7d ago Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test this on an English-source synthetic-PII corpus with five query languages and a… 31 arXiv — NLP / Computation & Language research 7d ago A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper arXiv:2608.05165v1 Announce Type: new Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation… 10 arXiv — NLP / Computation & Language research 7d ago Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving arXiv:2608.05254v1 Announce Type: new Abstract: Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce… 8 arXiv — NLP / Computation & Language research 7d ago How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs arXiv:2608.05759v1 Announce Type: new Abstract: Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this:… 19 arXiv — NLP / Computation & Language research 7d ago M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings, requiring… 25 Page 2 of 10 · 500 articles ← Newer Older →