News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow r/MachineLearning community 1d ago GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N] A researcher has reported a jailbreak of GPT-6 Astra within a day after release. The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques. TIP attacks exploit the model’s reasoning/instruction-following… 17 r/LocalLLaMA community 1d ago Any resource on using Blender with local models, and which models work best? Hey all, I've seen some really fun looking things with people having their local models drive Blender to create pretty cool looking world scenes. Is there a good tutorial on setting up Blender yo be driven by your model? For example, what programming harness, do you use a MCP… 22 r/LocalLLaMA community 1d ago The gap has closed, open source will win I've been trying the latest models from the frontier labs and honestly, after extensive testing I can not tell the difference between the best open source options. I think the differences are now marginal but the labs are doing heavy marketing to convince the public into paying… 36 Hacker News — AI on Front Page community 1d ago Actively exploited sandbox RCE in all Chromium versions Article URL: https://nvd.nist.gov/vuln/detail/cve-2026-85046 Comments URL: https://news.ycombinator.com/item?id=49570669 Points: 236 # Comments: 135 9 Hacker News — AI on Front Page community 2d ago Show HN: Open-Source eInk Bike Computer Hey all, i just launched my Eink Bike computer project and think it is cool. Another tidbit, in the crazy things that AI has done... It has helped create a ANT (common sensor wireless protocol used in workout/biking) implementation for ESP32 by messing around with undocumented… 37 Hacker News — AI on Front Page community 2d ago Corporate America is getting hooked on open-source AI https://archive.is/kmOqm Comments URL: https://news.ycombinator.com/item?id=49566137 Points: 216 # Comments: 197 24 r/LocalLLaMA community 2d ago I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked? This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe… 16 The Information — AI news-outlet 2d ago Saudi Firm Humain Unveils Arabic Language Model Developed With China’s MiniMax Humain, Saudi Arabia’s state-owned AI company, announced on Thursday an Arabic large language model developed based on Chinese AI firm MiniMax’s model. The humain-m3 model, built on MiniMax’s M3 open-source model, was pre-trained on more than 1 trillion tokens of Arabic content,… 30 r/LocalLLaMA community 2d ago We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0) I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public. It's an inference engine in Rust and C++… 25 arXiv — Machine Learning research 2d ago LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference arXiv:2609.03079v1 Announce Type: new Abstract: On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload… 36 arXiv — Machine Learning research 2d ago FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience arXiv:2609.03241v1 Announce Type: new Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or… 36 arXiv — NLP / Computation & Language research 2d ago From Zero to Hero: An Open LLM Ecosystem for Armenian arXiv:2609.03350v1 Announce Type: cross Abstract: Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release… 31 arXiv — Machine Learning research 2d ago TraveL: Transformer-based Multi-view Path Distributional Representation Learning arXiv:2609.03427v1 Announce Type: new Abstract: Path representation learning (PRL) for road networks has received increasing research attention, due to various path-related applications. Existing works on PRL typically exploit the co-occurrence relationship among road segments… 10 arXiv — Machine Learning research 2d ago Beyond Straightness: Non-Crossing Flow Matching via Quantile AlignTree Coupling arXiv:2609.03443v1 Announce Type: new Abstract: The performance of Flow Matching largely depends on the quality of the coupling between the source and target distributions. However, independent coupling often leads to path crossings and local velocity ambiguity, while OT-based… 9 arXiv — Machine Learning research 2d ago A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds arXiv:2609.03457v1 Announce Type: new Abstract: Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches… 38 arXiv — Machine Learning research 2d ago RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models arXiv:2609.04007v1 Announce Type: new Abstract: Despite strong performance on held-out electroencephalography (EEG) data, seizure detectors may fail under real-world acquisition variability, artifacts, and adversarial inputs. We introduce RobustSeiz, an open-source,… 37 arXiv — Machine Learning research 2d ago LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games arXiv:2609.02931v1 Announce Type: cross Abstract: Scripted and rule-based non-player characters (NPCs) in combat video games often exhibit predictable behaviors that experienced players can exploit, while reinforcement learning (RL) agents typically retain a fixed policy after… 11 arXiv — Machine Learning research 2d ago Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction arXiv:2609.02969v1 Announce Type: cross Abstract: We present, to the best of our knowledge, the first publicly available resource for predicting left ventricular ejection fraction (EF) from parasternal long-axis (PLAX) echocardiography. Because no PLAX-EF datasets previously… 23 arXiv — NLP / Computation & Language research 2d ago Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers arXiv:2609.03273v1 Announce Type: new Abstract: Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct… 27 arXiv — NLP / Computation & Language research 2d ago To What Extent Do Large Language Models Understand Bangla Idioms? arXiv:2609.03410v1 Announce Type: new Abstract: Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first… 21 arXiv — NLP / Computation & Language research 2d ago Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech arXiv:2609.03502v1 Announce Type: new Abstract: In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a… 33 arXiv — NLP / Computation & Language research 2d ago Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation arXiv:2609.03619v1 Announce Type: new Abstract: Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as… 23 arXiv — NLP / Computation & Language research 2d ago Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks arXiv:2609.03734v1 Announce Type: new Abstract: BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit… 30 arXiv — NLP / Computation & Language research 2d ago IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks arXiv:2609.03781v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally… 28 arXiv — NLP / Computation & Language research 2d ago Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation arXiv:2609.04048v1 Announce Type: new Abstract: Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic… 9 arXiv — NLP / Computation & Language research 2d ago Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory arXiv:2609.03450v1 Announce Type: cross Abstract: An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both.… 19 Hugging Face Daily Papers research 2d ago An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems Abstract Non-conversational domain signals can generate synthetic dialogue data that outperforms zero-shot and scarce real-data baselines for bootstrapping conversational recommender systems. Generated by thinkingmachines/Inkling-Small Conversational Recommender Systems (CRS)… 19 Hugging Face Daily Papers research 2d ago Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Abstract Fully quantizing hybrid LLMs—including recurrent Gated DeltaNet layers—to 4-bit NVFP4 preserves accuracy across long-context and reasoning benchmarks by localizing outliers and exploiting robust delta-rule dynamics. Generated by thinkingmachines/Inkling-Small Hybrid… 30 llama.cpp releases dev-tools 2d ago b10793 llama: fix whole source code rebuilt on each new commit ( #28278 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/45114012 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)… 22 r/LocalLLaMA community 3d ago Introducing Quartermaster, an open source local AI platform designed for ease of use that does not sacrifice customizability It started as a fork of llama-swap , but I have been building it out for myself since then as a convenient tool for all my local AI needs, and by now it has drifted far enough to be its own thing. The main idea is that you point it at your models folder and it configures things… 6 arXiv — Machine Learning research 3d ago SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework arXiv:2609.02203v1 Announce Type: new Abstract: Time series representation learning (TSRL) has attracted growing research interests in recent years. Two recent explorations in TSRL are: i) exploiting a transformer-based framework to learn time series; ii) instead of using only… 29 arXiv — Machine Learning research 3d ago DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models arXiv:2609.02468v1 Announce Type: new Abstract: We explore predicting eCommerce user preferences for product aspects such as brand, size, and color - a task we define as Aspect Affinity. Solving this task improves customer understanding and enables fine-grained personalization… 10 arXiv — Machine Learning research 3d ago RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection arXiv:2609.02497v1 Announce Type: new Abstract: Zero-shot graph anomaly detection seeks to deploy a detector trained on source graphs to unseen, unlabeled targets, yet domain shift can make source-derived notions of normality unreliable. We introduce RINSE (Robust Iterative… 13 arXiv — Machine Learning research 3d ago Source Distribution Estimation by Posterior Averaging arXiv:2609.02622v1 Announce Type: new Abstract: Simulation-based science often requires a distribution over simulator parameters whose push-forward reproduces a set of real observations: this is the source distribution estimation (SDE) problem. Existing methods fit the source… 4 arXiv — Machine Learning research 3d ago Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models arXiv:2609.02638v1 Announce Type: new Abstract: Knowledge graphs have become an important source of structured knowledge for Web applications, including search, question answering, and recommender systems. In these applications, link prediction can serve either as a prediction… 38 arXiv — Machine Learning research 3d ago H3DNAS: Hardware-Aware ONNX-Native 3D Point Cloud Model Compression arXiv:2609.02684v1 Announce Type: new Abstract: Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained by compute and memory budgets. Existing compression methods require access to the model's original source code, rendering… 7 arXiv — Machine Learning research 3d ago When Literature Data Mislead Artificial Intelligence in Materials Discovery arXiv:2609.01621v1 Announce Type: cross Abstract: Artificial intelligence (AI) increasingly treats scientific literature as a data source for building databases, training predictive models, and guiding discovery. Yet literature-derived datasets often assume that reported… 8 arXiv — Machine Learning research 3d ago Marginal Expected Revenue for Jointly Ranking Auction and Fixed-Price Listings in E-Commerce Sponsored Search arXiv:2609.01628v1 Announce Type: cross Abstract: E-commerce search ranking must balance multiple objectives--relevance, user engagement, and platform revenue--when allocating impression slots to competing listings. Estimating the expected revenue component is well understood… 11 arXiv — NLP / Computation & Language research 3d ago SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition arXiv:2609.01737v1 Announce Type: new Abstract: Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a… 26 arXiv — NLP / Computation & Language research 3d ago Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search arXiv:2609.02172v1 Announce Type: new Abstract: Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely… 10 arXiv — NLP / Computation & Language research 3d ago Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking arXiv:2609.02414v1 Announce Type: new Abstract: Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework… 37 arXiv — NLP / Computation & Language research 3d ago HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks arXiv:2609.02772v1 Announce Type: new Abstract: Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target author using only a few reference examples while preserving the original meaning. Existing methods often struggle to achieve… 33 The Information — AI news-outlet 3d ago Moonshot AI Files Confidentially for Hong Kong IPO Moonshot AI, whose Kimi K3 open-source model has become a global hit, has filed confidentially an application for a Hong Kong initial public offering, Chinese media outlet LatePost reported Wednesday. Moonshot’s path to its IPO hasn’t been easy. The company went through an… 8 The Information — AI news-outlet 4d ago Why I’m Launching The Information’s AI Show, AI Deep Dive The Information has become the most trusted source on the AI business by breaking news about the biggest deals in the industry. Now we’re expanding that coverage with a new show called “AI Deep Dive” that will give you a more sophisticated understanding of the technology itself.… 19 Hugging Face Daily Papers research 4d ago Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You Abstract CASTER enables gradient-free test-time adaptation for frozen models by transporting source class distributions via affine transformations in a discriminative subspace, with a transportability certificate to gate unreliable updates. Generated by… 10 r/MachineLearning community 4d ago Most open-source AI detectors can't hold a 0.5% false-positive rate [P] We needed to know where the open-source AI-detection field actually stands, so we ran every notable open detector through the same protocol. Setup: - Public data only: Jabarian & Imas 2025 (NBER), Liang 2023 TOEFL essays, a 1,060-text frontier set (GPT-5.x, Claude Opus 5, Gemini… 14 llama.cpp releases dev-tools 4d ago b10754 opencl: fix out‐of‐bound reads in the Adreno image kernels ( #27632 ) opencl: clamp the q4_K decode GEMV's fetch row on a padded x-grid opencl: enforce the tiling contract of the image KQ/KQV GEMMs opencl: decide the image KQ/KQV split at the dispatch, not from strides Website:… 26 arXiv — Machine Learning research 4d ago Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training arXiv:2609.00047v1 Announce Type: new Abstract: Graph prompt learning is an effective paradigm to adapt pre-trained graph models to downstream tasks in low-resource scenarios. However, existing multi-task graph pre-training frameworks generally use randomly initialized prompts,… 7 arXiv — Machine Learning research 4d ago REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent arXiv:2609.00049v1 Announce Type: new Abstract: Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain… 31 arXiv — Machine Learning research 4d ago Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding arXiv:2609.00097v1 Announce Type: new Abstract: The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the… 6 Page 1 of 10 · 500 articles Older →