News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — NLP / Computation & Language research 1mo ago Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR arXiv:2607.11163v1 Announce Type: new Abstract: Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue,… 36 arXiv — NLP / Computation & Language research 1mo ago The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition arXiv:2607.11341v1 Announce Type: new Abstract: This paper addresses the challenges of using sign language within shared mobility services, such as taxis, carpools, or ride-sharing platforms. The use of sign language recognition (SLR) in real-world, confined environments,… 18 arXiv — NLP / Computation & Language research 1mo ago JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes arXiv:2607.11715v1 Announce Type: new Abstract: Large-scale, richly annotated career trajectory data underpins workforce planning, job recommendation, and labour market analysis, yet publicly available datasets are either small, closed to independent use, or built from… 15 arXiv — NLP / Computation & Language research 1mo ago STEP: Career-Path Recommendation via Temporal and Educational Trajectory Modeling arXiv:2607.11722v1 Announce Type: new Abstract: Career paths encode decades of skill acquisition, role transitions, and educational investment, and understanding them at scale underpins workforce planning, labor market policy, and job recommendation. Resumes are a rich source of… 24 r/LocalLLaMA community 1mo ago Why aren't any American open-source AI labs even close to Chinese ones on benchmarks yet? I know there are a few american labs working on open-source AI but none of them show up in the benchmarks like Chinese open source does, why haven't any American labs been able to reach top open source benchmarks yet?   submitted by   /u/Lost_Foot_6301 [link]  … 28 r/LocalLLaMA community 1mo ago This is why we need local models and opensource harnesses   submitted by   /u/Comfortable-Rock-498 [link]   [comments] 19 r/LocalLLaMA community 1mo ago If Frontier AI is so Dangerous, Why should private companies be allowed to develop it? There's a push by openAI and anthropic primarily. Well more than a push just straight fear mongering about open source ai and its dangers. If it's so dangerous why would the US gov. in particular continue to allow private industry to develop and release. If I had a company… 22 Ars Technica — AI news-outlet 1mo ago Now, defenders are embracing the prompt injection, too "Context bombing" tricks hacking agents into shutting down before they can do harm. 14 r/MachineLearning community 1mo ago Hundreds of papers hit arXiv every day and maybe 3 matter to my research, so I built an open-source tool that finds them [P] Left: Telegram digest (optional); Right: detailed digest on HTML Like probably everyone here, my to-read list only grows. Skimming arXiv listings or my feeds takes 30-60 minutes a day, 95% of it is irrelevant to what I actually work on, and newsletters don't really help: they… 8 r/LocalLLaMA community 1mo ago Zhipu founder backs open-source AI as global security debate intensifies   submitted by   /u/computeruser420 [link]   [comments] 25 Hugging Face Daily Papers research 1mo ago A Sovereign, Open-Source Foundation Model for German and English Abstract We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows,… 27 r/LocalLLaMA community 1mo ago Compressed Version of Qwen-3.6-27B coming from PrismML - Khosla-Backed Startup Claims Breakthrough With Largest-Ever AI Model on an iPhone The startup, PrismML, said it has shrunk down Qwen 3.6 , an open-source large language model developed by Chinese internet giant Alibaba , to run on an iPhone 17 Pro. The model has 27 billion parameters , which are roughly similar to the synapses in a brain and can help… 33 arXiv — Machine Learning research 1mo ago Pitfalls and Remedies for Multi-Task Bayesian Optimization arXiv:2607.09073v1 Announce Type: new Abstract: Bayesian optimization routinely warm-starts a target experiment with data from related source tasks, and the multi-task Gaussian process is the textbook surrogate for the job. We revisit this default in a controlled setting and… 6 arXiv — Machine Learning research 1mo ago Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning arXiv:2607.09422v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, such as the tuning of electrostatically-defined quantum-dot arrays. However, if parameter… 6 arXiv — Machine Learning research 1mo ago Control Laguerre Tessellation: Semi-discrete Optimal Transport Over Control Systems arXiv:2607.09139v1 Announce Type: cross Abstract: We study the optimal transport of optimally controlled agents from a compactly supported absolutely continuous source to a discrete target measure. The ground cost for the transport is induced by the optimal cost of the agents'… 18 arXiv — NLP / Computation & Language research 1mo ago WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning arXiv:2607.09328v1 Announce Type: new Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed… 16 arXiv — NLP / Computation & Language research 1mo ago A Sovereign, Open-Source Foundation Model for German and English arXiv:2607.09424v1 Announce Type: new Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference… 15 arXiv — NLP / Computation & Language research 1mo ago Topic model based on co-occurrence word networks for unbalanced short text datasets arXiv:2311.02566v2 Announce Type: replace Abstract: We propose a straightforward solution for detecting scarce topics in unbalanced short-text datasets. Our approach, named CWUTM (Topic model based on co-occurrence word networks for unbalanced short text datasets), addresses the… 21 Hugging Face Daily Papers research 1mo ago PanoWorld: Real-World Panoramic Generation Abstract In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implicit geometric transformation.Building on this… 4 Hacker News — AI on Front Page community 1mo ago Show HN: Juggler – an open-source GUI coding agent, by the creator of JUCE Hello HN, I don't post on here much, but wanted to get some eyes on a new project I'm just launching. I think we definitely need one more AI code agent.. I'm a long-term C++ dev, and over 30+ years I've created some successful audio dev tools (JUCE, the Tracktion DAW, the Cmajor… 11 Interconnects (Nathan Lambert) research 1mo ago 6 months to live for open models The most serious test to date of open source AI’s viability is happening right now. 32 r/LocalLLaMA community 1mo ago Built EverFern, an alternative to cowork open-source(LangGraph + Electron). Tested it against Qwen3-8B, looking for feedback I built EverFern, an open-source (MIT) desktop agent that does computer use, browser control, and file/code tasks, aimed at people who don't want to pay $20–200/mo for Claude Cowork or Manus and send everything to the cloud. Everything runs locally, config and history live in… 29 r/LocalLLaMA community 1mo ago The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico Politico: Wall Street’s new obsession: Which CEOs have Trump’s ear?: https://www.politico.com/newsletters/politico-influence/2026/07/10/wall-streets-new-obsession-which-ceos-have-trumps-ear-00993324   submitted by   /u/Nunki08 [link]   [comments] 29 r/LocalLLaMA community 1mo ago Can we reconstruct a closed-source LLM tokenizer using only two oracles from the chat API? From the chat APIs alone, we can extract two useful oracles: 1. Token length oracle: Given any string s, return len(tokenize(s)). Prefix token oracle: Given a string s and integer n, return the string decoded from the first n tokens of tokenize(s). The first oracle (token count)… 33 r/LocalLLaMA community 1mo ago MIT LLM Serve Dashboard I am making open source A single-file, dependency-free live dashboard for your local LLM serving box — GPU utilization, per-model throughput, KV/context fill, and system stats for llama.cpp and vLLM , in one green terminal-styled page. No framework, no build step, no external requests. The frontend is… 31 TechCrunch — AI news-outlet 1mo ago Open source AI matters more than ever, according to Hugging Face’s Clem Delangue Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent years, where AI builders can share and download open models and datasets, now used by roughly half the Fortune 500. Delangue… 20 r/LocalLLaMA community 1mo ago NVIDIA Readies GeForce RTX 5090 SE Graphics Card - TPU   submitted by   /u/panchovix [link]   [comments] 14 Ars Technica — AI news-outlet 1mo ago Disable auto-play and infinite scroll or risk massive fines, EU tells Meta Digital Services Act may force Meta to make big changes on its platforms. 13 TechCrunch — AI news-outlet 1mo ago Hugging Face’s CEO on why companies are done renting their AI Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent years, where AI builders can share and download open models and datasets, now used by roughly half the Fortune 500. Delangue… 20 Hacker News — AI on Front Page community 1mo ago EU Commission: addictive design Instagram and Facebook in breach of the DSA Article URL: https://ec.europa.eu/commission/presscorner/home/en Comments URL: https://news.ycombinator.com/item?id=48858292 Points: 211 # Comments: 144 19 arXiv — Machine Learning research 1mo ago Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation arXiv:2607.07748v1 Announce Type: new Abstract: Large Language Models achieve strong code generation for high resource languages like Python and Java but suffer sharp performance drops on Low-Resource Programming Languages~(LRPLs) such as Julia. Improving Small Language… 16 arXiv — Machine Learning research 1mo ago A Self-Supervised Approach for Minimal-Annotation Hydroacoustic Data Exploration arXiv:2607.07733v1 Announce Type: cross Abstract: Passive hydroacoustic monitoring often generates large volumes of continuous recordings that are only partially exploited due to the cost of manual annotation. Supervised detection methods perform well but require large labeled… 10 arXiv — NLP / Computation & Language research 1mo ago Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution arXiv:2607.08700v1 Announce Type: new Abstract: Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trusted, we need to know how capable the judge must be… 10 arXiv — NLP / Computation & Language research 1mo ago From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents arXiv:2607.08028v1 Announce Type: cross Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization adds requirements for source boundaries, entity routing, answer contracts,… 37 r/LocalLLaMA community 1mo ago Meta are apparently working on an open source variant of Muse Spark. No real details or timescales yet, but this article has confirmation from Alexandr Wang that Meta are working on an open source variant of Muse Spark. One to keep an eye on. https://www.cnbc.com/2026/07/09/meta-jumps-into-ai-coding-market-to-chase-anthropic-and-openai.html  … 6 OpenAI Python SDK releases dev-tools 1mo ago v2.45.0 2.45.0 (2026-07-09) Full Changelog: v2.44.0...v2.45.0 Features api: gpt-5.6-sol updates ( 039d1fe ) Bug Fixes api: restore beta resource accessors ( 2dfc130 ) Chores retrigger release automation ( 7b61351 ) 21 TechCrunch — AI news-outlet 1mo ago Popular open source AI developer tool Ollama raises $65M, grows to nearly 9M users Benchmark-backed Ollama has amassed 176,000 stars, and nearly 17,000 forks on Github by helping developers easily run AI on their PCs. 12 Hugging Face Daily Papers research 1mo ago Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo Abstract Large language models can be adapted for low-resource programming languages through specialized training pipelines and benchmarks, achieving superior code completion performance compared to general-purpose models. Generated by Qwen/Qwen2.5-Coder-32B-Instruct Large… 10 arXiv — Machine Learning research 1mo ago Does Demand Response Increase Vulnerability to Cyber Attacks by Adversarial Data Modifications? arXiv:2607.06632v1 Announce Type: new Abstract: Adversarial attacks are crafted data manipulations that aim to deteriorate the outcomes of prediction or decision-making algorithms. In the energy systems literature, adversarial attacks have been studied with a focus on problems… 4 arXiv — Machine Learning research 1mo ago Intrinsic-Noise Consolidation: A Doob-Barrier-Conditioned Diffusion Turns Analog Device Noise into a Continual-Learning Resource arXiv:2607.06924v1 Announce Type: new Abstract: On analog neuromorphic hardware, intrinsic device noise is normally an accuracy tax. We ask whether it can instead consolidate memories. We cast per-synapse consolidation as a Doob h-transform: condition each weight's stochastic… 18 arXiv — Machine Learning research 1mo ago Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery arXiv:2607.06930v1 Announce Type: new Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. Real-world datasets often exhibit complex latent structures composed of multiple subgroups… 18 arXiv — Machine Learning research 1mo ago Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting arXiv:2607.07016v1 Announce Type: new Abstract: Accurate forecasting of cellular network traffic is essential for network planning, resource allocation, and quality-of-service assurance in modern mobile communication systems. Real-world traffic often exhibits bursty endogenous… 28 arXiv — Machine Learning research 1mo ago Intrinsic Green's Learning: Supervised Learning on Manifolds via Inverse PDE arXiv:2607.07034v1 Announce Type: new Abstract: We introduce Intrinsic Green's Learning (IGL), a framework that models a target function on a manifold as the solution to a linear PDE whose source term is learned from data. Rather than approximating the target directly, IGL… 35 arXiv — Machine Learning research 1mo ago Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection arXiv:2607.07146v1 Announce Type: new Abstract: The Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time is the resource the effort exists to conserve. Because attention is the cost of… 29 arXiv — Machine Learning research 1mo ago FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation arXiv:2607.07314v1 Announce Type: new Abstract: Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training process and the learned model itself. Recently, centralized Taking Away Training Data (TATD)… 33 arXiv — Machine Learning research 1mo ago On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces arXiv:2607.07375v1 Announce Type: new Abstract: Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on… 4 arXiv — Machine Learning research 1mo ago Multi-Class vs. Multi-Label BERT for CVE-to-CWE Mapping: How Taxonomy Structure Shapes the Errors arXiv:2607.07573v1 Announce Type: new Abstract: Assigning Common Weakness Enumeration (CWE) categories to Common Vulnerabilities and Exposures (CVE) records remains an important but largely manual step in vulnerability analysis. We study this task as a text classification… 10 arXiv — NLP / Computation & Language research 1mo ago Ad Headline Generation using Self-Critical Masked Language Model arXiv:2607.06818v1 Announce Type: new Abstract: For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creative quality bar of the website, especially at a large scale. We thus propose a programmatic… 12 arXiv — NLP / Computation & Language research 1mo ago SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation arXiv:2607.07469v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to… 17 arXiv — NLP / Computation & Language research 1mo ago Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems arXiv:2512.17648v2 Announce Type: replace Abstract: Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech under strict latency constraints, demanding models that balance low latency with high translation quality.… 7 Page 10 of 10 · 500 articles ← Newer