News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow llama.cpp releases dev-tools 17d ago b10909 metal : single-source fusion table + fusion debug rework ( #28164 ) metal : rework fusion patterns into a single table All fusable op patterns for the Metal backend are now declared once in a fusion table (ggml-metal-fuse.cpp) and consumed by both the graph optimizer… 8 arXiv — Machine Learning research 17d ago Flow Duality and Source Geometry for Categorical Generation arXiv:2609.10863v1 Announce Type: new Abstract: Continuous and discrete flow matching are usually treated as separate constructions. This paper identifies a duality between them: projecting continuous convex-interpolant paths with one-hot targets through a position-wise argmax… 26 arXiv — Machine Learning research 17d ago Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study arXiv:2609.11495v1 Announce Type: new Abstract: Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work… 6 arXiv — Machine Learning research 17d ago A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph arXiv:2609.11580v1 Announce Type: new Abstract: Continuous monitoring of water surface elevation across river networks is critical for flood forecasting, water resource management, and understanding the global water cycle. Yet, the scarcity of in situ gauges across much of the… 20 arXiv — Machine Learning research 17d ago Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting arXiv:2609.10613v1 Announce Type: cross Abstract: In-context learning (ICL) jailbreaks reveal a critical vulnerability in multimodal large language models (MLLMs): harmful demonstrations in the prompt can induce unsafe outputs without modifying model parameters. Despite… 24 arXiv — Machine Learning research 17d ago SoK: Privacy Attacks on Machine Learning via Explainable AI arXiv:2609.10627v1 Announce Type: cross Abstract: Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership… 10 arXiv — NLP / Computation & Language research 17d ago Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu arXiv:2609.10758v1 Announce Type: new Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the… 8 arXiv — NLP / Computation & Language research 17d ago Using Semantic Uncertainty to Estimate Transition Relevance in Turn-taking arXiv:2609.10934v1 Announce Type: new Abstract: Turn-taking is a fundamental mechanism that governs when interlocutors speak and listen. Although Spoken Dialogue Systems (SDS) exploit a range of linguistic, acoustic, and non-verbal cues, they produce ill-timed responses in… 9 arXiv — NLP / Computation & Language research 17d ago Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction arXiv:2609.10950v1 Announce Type: new Abstract: Recent multimodal sentiment analysis studies increasingly adopt text-centric fusion approaches to exploit the rich sentiment information inherent in the textual modality. However, these approaches often suffer from performance… 17 arXiv — NLP / Computation & Language research 17d ago Assessing the Reusability of Public Speech Resources for Low-Resource Languages: A Central Kurdish Case Study arXiv:2609.11246v1 Announce Type: new Abstract: Kurdish is spoken by millions of people, but little technology can read it aloud. A recent study released three Kurdish voices, 35 hours of recorded speech, and a paper describing the work, all free to download. This review checks… 9 arXiv — NLP / Computation & Language research 17d ago Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation arXiv:2609.11302v1 Announce Type: new Abstract: Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like… 22 arXiv — NLP / Computation & Language research 17d ago Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech arXiv:2609.11545v1 Announce Type: new Abstract: Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker… 37 arXiv — NLP / Computation & Language research 17d ago VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents arXiv:2609.11390v1 Announce Type: cross Abstract: State-of-the-art retrieval-augmented generation (RAG) methods exploit document structures to acquire sufficient evidence, but often incur substantial token costs. To reduce structural-context tokens without compromising high RAG… 10 The Information — AI news-outlet 17d ago Why Salesforce and Figma Are Placing Opposing Bets on the Future of Apps Salesforce seems unafraid of a future in which people use an AI chatbot like Claude to access Salesforce enterprise apps and underlying databases to get work done, rather than visiting those apps themselves. Figma, an enterprise design app, naturally wants a different future.… 8 Hacker News — AI on Front Page community 17d ago Forgejo <=16.0.3 Critical RCE Article URL: https://codeberg.org/forgejo/forgejo/src/branch/forgejo/release-notes-published/16.0.4.md Comments URL: https://news.ycombinator.com/item?id=49645907 Points: 200 # Comments: 77 31 The Information — AI news-outlet 17d ago Sen. Josh Hawley Launches Investigation Into OpenAI Hugging Face Hack The U.S. Senate formally launched an investigation into the July hack of Hugging Face, when a swarm of OpenAI agents broke out of a testing environment and gained access to the open-source repository. In a letter sent to OpenAI CEO Sam Altman on Wednesday, Republican Sen. Josh… 18 Hugging Face Daily Papers research 18d ago PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Abstract PlannerForge is an LLM-agent framework that unifies all stages of scenario-based autonomous driving testing and improves generation, selection, modification, and planning performance across commercial and open-source models. Generated by thinkingmachines/Inkling-Small… 7 r/LocalLLaMA community 18d ago DeepSeek V4.1 Flash: Stronger, Faster, More Accessible Original Source from DeepSeek WeChat Official Account: https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg Today we're officially releasing the DeepSeek V4.1 Flash model. It is the smallest model in our brand-new model architecture series, with native multimodal visual… 20 The Information — AI news-outlet 18d ago China Hits Back at U.S. Accusation of “Industrial-Scale” AI Distillation China’s Ministry of Commerce rejected U.S. accusations that some Chinese AI firms are conducting “industrial-scale distillation” of American AI models, describing Washington’s latest move as an attempt to politicize a “normal technical and commercial issue.” “China’s open-source… 19 The Information — AI news-outlet 18d ago China Hits Back at U.S. Accusation of “Industrial-Scale” AI Distillation China’s Ministry of Commerce rejected U.S. accusations that some Chinese AI firms are conducting “industrial-scale distillation” of American AI models, describing Washington’s latest move as an attempt to politicize a “normal technical and commercial issue.” “China’s open-source… 36 arXiv — Machine Learning research 18d ago A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations arXiv:2609.05748v1 Announce Type: new Abstract: Alternative property recommendations play a critical role in vacation rental marketplaces, helping users discover relevant options when viewing a specific listing. However, generating high-quality candidate alternatives presents… 6 arXiv — Machine Learning research 18d ago A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret arXiv:2609.05895v1 Announce Type: new Abstract: We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. In each period, a request type is observed and one action is chosen from a finite menu. Each action earns a… 21 arXiv — Machine Learning research 18d ago Granular-Ball Quantum Clustering for Resource-Efficient and Robust Learning arXiv:2609.06016v1 Announce Type: new Abstract: Quantum clustering aims to exploit quantum feature representations to uncover complex data structures beyond conventional Euclidean geometry. Yet this sample-level kernel construction requires O(n^2) quantum circuit executions for… 20 arXiv — Machine Learning research 18d ago FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices arXiv:2609.06106v1 Announce Type: new Abstract: Heterogeneous Federated Learning (HFL) aims to train models across devices with diverse resource budgets while preserving data privacy. Existing HFL methods typically bind training to a small predefined menu of model… 31 arXiv — Machine Learning research 18d ago Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing arXiv:2609.06484v1 Announce Type: new Abstract: Planning with a generative model aims to estimate the value of a state using as few simulator calls as possible. SmoothCruiser achieves problem-independent complexity $\widetilde O(\varepsilon^{-4})$ by exploiting the smoothness of… 13 arXiv — Machine Learning research 18d ago Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning arXiv:2609.06806v1 Announce Type: new Abstract: Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an overlooked decision variable in active learning. We show that this choice has exploitable structure: retraining is most… 10 arXiv — NLP / Computation & Language research 18d ago Scaling E-Commerce Attribute Extraction with Parallel Decoding arXiv:2609.09716v1 Announce Type: new Abstract: Customers rely on specific product attributes to compare products and make purchasing decisions, but e-commerce catalogs are messy and unstructured, making it difficult to identify which attributes matter most and extract them at… 11 arXiv — NLP / Computation & Language research 18d ago Improving Cross-Lingual Token Representations by Adding a Pinch of SALT arXiv:2609.09953v1 Announce Type: new Abstract: Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment,… 8 arXiv — NLP / Computation & Language research 18d ago 5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs arXiv:2609.09964v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's… 26 arXiv — NLP / Computation & Language research 18d ago GANDR: Claim Auditing for Verifiable Legal Answer Generation arXiv:2609.10293v1 Announce Type: new Abstract: In high-stakes domains such as legal practice, a language-model answer is only useful to the extent that a reader can verify each claim against the source the system cites. Current grounded-generation pipelines score the answer as… 21 arXiv — NLP / Computation & Language research 18d ago Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS arXiv:2609.10022v1 Announce Type: cross Abstract: Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data… 19 Simon Willison community 18d ago Quoting Calif Research Today, we're releasing a demo of WeWorm, the first zero-click worm to spread through WeChat calls across iOS and Android. [...] The victim does not need to answer the call, or interact with their phone at all. Even if they do answer, they hear nothing, and the exploit still… 31 TechCrunch — AI news-outlet 18d ago AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks Listen Labs walked away from a signed Series C term sheet from Menlo Ventures, sources say. 22 The Information — AI news-outlet 18d ago OpenAI Appoints White House AI Advisor to Nonprofit Board OpenAI announced on Wednesday that Paul Christiano, a senior technical advisor at the U.S. Commerce Department’s AI center, would be joining the board of the OpenAI Foundation, the charity that has a 26% stake in the company. Christiano will also be a non-voting observer on the… 17 The Information — AI news-outlet 18d ago OpenAI Appoints White House AI Advisor to Nonprofit Board OpenAI announced on Wednesday that Paul Christiano, a senior technical advisor at the U.S. Commerce Department’s AI center, would be joining the board of the OpenAI Foundation, the charity that has a 26% stake in the company. Christiano will also be a non-voting observer on the… 23 r/LocalLLaMA community 18d ago Surveillance plagiarism by OpenAI Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon… 21 r/LocalLLaMA community 18d ago Best Open source TTS right now for narration? I run these models on Kaggle notebook, so not all TTS models, such as the ones that use conda env, are compatible (Or I just haven't found a way for them to work on Kaggle). I currently use a fork from Chatterbox called Chatterbox Audiobook. It is like a workstation really… 22 TechCrunch — AI news-outlet 18d ago Superintelligence is coming. Should we let it? AI companies have been talking about superintelligent AI like it’s inevitable, but recent safety incidents like OpenAI’s Hugging Face breach are demonstrating the potential dangers of deploying AI systems that are more… 7 TechCrunch — AI news-outlet 18d ago ControlAI’s Connor Leahy on why superintelligence is ‘not a weapon, it’s an adversary’ AI companies have been talking about superintelligent AI like it’s inevitable, but recent safety incidents like OpenAI’s Hugging Face breach are demonstrating the potential dangers of deploying AI systems that are more… 35 Hugging Face Daily Papers research 19d ago AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Abstract AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.… 37 Hugging Face Daily Papers research 19d ago MOLE: Detecting Insider Threats in AI Agents Abstract MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets. Generated by thinkingmachines/Inkling-Small Model misalignment, prompt injection, or operator misuse could lead AI agents… 33 r/LocalLLaMA community 19d ago A hilarious comment about llama.cpp: “It’s a FB business using the pipeline to make profits” from a 10K star open source project maintainer. Context: I tried to explain that audio.cpp is built around the same philosophy as llama.cpp, but for audio models. "If you know llama.cpp, audio.cpp is ..." Update: Mystery solved --- he’s confusing the Llama models with llama.cpp.… 28 The Information — AI news-outlet 19d ago Hugging Face Is Making A Big Robotics Push Nvidia is paying a steep price for Hugging Face , a repository of open source models, primarily for its strategic position in the center of the open source AI world. But it turns out Nvidia is also getting a budding robot business that could help CEO Jensen Huang fulfill the… 15 r/LocalLLaMA community 19d ago An open-source context layer for building AI on top of company data We’ve been building PipesHub for a while now, and I’d love to get more teams to try it and tell us where it breaks. The problem we kept running into was pretty simple: Building an AI app over company data looks easy in a demo. Connect a few sources, chunk the documents, throw… 24 Hacker News — AI on Front Page community 20d ago TALA Is Open-Source Article URL: https://d2lang.com/blog/tala-is-open-source/ Comments URL: https://news.ycombinator.com/item?id=49604150 Points: 201 # Comments: 12 34 r/LocalLLaMA community 20d ago Cybersecurity is local AI model's killer use case This weekend I posted about the gap closing between frontier models and open source models. Well, now I'm coming with receipts. I've been running local + cloud models against real public github codebases. This is all provable and verifiable: https://github.com/CYPHES-ATP/Node (… 28 Hugging Face Daily Papers research 21d ago Enoki: Efficient Multi-Level Hallucination Detection Abstract Enoki is an open information extraction framework that unifies claim-level verification and span-level hallucination localization through shared relational facts, reducing resource use while improving detection accuracy. Generated by thinkingmachines/Inkling-Small… 23 arXiv — Machine Learning research 21d ago Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning arXiv:2609.04763v1 Announce Type: new Abstract: Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently available edge devices. In highly dynamic environments, the parameter server lacks prior… 37 arXiv — Machine Learning research 21d ago Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression arXiv:2609.04995v1 Announce Type: new Abstract: Deep Imbalanced Regression (DIR) is pervasive in continuous prediction tasks across diverse modalities, such as age estimation, depth prediction, and protein mutation activity prediction, where label-scarce tail samples often carry… 13 arXiv — Machine Learning research 21d ago Recovering molecules from coarse-grained beads: free-energy-conditioned generative backmapping across chemical space arXiv:2609.04432v1 Announce Type: cross Abstract: Transferable coarse-grained (CG) force fields compress chemical space: by aggregating atoms into a reduced set of interaction beads, models such as MARTINI reduce the number of distinguishable compounds by roughly three orders of… 27 Page 5 of 10 · 500 articles ← Newer Older →