News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — NLP / Computation & Language research 12d ago Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion arXiv:2609.16777v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a… 13 arXiv — NLP / Computation & Language research 12d ago Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising arXiv:2609.16997v1 Announce Type: new Abstract: Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis sentiment dataset with… 8 arXiv — NLP / Computation & Language research 12d ago CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine arXiv:2609.16301v1 Announce Type: cross Abstract: Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to… 13 arXiv — NLP / Computation & Language research 12d ago CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection arXiv:2609.16582v1 Announce Type: cross Abstract: Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation… 25 r/LocalLLaMA community 12d ago Open Source Appreciation Post It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off. Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO… 18 Vercel — AI dev-tools 12d ago Is Agentic now tailors its audit by site type Is Agentic reports now let you view your checks through one of four site types: Docs & content, Business, App, or Commerce. For example, the Commerce view highlights payment and checkout standards like x402, UCP, and ACP, while the App view highlights API discovery,… 21 The Information — AI news-outlet 12d ago Salesforce Unveils New AI Model, Agent Security Tools at Dreamforce At the opening of its annual Dreamforce customer conference, Salesforce signaled its involvement in two major trends in enterprise AI: Open-source AI models and software that helps corporate customers safely run AI agents on their data. Salesforce unveiled a new AI model, called… 34 r/LocalLLaMA community 12d ago Closed source AI is more dangerous than open source AI. Security through obscurity is no form of security.   submitted by   /u/keequalshalfmvsqrd [link]   [comments] 18 TechCrunch — AI news-outlet 13d ago Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear Salesforce Koa is built on Nvidia's open-weight Nemotron model and is trained to do sales, marketing and customer-support tasks. 14 r/LocalLLaMA community 13d ago jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar. For years, the JVM has watched the AI revolution from the bench. Every model, AI framework, every breakthrough, built with/for Python. jinfer is an inference engine built for the JVM from first principles: chat, vision, audio transcription, embeddings, reranking, and TTS. No… 7 Hugging Face Daily Papers research 13d ago Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks Abstract Backdoor vulnerability in fine-tuned language models varies drastically with poison selection, and a learned set-scoring method improves worst-case attack success by identifying high-impact poisoned examples. Generated by thinkingmachines/Inkling-Small Backdoor… 15 arXiv — Machine Learning research 13d ago BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents arXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an… 32 arXiv — Machine Learning research 13d ago Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction arXiv:2609.13701v1 Announce Type: new Abstract: Accurate job runtime prediction can improve scheduling-aware resource management in grid and distributed computing environments, but prediction models must be evaluated under realistic deployment constraints. This paper revisits… 19 arXiv — Machine Learning research 13d ago Bayesian optimization with kernel ensembles and disagreement-based acquisition for source localization and acoustic inversion arXiv:2609.14262v1 Announce Type: new Abstract: Joint source localization and geoacoustic inversion requires optimizing an objective built from an expensive normal mode propagation model. Bayesian optimization (BO) with a Gaussian process (GP) surrogate can obtain accurate… 29 arXiv — Machine Learning research 13d ago Learning Source Acquisition Policies by Offline Planning arXiv:2609.14299v1 Announce Type: new Abstract: Predicting under an acquisition budget requires choosing feature groups whose value can depend on later queries. O-MPAC transfers finite-horizon risk-cost targets from complete training records into a shared source-action scorer.… 14 arXiv — NLP / Computation & Language research 13d ago Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning arXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by… 38 arXiv — NLP / Computation & Language research 13d ago The University of Melbourne WMT 2026 CreoleMT Submission: A Domain-Balanced Approach to Low-Resource Pacific Creole Machine Translation arXiv:2609.13615v1 Announce Type: new Abstract: For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bislama, and Solomon Pijin, with particular attention to broad domain performance.… 4 arXiv — NLP / Computation & Language research 13d ago Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice arXiv:2609.13841v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically… 6 arXiv — NLP / Computation & Language research 13d ago Measuring the Cost of Variety Conflation in Multilingual MT Evaluation: Adding Mozambican Xichangana, Nyanja and Sena to FLORES+ arXiv:2609.13847v1 Announce Type: new Abstract: In this paper, we extend FLORES+ with Portuguese-source evaluation sets for three Mozambican Bantu varieties: Xichangana, Mozambican Nyanja, and Sena. We compare Xichangana with the existing Tsonga reference and Mozambican Nyanja… 10 arXiv — NLP / Computation & Language research 13d ago Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection arXiv:2609.14570v1 Announce Type: new Abstract: Multi-agent LLM systems combine multiple inference calls, but prior work often confounds how calls are connected with how they are diversified. We study these factors independently: inference topology and source of inter-agent… 37 arXiv — NLP / Computation & Language research 13d ago Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation arXiv:2609.14963v1 Announce Type: new Abstract: As large language models become capable translators of classical texts, a key challenge is deciding which outputs need expert review when no human reference exists. This study tests reference-free error triage through… 21 arXiv — NLP / Computation & Language research 13d ago Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains arXiv:2609.14969v1 Announce Type: new Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform… 37 arXiv — NLP / Computation & Language research 13d ago Salesforce Koa: An Enterprise Language Model for Agentic Tool Use arXiv:2609.15066v1 Announce Type: new Abstract: We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is… 38 Hugging Face Daily Papers research 13d ago PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Abstract PhysBrain 1.5 unifies physical environment understanding, action generation, and future state prediction via joint autoregressive training on discrete vision-language, motion, and visual target sequences, achieving state-of-the-art open-source embodied performance.… 29 r/LocalLLaMA community 13d ago What are Open-Source Views on 'Slowing Down AI'? Personally, I think the whole "AI (LLMs) is going to take over" is just BS marketing. They've been pushing this doom-and-gloom since 2019, and we all know the models back then were nowhere near as capable as today's. What does everyone here think, since it's such a popular topic… 34 The Information — AI news-outlet 13d ago Wall Street Delivers Verdict on AI Slowdown’s Winners and Losers Guess who likes the idea of an AI development slowdown? Software investors! Stocks of enterprise software firms, such as Salesforce, Shopify, ServiceNow and Figma, all gained ground on Monday—ServiceNow jumped 7%, for instance, while Salesforce rose nearly 5%. Software firms, of… 12 The Information — AI news-outlet 13d ago Refounding America: The Tax Code Needs to Change in the Age of AI Vinod Khosla is founder of Khosla Ventures and has spent the last 40 years in the business of innovation. He was the co-founder of Sun Microsystems. Khosla Ventures is an investor in OpenAI. Major technological revolutions should force a renegotiation of society’s basic economic… 8 r/LocalLLaMA community 13d ago Are there any organizations that are lobbying in favor of open source AI? So we’re seeing how Anthropic and OpenAI are gunning for regulations. I think most of us realize that this is a ploy for them to achieve regulatory capture, thus securing their moat and kicking out open source. The thing is, there’s so much vested corporate interest in ensuring… 36 r/LocalLLaMA community 13d ago Xi promotes open source AI zone among BRICS countries Xi emphasised the need to step up cooperation. https://www.hindustantimes.com/india-news/china-will-lead-creation-of-open-source-ai-for-brics-xi-jinping-101789300908632.html… 17 Hacker News — AI on Front Page community 13d ago OpenAI bots knew about the RubyGems caching vulnerability Article URL: https://tenderlovemaking.com/2026/09/11/what-a-time-to-be-alive/ Comments URL: https://news.ycombinator.com/item?id=49695876 Points: 211 # Comments: 212 13 r/LocalLLaMA community 14d ago The new k2 horizon models seem like an absolute beast Especially the 7B one seems very interesting, it casually destroys muse glimmer with a way smaller size. And they open source literally everything, every step of the way. Anyone tried that model? It can be a new milestone if 7b and 3.7b ones are actually good, and not just… 11 r/LocalLLaMA community 14d ago Muse spark 1.3 is great, but will they open source it? Fingers crosses, this model seems really good when it comes to coding, and the speeds are nice. They talked about making it open source but what happened? I still hope they release 1.2/1.3.   submitted by   /u/formatme [link]   [comments] 33 arXiv — Machine Learning research 14d ago Efficient AI Model Deployment Using Quantization Analysis Tool arXiv:2609.11954v1 Announce Type: new Abstract: As deep learning models are increasingly deployed on resource constrained devices, the demand for efficient model optimization techniques continues to grow. Effective deployment of AI models on edge and low power platforms requires… 17 arXiv — Machine Learning research 14d ago FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences arXiv:2609.11993v1 Announce Type: new Abstract: Machine learning research in financial services is limited by the scarcity of representative open-source datasets. Existing resources are often narrowly focused on a single modality or task and fail to reflect the structured,… 11 arXiv — Machine Learning research 14d ago MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling arXiv:2609.13048v1 Announce Type: new Abstract: Efficient microservice scheduling is crucial for maintaining load balance across nodes in data centers and ensuring high quality of service. However, achieving this in practice remains challenging due to dynamic resource imbalance… 4 arXiv — NLP / Computation & Language research 14d ago The Cost of Compression: A Rate-Distortion Limit on Factual Hallucination arXiv:2609.12111v1 Announce Type: new Abstract: Factual hallucination in closed-book question answering is often treated as a coverage problem: a model fails because the relevant fact is absent from its internal memory. This view misses a second source of error. Even when a fact… 19 arXiv — NLP / Computation & Language research 14d ago EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development arXiv:2609.12268v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) can improve knowledge-intensive question answering, but the first design choice is easy to overlook: how should the source corpus be partitioned into retrievable units? Fixed-size chunks often… 21 arXiv — NLP / Computation & Language research 14d ago I Am No One: Style-Aware Paraphrasing for Text Anonymization arXiv:2609.12341v1 Announce Type: new Abstract: Authorship attribution models can re-identify users from seemingly anonymized text by exploiting stable stylistic fingerprints, even after explicit identifiers are removed, posing a growing privacy risk for text publishing and… 17 arXiv — NLP / Computation & Language research 14d ago Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, these models struggle with predicting high-bitrate… 27 arXiv — NLP / Computation & Language research 14d ago Cortex: Content Analysis Support Software, a Resource for Qualitative Research arXiv:2609.11970v1 Announce Type: cross Abstract: Qualitative research is widely used in the human and social sciences, characterized by a deep understanding of phenomena through the interpretation of meanings and contexts. Among qualitative data analysis methods, content… 9 arXiv — NLP / Computation & Language research 14d ago UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking arXiv:2505.15063v3 Announce Type: replace Abstract: The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are… 21 r/LocalLLaMA community 14d ago Right to Intelligence. Protect your right to run local AI. With all the recent drama surrounding AI safety. It’s obvious that open source could be caught in the crossfire.   submitted by   /u/Euphoric_Ad9500 [link]   [comments] 30 r/LocalLLaMA community 14d ago Decided to build a game, and test the ceiling of Qwen3.8 27b This took roughly 5 hours to create, using 2 different configured harnesses, same model. RTX 3090, overclocked +12% gain (MSI Afterburner), Q4KM - built this for fun, will be throwing it on GitHub, opensource for people to get an idea of a project created to the near ceiling of… 35 r/MachineLearning community 14d ago Pacing the Frontier – Tahuna: AI Training Infrastructure, Now Open Source [P] Dario says we need to pace the frontier. Good news: we’ve been pacing Tahuna for months. Today, Tahuna is open source—as promised back in April. We built it so small teams could train models, run inference, orchestrate GPUs, and experiment with autonomous research without first… 32 r/LocalLLaMA community 15d ago Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen ( https://github.com/peonist-ai/halogen-flash-server ) that… 7 r/LocalLLaMA community 15d ago Looks like a coordination to stop distribution of intelligence Coxon, bernie and now this First https://x.com/DarioAmodei/status/2098773920774074715 Then https://x.com/elonmusk/status/2098789109980332057 Then https://x.com/sama/status/2098811563415150910 I think fear mongering approaching and they will try to slow down open source "They"… 4 r/LocalLLaMA community 15d ago For those of you forced to only use open models from Western labs in production, what are you deploying? First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects and for school, however, I’m also responsible for deploying local AI on my organization’s H100s, and we are forbidden by… 26 r/LocalLLaMA community 16d ago This is why we need open-source harnesses + local models i've been thinking about this more after trying different agent setups. the model isn't the only thing that determines how well an agent performs. The harness around the model matters a lot too. With a managed agent setup, you're often giving up control over things like the… 21 The Information — AI news-outlet 16d ago Exclusive: Salesforce to Drop ‘Agentforce’ Branding From Some Products Salesforce, ahead of its annual Dreamforce customer conference next week, is removing the word “Agentforce” from the names of some of its top products, according to two people with direct knowledge of the matter. The reversal comes after Salesforce made Agentforce a centerpiece… 14 Interconnects (Nathan Lambert) research 16d ago Open-Source AI & Open Models Reading List How to get up to speed on open models and their implications. 4 Page 4 of 10 · 500 articles ← Newer Older →