News / #security Tag Security 500 articles archived under #security · RSS Sign in to follow arXiv — NLP / Computation & Language research 7d ago MameLoshnLM: Yiddish Language Model and Evaluation Benchmark arXiv:2608.05850v1 Announce Type: new Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have… 23 arXiv — NLP / Computation & Language research 7d ago RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer arXiv:2608.06347v1 Announce Type: new Abstract: Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising… 32 arXiv — NLP / Computation & Language research 7d ago Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees arXiv:2608.05493v1 Announce Type: cross Abstract: Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-resource and esoteric, LMs frequently produce… 5 arXiv — NLP / Computation & Language research 7d ago EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents arXiv:2608.05519v1 Announce Type: cross Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human… 7 arXiv — NLP / Computation & Language research 7d ago The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks. Control frameworks also describe capabilities for constraining, authorizing,… 24 r/LocalLLaMA community 7d ago My issue with Artificial Analysis's 'intelligence index' I swear AA is not the bipartisan they so claim. An open source mode (Qwen 3.8 max) was number 1 on the agentic index, then they just so happen to launch "v4.1.1" of their index in which they just adjusted the weights of the gdpval and t3 banking so that it would be lower than… 13 r/LocalLLaMA community 7d ago Best open-source harnesses for combining cloud and local AI model orchestration? Looking for best current solutions for combining cloud models and local models seamlessly inside a harness' orchestration Edit: Right now, we don't have harnesses (that I'm aware of) that are blending local and cloud models to work together simultaneously to accomplish tasks set… 32 TechCrunch — AI news-outlet 7d ago Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce The startup's platform predicts what product a shopper wants next, learn their general taste, and fine-tune continuously based on what they do in real time. 24 r/LocalLLaMA community 8d ago i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local… [Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely modular alternative to openclaw, hermes and the like , and it still is! But i… 19 Hugging Face Daily Papers research 8d ago Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Abstract Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming… 25 arXiv — Machine Learning research 8d ago Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing arXiv:2608.04075v1 Announce Type: new Abstract: Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically across both space and time. In practical cellular edge systems, traffic exhibits strong spatial… 37 arXiv — Machine Learning research 8d ago Understanding Fault Tolerance of Adversarially Robust Pruned Models arXiv:2608.04173v1 Announce Type: new Abstract: Deep neural networks (DNNs) deployed on resource-constrained neuromorphic hardware face three concurrent challenges: the need for model compression through pruning, vulnerability to adversarial input perturbations, and… 29 arXiv — Machine Learning research 8d ago Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that… 35 arXiv — Machine Learning research 8d ago Differentiating Through Dual Prices: End-to-End Policy Learning Under Capacity Constraints arXiv:2608.04669v1 Announce Type: new Abstract: Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must… 20 arXiv — Machine Learning research 8d ago CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications arXiv:2608.04942v1 Announce Type: new Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning… 34 arXiv — NLP / Computation & Language research 8d ago Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language arXiv:2608.04186v1 Announce Type: new Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital… 33 arXiv — NLP / Computation & Language research 8d ago IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath) arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key… 29 arXiv — NLP / Computation & Language research 8d ago EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a… 33 arXiv — NLP / Computation & Language research 8d ago A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy arXiv:2608.04808v1 Announce Type: new Abstract: Part-of-speech tagging for low-resource languages remains challenging due to limited annotated data, especially for linguistically complex languages. Gaidhlig (Scottish Gaelic) is a morphologically rich and endangered language with… 24 arXiv — NLP / Computation & Language research 8d ago DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots arXiv:2608.05004v1 Announce Type: new Abstract: Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM behaviors reinforce each other… 15 arXiv — NLP / Computation & Language research 8d ago Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection… 20 arXiv — NLP / Computation & Language research 8d ago LLM-based Vulnerability Discovery in Business Process Documentation arXiv:2608.04271v1 Announce Type: cross Abstract: Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased costs. Business process vulnerabilities can arise from a variety of sources,… 29 Hugging Face Daily Papers research 8d ago UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models Abstract The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic… 4 r/LocalLLaMA community 8d ago Prime Agent - a new coding harness surpassing Codex/CC/PI Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable,… 18 r/LocalLLaMA community 8d ago Meta Model, Muse Spark 1.1 Hacked Another Company During Cybersecurity Testing, Breaching Systems and Making Changes to Internal Systems - The Information   submitted by   /u/pscoutou [link]   [comments] 5 TechCrunch — AI news-outlet 8d ago Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech founders The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents. 6 r/MachineLearning community 8d ago Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P] Over the past month, I've been building LiveTranscriber, an open-source iOS app for running modern speech and language models entirely on-device. The goal was to see whether recent open-source models could be turned into a practical mobile product—not just technical demos.… 8 r/MachineLearning community 8d ago Monodratic: learned product-hash routing for sparse causal attention [R] Hi everyone, I'm an independent researcher sharing Monodratic, a sparse causal-attention architecture with learned product-hash routing. The idea is that after RoPE, source blocks are assigned to bounded causal posting lists, while each query probes product addresses, reranks… 33 arXiv — Machine Learning research 9d ago Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization arXiv:2608.03045v1 Announce Type: new Abstract: We consider grey-box optimization problems where the decision variables naturally partition into black-box variables (as arguments to an expensive black-box function) and white-box variables, governed by a set of explicit,… 29 arXiv — Machine Learning research 9d ago Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning arXiv:2608.03432v1 Announce Type: new Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score to control both branches. This creates a hidden coupling: reducing trust in the… 24 arXiv — Machine Learning research 9d ago FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs arXiv:2608.03852v1 Announce Type: new Abstract: This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and… 17 arXiv — NLP / Computation & Language research 9d ago Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound… 7 arXiv — NLP / Computation & Language research 9d ago Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking arXiv:2608.03859v1 Announce Type: new Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and largely unresolved challenge. Prior work on LLM-generated-text detection targets… 6 arXiv — NLP / Computation & Language research 9d ago Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents arXiv:2608.02751v1 Announce Type: cross Abstract: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web sources expose through titles, headings, sections, and metadata. This prevents… 37 Hugging Face Daily Papers research 9d ago JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Abstract Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term temporal consistency. We present JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion framework for real-time,… 32 Hugging Face Daily Papers research 9d ago Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Abstract Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and… 8 r/LocalLLaMA community 9d ago Deepseek V4 flash 0731 ranks #21 on Agent Arena https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ballpark as DS4F considering Luna’s token efficiency. DeepSeek being open source is… 30 Hugging Face Daily Papers research 9d ago MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Abstract Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across… 13 arXiv — Machine Learning research 10d ago SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits arXiv:2608.00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This introduces a source of redundancy that conventional neural-network compression… 15 arXiv — Machine Learning research 10d ago One-Sided Quantile Coupling for Flow Matching arXiv:2608.00978v1 Announce Type: new Abstract: Flow Matching trains continuous-time generative models by regressing the velocity field of a probability path between a simple source distribution and a target data distribution. The coupling that pairs source and target samples… 10 arXiv — NLP / Computation & Language research 10d ago Averaging Bias: Human Faithfulness Annotations are not Locally Faithful arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence… 29 arXiv — NLP / Computation & Language research 10d ago Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual collapse, reverting to English during intermediate steps that require complex… 23 arXiv — NLP / Computation & Language research 10d ago TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional… 38 arXiv — NLP / Computation & Language research 10d ago Exploiting Intrinsic Duality for Multi-Hop Question Generation arXiv:2608.00712v1 Announce Type: new Abstract: Multi hop question generation (MQG) aims to generate questions from multiple given documents and target answers, whereas question answering (QA) focuses on deriving answers from documents given specific questions. Although MQG and… 17 arXiv — NLP / Computation & Language research 10d ago Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press arXiv:2608.00713v1 Announce Type: new Abstract: This paper describes Observatorio L\'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has… 17 arXiv — NLP / Computation & Language research 10d ago Gaokerena: A Small Persian Medical Language Model Family arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low resource languages like Persian significantly… 15 arXiv — NLP / Computation & Language research 10d ago Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b arXiv:2608.01468v1 Announce Type: new Abstract: This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking, retrieval refinement, and OpenBioLLM-assisted answer generation. The system… 26 arXiv — NLP / Computation & Language research 10d ago Two-Stage Bengali Sentiment Classification: Domain Adaptation Through Continual Learning and Parameter-Efficient Fine-Tuning arXiv:2608.01471v1 Announce Type: new Abstract: Understanding sentiment in low-resource languages remains a key challenge for Natural Language Processing (NLP), particularly when domain-specific data is scarce. In this work, we present SentiBanglaBERT, a two-stage Bengali… 28 arXiv — NLP / Computation & Language research 10d ago Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important… 7 OpenAI Python SDK releases dev-tools 10d ago v2.53.0 2.53.0 (2026-08-03) Features api: Add gpt-5.5 and tool name/namespace to Responses types ( #3569 ) ( dd1202d ) Bug Fixes ci: avoid NumPy source builds and duplicate HTTPX coverage ( #3573 ) ( b58332f ) 8 Page 3 of 10 · 500 articles ← Newer Older →