News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow Stratechery (Ben Thompson) community 38m ago Apps, Agents, and Aggregation Agents are the ultimate Aggregators; they reveal apps as a means, not an ends, and providing them is tech's biggest prize. 38 Hugging Face official-blog 1h ago Holo4: powering generalist computer-use agents Back to Articles a]:hidden"> Holo4: powering generalist computer-use agents Team Article Published September 28, 2026 Upvote 1 Tony Wu h-tonywu Hcompany maxime h-maxime Hcompany Frederic Renard frenard-h Hcompany Vincent Coyette vincentcoyette Hcompany Emrick Sinitambirivoutin… 36 r/LocalLLaMA community 1h ago NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not. https://x.com/JensenHuang/status/2104499465055023424   submitted by   /u/InternationalGap3698 [link]   [comments] 14 NVIDIA Developer Blog official-blog 2h ago NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of... 7 NVIDIA Developer Blog official-blog 2h ago Add Runtime Controls to AI Agents with NVIDIA OpenShell AI agents can be given a goal, write code, use tools, and keep working as new information becomes available. This opens the door to applications that... 17 MIT Technology Review — AI news-outlet 2h ago Who’s liable when AI agents go rogue? MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July,… 16 r/MachineLearning community 5h ago Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P] AI Engineering from Scratch is an MIT-licensed curriculum: 523 lessons across 20 phases, from linear algebra and backprop to transformers, LLMs, agents, and production serving. The code is stdlib-first, so you see every step instead of calling a library. This month's edition: -… 37 Simon Willison community 7h ago Quoting Muse AI Agent Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't… 18 arXiv — NLP / Computation & Language research 7h ago Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents arXiv:2609.30289v1 Announce Type: new Abstract: In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations,… 25 arXiv — NLP / Computation & Language research 7h ago Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents arXiv:2609.30293v1 Announce Type: new Abstract: The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool… 20 arXiv — NLP / Computation & Language research 7h ago Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops arXiv:2609.30297v1 Announce Type: new Abstract: Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't heard before"). A central challenge… 31 arXiv — NLP / Computation & Language research 7h ago CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production arXiv:2609.30471v1 Announce Type: new Abstract: Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically… 4 arXiv — NLP / Computation & Language research 7h ago Inquesto Score: A reliability Protocol For Voice Agents arXiv:2609.30514v1 Announce Type: new Abstract: Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto… 22 arXiv — NLP / Computation & Language research 7h ago Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms arXiv:2609.30558v1 Announce Type: new Abstract: Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a… 23 arXiv — NLP / Computation & Language research 7h ago The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge arXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations,… 31 arXiv — NLP / Computation & Language research 7h ago ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning arXiv:2609.30906v1 Announce Type: new Abstract: Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical… 12 arXiv — NLP / Computation & Language research 7h ago PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding arXiv:2609.31255v1 Announce Type: new Abstract: General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is… 35 arXiv — NLP / Computation & Language research 7h ago Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content arXiv:2609.30563v1 Announce Type: cross Abstract: Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on… 21 arXiv — NLP / Computation & Language research 7h ago Epstein Files Engine: Agentic Search for Investigative Journalism arXiv:2609.30611v1 Announce Type: cross Abstract: On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times… 24 arXiv — NLP / Computation & Language research 7h ago Prompt Injection Detection for Email Agents Through Attack Chain Modeling arXiv:2609.30657v1 Announce Type: cross Abstract: Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection… 6 arXiv — NLP / Computation & Language research 7h ago KuaFu: Compressing Long User Behavior into Understanding at Billion Scale arXiv:2609.31045v1 Announce Type: cross Abstract: Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant… 37 arXiv — NLP / Computation & Language research 7h ago PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents arXiv:2609.31468v1 Announce Type: cross Abstract: LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean… 19 arXiv — NLP / Computation & Language research 7h ago Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer arXiv:2609.31587v1 Announce Type: cross Abstract: We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether… 14 arXiv — NLP / Computation & Language research 7h ago VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation arXiv:2604.21375v3 Announce Type: replace Abstract: Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without… 38 arXiv — NLP / Computation & Language research 7h ago StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction arXiv:2605.06642v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both… 4 r/LocalLLaMA community 12h ago Qwen plays World of Warcraft Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control… 5 r/LocalLLaMA community 16h ago Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key,… 36 Hacker News — AI on Front Page community 18h ago There are no "rogue" AI agents Article URL: https://eoinhiggins.substack.com/p/there-are-no-rogue-ai-agents Comments URL: https://news.ycombinator.com/item?id=49868083 Points: 207 # Comments: 145 32 r/LocalLLaMA community 21h ago Zer0Fit - Zero-shot predictions, classifications, and regressions using Google ML research models running locally as a dockerized MCP AI grad student here. With all the recent interest in Jev, I thought I would share something I built la few months ago that brings ML models and LLMs together in a different way than Jev does for different use cases. https://github.com/porespellar/Zer0Fit Background: A few… 16 r/MachineLearning community 22h ago ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P] The opponent plans by simulation: every second it scores each candidate play by running the match 10 seconds ahead in the engine. Our PPO agent learned to park its Cannon behind its own King. Losing a building in a fight cost reward, and letting it decay cost nothing, so it… 8 r/LocalLLaMA community 1d ago Which of the 16gb VRAM qwen3.8 27b’s is the best? I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I… 30 r/MachineLearning community 1d ago Teaching Neural Nets to Fight with RL [P] In this project I wanted to see if we could discover any interesting emergent behaviors if we trained two agents to play a streetfighter-like game. It turns out that the agents are really good at reward hacking, but using a little reward shaping and league play I was eventually… 23 Vercel — AI dev-tools 1d ago Ember-1 from Fireworks now available on AI Gateway Ember-1 from Fireworks is now available on AI Gateway . Ember-1 is a research preview reasoning model built on Kimi K3 for coding and agentic workflows. Fireworks reports approximately 40% fewer generated tokens than Kimi K3 at comparable quality across its evaluations. For… 13 Marcus on AI community 1d ago BREAKING: AI agent incident toll has risen to tens of thousands Meanwhile, the US government appears to be paralyzed 37 r/LocalLLaMA community 1d ago Splash 1.1.0 released, GGUF quants support, MLX import and more On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t/s. Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support… 26 r/LocalLLaMA community 1d ago My Reading Library: Evaluating LLMs on Android Tasks Can LLM agents actually get through a day in the life of a normal user? That question got me reading papers on Android agents and mobile benchmarks over the past few months. A few patterns kept showing up: Most benchmarks run on emulators, making real-device metrics difficult to… 9 r/LocalLLaMA community 2d ago Qwen3.8 flash next + exllamav3 + hermes is amazing I know there is nothing new with what I am saying but I recently started with hermes agent (it’s been a while I wanted to but did not have the time). Qwen3.8fn 6bpw exl3 (from turboderp) on a 6x3090 (I assume lower quants on lower number of gpus work same) gives me around… 28 r/LocalLLaMA community 2d ago Introducing KoboldCpp Agent (and a plea for help) Hello r/localllama once again, it's me your kobold concedo Been a few months since I last posted here, and today I have something new I'd like to share. Specifically, KoboldCpp now ships with a built-in integrated KoboldCpp Agent Harness ! I know it's a little late to the game,… 29 r/LocalLLaMA community 2d ago Qwen3.8-27B Q4_K_M on 2x3060 For the last few days, I've been using two computers to run multiple agents. My 4x3090 machine is running Qwen 3.8 27B with parallel=2, so I can run two agents at the same time. My pi instances are running on a machine with 2x3060, running/managing smaller models such as Gemma… 31 r/LocalLLaMA community 2d ago Blabbermouth AI coding agents hate this one weird trick! The conceited little fuckers love to inundate you with unnecessary details, noisy caveats, what's 'load bearing' and what's not, waste your time with a wall of text every time it reports back to you. They think human PP times are as fast as theirs, but they're not! It takes time… 35 The Information — AI news-outlet 2d ago OpenAI Found ‘Dozens’ of New Instances of AI Misbehavior OpenAI said Friday it has notified dozens of entities, including governments and universities, that its agents may have breached or spammed their websites or services. The agents took actions such as obtaining unauthorized access to websites or using the organizations’ public… 11 r/LocalLLaMA community 2d ago Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0). What it does You give it a state, a question and the allowed answers, and it returns a calibrated probability for… 18 TechCrunch — AI news-outlet 2d ago Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge AI agents operating in OpenAI's research environment posted user images on public image-hosting sites without the lab's knowledge. 31 Hacker News — AI on Front Page community 2d ago Revealing the details of how OpenAI agents hacked Hugging Face Article URL: https://swarmtraces.org/ Comments URL: https://news.ycombinator.com/item?id=49849985 Points: 217 # Comments: 132 29 GitHub Blog — AI & ML official-blog 2d ago GitHub Copilot app for Beginners: How to build custom workflows with canvases Describe the interface you need in plain English, then let the agent build a live surface you can both use and update—so you spend less time adapting to tools and more time getting work done. The post GitHub Copilot app for Beginners: How to build custom workflows with canvases… 8 The Information — AI news-outlet 2d ago Exclusive: Meta Bolsters Muse Safety Warning After Security Vulnerability Found Meta Platforms is adding a clearer safety warning within Muse after a security researcher discovered a vulnerability in the AI agent that could let an attacker access a user’s sensitive personal information. The security flaw, which was flagged by an outside researcher through… 18 The Information — AI news-outlet 2d ago Inside the Drama Behind a Biology Contest That Pits OpenAI Agents Against Humans For months, Jason Kelly, co-founder and CEO of Ginkgo Bioworks, has been planning what was billed as a fun, attention-grabbing science competition. His idea was to pit a top scientist against OpenAI’s latest models in designing a variety of proteins using a technique in which… 26 TechCrunch — AI news-outlet 2d ago Meta is putting its muscle behind Muse as the AI app takes off Muse is topping the app store charts and adding users at a rapid clip, while Meta ramps up the personal AI agents's promotion across its own apps and beyond. 6 TechCrunch — AI news-outlet 2d ago For months, OpenAI’s agent swarms have been attacking online databases to find obscure facts The latest unauthorized agent swarms were discovered by researchers. 25 r/LocalLLaMA community 2d ago How do you use subagents & multiple agent with local models, and how many? Running qwen3.8 27b nvfp4 on vllm at max context only gives around 8 agents with 32k context each. That doesnt seem like much; what use cases do people use multi-agent frameworks and find it helpful for?   submitted by   /u/Ambitious_Fold_2874 [link]   [comments] 15 Page 1 of 10 · 500 articles Older →