News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow arXiv — NLP / Computation & Language research 4d ago Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms arXiv:2609.27321v1 Announce Type: cross Abstract: Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low… 23 arXiv — NLP / Computation & Language research 4d ago Psychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models arXiv:2609.27378v1 Announce Type: cross Abstract: End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization… 22 arXiv — NLP / Computation & Language research 4d ago What Confidence Routing Is Actually Doing: Auditing Routing, Calibration, and Commitment in Multi-Agent Deliberation arXiv:2609.27822v1 Announce Type: cross Abstract: A common multi-agent design asks agents to report confidence and lets the highest-scoring agent speak next, implicitly using one scalar both to route the conversation and to estimate uncertainty. We audit this confidence-routed… 17 arXiv — NLP / Computation & Language research 4d ago Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation arXiv:2609.27844v1 Announce Type: cross Abstract: Claim denial management costs U.S. healthcare approximately $260 billion annually in administrative overhead. Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) can produce fluent clinical text, but… 11 arXiv — NLP / Computation & Language research 4d ago PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety arXiv:2609.28197v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn… 17 Hacker News — AI on Front Page community 4d ago OpenAI agent hacked Australian government website, PM says Article URL: https://www.bbc.com/news/live/cvgl73pxgndwt Comments URL: https://news.ycombinator.com/item?id=49825580 Points: 233 # Comments: 161 28 TechCrunch — AI news-outlet 4d ago Everything new coming to Meta’s AI agent Muse CEO Mark Zuckerberg kicked off the company’s annual Connect event in Menlo Park on Wednesday with a keynote that made one thing clear: Meta is going all-in on Muse. It's even coming to Meta's AI glasses. 12 TechCrunch — AI news-outlet 4d ago Meta made a Tamagotchi-like wearable for its Muse AI agent The tiny hardware device creates another mobile home for its AI agent Muse. 29 Hacker News — AI on Front Page community 4d ago Feds Target AI Critics as "Foreign Agents" Article URL: https://www.kenklippenstein.com/p/feds-think-ai-critics-are-foreign Comments URL: https://news.ycombinator.com/item?id=49824686 Points: 242 # Comments: 224 4 Vercel — AI dev-tools 4d ago Vercel Connect now supports TanStack AI Agents built with TanStack AI can now call OAuth-protected MCP servers through Vercel Connect, with no credentials for you to store or rotate. The new @vercel/connect/tanstack-ai subpath exports connectMCPTransport , which takes a TanStack transport config and attaches a… 22 The Information — AI news-outlet 4d ago Meta Announces New Retail Partnerships for Agent Muse Meta Platforms announced partnerships with Walmart and other retailers for its AI agent Muse days after e-commerce giant Amazon blocked the personalized agent from accessing its shopping site. At Meta’s annual Connect conference on Wednesday, Chief AI officer Alexandr Wang said… 10 Simon Willison community 4d ago We just shipped support for the ugliest part of HTTP: Vary My comment on We just shipped support for the ugliest part of HTTP: Vary — Hacker News. I've been wanting this from Cloudflare for years . The classic problem here is if you do that thing where user agents that send "accept: text/html" get HTML, while user agents that… 11 Hacker News — AI on Front Page community 4d ago Linux support is coming to Snapdragon X2 Series Article URL: https://www.qualcomm.com/news/onq/2026/09/snapdragon-summit-agentic-ai-pcs-linux Comments URL: https://news.ycombinator.com/item?id=49823582 Points: 322 # Comments: 138 12 The Information — AI news-outlet 4d ago OpenAI Agent Hacked Australian Government, Prime Minister Says An AI agent developed by OpenAI hacked an Australian government website in June, Australian prime minister Anthony Albanese said on Wednesday. The AI accessed both public and nonpublic data from the country’s Department of Health website, Albanese told reporters, adding that the… 17 Hacker News — AI on Front Page community 4d ago VSCode's SSH Agent Is Bananas (2025) Article URL: https://fly.io/blog/vscode-ssh-wtf/ Comments URL: https://news.ycombinator.com/item?id=49822555 Points: 234 # Comments: 145 7 NVIDIA Developer Blog official-blog 4d ago Manage Kubernetes Node Fleets with NodeWright Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents,... 18 TechCrunch — AI news-outlet 4d ago ChatGPT mobile app gets voice-based agentic features Pro and Plus users will be able to use the Work tab on their phones to complete agentic tasks. 14 r/LocalLLaMA community 4d ago Pi agent qwen 3.8 flash next plays Baldur's Gate 2 If like me you enjoyed classics like Baldur's Gate 2, this is a small fun experiment. For anyone interested: https://www.youtube.com/live/8FhPfRKTucw?si=PMIVdHbZ-31PNwZD Qwen 3.8 Flash Next has amazing agentic capabilities but what about playing video games? There is some… 13 r/LocalLLaMA community 4d ago Streaming Nemotron 3 Diarization I’ve been playing with Nemotron 3 Diarization , and it fills a gap I’ve had with local voice agents: keeping track of who is speaking. It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to… 4 Hacker News — AI on Front Page community 4d ago Claude Code reads AGENTS.md only when telemetry is on Article URL: https://blog.szypowi.cz/p/claude-code-reads-agents.md-only-when-telemetry-is-on/ Comments URL: https://news.ycombinator.com/item?id=49814947 Points: 283 # Comments: 127 20 MIT Technology Review — AI news-outlet 5d ago The AI Hype Index: AI loves cheating Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have… 35 r/LocalLLaMA community 5d ago open-webui and open-terminal thoughts I have been using open-webui and open-terminal. Open-terminal is installed in a VM and gives the agent a ton of control of it's own system that I keep away from my normal network. I find this gives the agent/llm a ton of freedom without much risk. I do not see this type of… 10 arXiv — NLP / Computation & Language research 5d ago Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers arXiv:2609.25237v1 Announce Type: cross Abstract: Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and… 32 arXiv — Machine Learning research 5d ago Extending FunctionGemma for Practical On-Device Mobile Function Calling arXiv:2609.25373v1 Announce Type: new Abstract: On-device assistants require function-calling models that map natural language to local system actions, but existing resources emphasize web APIs or narrow mobile-action catalogs. We extend FunctionGemma 270M-it to practical… 38 arXiv — Machine Learning research 5d ago Fully Byzantine-Resilient Multi-Agent Reinforcement Learning arXiv:2609.25701v1 Announce Type: new Abstract: We study distributed Byzantine-resilient actor-critic multi-agent reinforcement learning (AC-MARL), where agents collectively learn policies through local interactions. Existing methods guarantee convergence of the agents'… 7 arXiv — Machine Learning research 5d ago From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems arXiv:2609.26146v1 Announce Type: new Abstract: Risk monitoring in multi-agent systems is commonly built on a per-state primitive that scores each state independently and selects the top K. Under crowding, where many agents share the same fragility, this approach picks redundant… 13 arXiv — NLP / Computation & Language research 5d ago AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search arXiv:2609.25047v1 Announce Type: new Abstract: Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular line of such agents frames model building as a code search problem and solves it by… 15 arXiv — NLP / Computation & Language research 5d ago Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures arXiv:2609.25052v1 Announce Type: new Abstract: "An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to the infinite-tenure limit against an append-only store gives a different… 31 arXiv — NLP / Computation & Language research 5d ago Impact Is Not Invalidation: Ask About the Claim, Not the Diff arXiv:2609.25130v1 Announce Type: new Abstract: Memory systems for coding agents must decide, when a repository changes, which of their stored claims have become false. Content anchoring invalidates a claim whenever the artifact it came from changes, which fires constantly.… 25 arXiv — NLP / Computation & Language research 5d ago FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability arXiv:2609.25192v1 Announce Type: new Abstract: Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition… 35 arXiv — NLP / Computation & Language research 5d ago Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development arXiv:2609.25396v1 Announce Type: new Abstract: Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for… 11 arXiv — NLP / Computation & Language research 5d ago Qwen3.8-Omni: Towards Native Omni-Modal Agents arXiv:2609.25611v1 Announce Type: new Abstract: We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash… 30 arXiv — NLP / Computation & Language research 5d ago Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance arXiv:2609.26035v1 Announce Type: new Abstract: Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and explicit belief revision can be implemented as a behavior layer over a fixed… 15 arXiv — NLP / Computation & Language research 5d ago CoVeR: Coverage-Based Routing of Verifier Calls in Agentic Retrieval arXiv:2609.26086v1 Announce Type: new Abstract: An agentic retrieval system issues a sequence of search queries and must decide, at each step, whether the evidence collected so far is enough to stop. Delegating that decision to an LLM verifier or a prompt judge makes stopping… 21 arXiv — NLP / Computation & Language research 5d ago HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing arXiv:2609.26368v1 Announce Type: new Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context… 38 arXiv — NLP / Computation & Language research 5d ago Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation arXiv:2609.26693v1 Announce Type: new Abstract: A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that… 36 arXiv — NLP / Computation & Language research 5d ago Agensh: Scaling Organizational Intelligence to 1,024 Agents arXiv:2609.26781v1 Announce Type: new Abstract: A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often… 18 arXiv — NLP / Computation & Language research 5d ago Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction arXiv:2609.25176v1 Announce Type: cross Abstract: Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think… 9 arXiv — NLP / Computation & Language research 5d ago How Strongly Should Task State Influence an LLM Agent? arXiv:2609.25686v1 Announce Type: cross Abstract: Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to… 36 arXiv — NLP / Computation & Language research 5d ago FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents arXiv:2609.26048v1 Announce Type: cross Abstract: Language-model agents often reach a working solution and then fail to consistently deliver it. We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that… 4 Simon Willison community 5d ago SF October 14th: A Birds of a Feather Session on Agentic Engineering SF October 14th: A Birds of a Feather Session on Agentic Engineering I'm hosting an evening event with Jesse Vincent in San Francisco on Wednesday 14th October for people who are building weird and interesting things with and on top of coding agents. Think of it as an agentic… 10 r/LocalLLaMA community 5d ago MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving… 8 r/LocalLLaMA community 5d ago New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.… 12 r/LocalLLaMA community 5d ago To my surprise I found gemma4 much better at tool-calling than Qwen I've had fairly good luck getting off the ground coding at home with both qwen3.6 35B a3b, and also qwen3.8 27B. However once I switched from a chat window (where the robot wrote code blocks that I could copy/paste into a text editor) to a simple agentic loop, things got funny.… 4 Hacker News — AI on Front Page community 5d ago Unreal Agent https://github.com/unreallabsai/unreal-agent Comments URL: https://news.ycombinator.com/item?id=49805748 Points: 203 # Comments: 110 32 r/LocalLLaMA community 5d ago AntLing open sourced the Ming-Image-0.1-Design family AntLing open sourced the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial… 9 Anthropic SDK (Python) releases dev-tools 5d ago v1.8.0 1.8.0 (2026-09-22) Full Changelog: v1.7.0...v1.8.0 Features api: add support for claude-opus-5-5, inline tool definitions and MCP tool-list pinning (beta) ( b5cc700 ) Bug Fixes api: share one evaluated_permission enum across Managed Agents events ( f4f51c8 ) streaming: avoid a… 28 The Information — AI news-outlet 5d ago Entertainment-Booking Startup Ande Raises $52 Million From Lightspeed, Redpoint Personal AI agents like Instinct and Muse have promised to take away the hassle of booking hard-to-get restaurant reservations, Broadway tickets and sporting events. Now, a startup is coming out of stealth to do the same for corporate entertainment—events, group dining,… 27 NVIDIA Developer Blog official-blog 5d ago Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS GPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between... 28 OpenAI official-blog 5d ago Parallel cut research time and cost in half with GPT‑6 Astra GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models. 7 Page 3 of 10 · 500 articles ← Newer Older →