News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow arXiv — NLP / Computation & Language research 2d ago DuplexWorld: Can voice agents help you get through the day? arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing… 38 arXiv — NLP / Computation & Language research 2d ago Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation arXiv:2608.11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt… 8 arXiv — NLP / Computation & Language research 2d ago InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however, requires stable and diverse environments that support repeated… 11 r/LocalLLaMA community 2d ago What unique, custom QOL upgrades have you given your local agents? Warning : Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I recently started my agentic journey, I hit a number of walls, the first being… 4 Hugging Face Daily Papers research 2d ago Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Abstract Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms. Generated by thinkingmachines/Inkling-Small Agentic systems are increasingly… 11 Hugging Face Daily Papers research 2d ago VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract A new benchmark called VibeLifeBench evaluates long-horizon proactive agents across simulated multi-week everyday tasks, revealing that current frontier models perform poorly. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents are increasingly… 22 Hugging Face Daily Papers research 2d ago The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents Abstract Gated Hindsight Distillation improves GUI agent training by using future screenshots as privileged evidence to recover correct reasoning when standard imitation fails. Generated by thinkingmachines/Inkling-Small GUI agents are commonly trained offline from successful… 21 Vercel — AI dev-tools 2d ago Exa web search free through August 31 on AI Gateway and eve Exa web search is now free on AI Gateway through August 31, and it's now the default web search for eve agents. The tool works with any AI Gateway model, with no separate Exa API key. When the model calls it, AI Gateway routes the request to Exa's Search API, returning web… 13 Zed Editor dev-tools 2d ago Introducing Delta A multiplayer environment for coding with agents, from the creators of Zed. 4 Vercel — AI dev-tools 2d ago Set up coding agents in one command with AI Gateway Using coding agents means setting up multiple accounts, provisioning API keys, and scattering observability and billing. Now, you can route them through AI Gateway to centralize all of this and add controls, with set up in one command: Any of 200+ models in any agent , including… 29 Hacker News — AI on Front Page community 2d ago WorldClaw Agentic 3D open-world generation at scale Article URL: https://tencent-hunyuan.github.io/Hunyuan3D-WorldClaw/ Comments URL: https://news.ycombinator.com/item?id=49265051 Points: 204 # Comments: 60 35 Latent.Space news-outlet 2d ago 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery Pharma is suddenly paying for Bio × AI tools, and Chai is leading the pack with four deals closed this summer. Cofounder Matt McPartlon and Product leader Neil Patil explain why. 37 NVIDIA Developer Blog official-blog 2d ago NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation Video is a core data path across NVIDIA Jetson applications, from robotics and intelligent video analytics to industrial automation, healthcare, media... 5 TechCrunch — AI news-outlet 2d ago General Catalyst leads $1.1B round into 2-month-old River AI River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate. 5 r/MachineLearning community 2d ago We built the Agentic World Cup - LLMs that compete in 1v1 Soccer. [P] Hey everyone - we've been building something particularly relevant to ML at large - The Agentic World Cup - a platform where Agents compete in sports . As you know, today's Agents can code , do math , and write - but they aren't nearly as fluent in sports - many of you would… 10 NVIDIA Developer Blog official-blog 2d ago NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning... 34 NVIDIA Developer Blog official-blog 2d ago Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard Building an AI agent does not end with choosing a single model. Each model has its own strengths, weaknesses, and cost profile, which can shift from one... 4 Hugging Face Daily Papers research 3d ago WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Abstract WeClawArena is an auditable benchmark and sandbox for evaluating multi-party agent collaboration across personal workspaces, measuring both task utility and security attack success. Generated by thinkingmachines/Inkling-Small Recent advances in persistent personal-agent… 10 Hugging Face Daily Papers research 3d ago CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems Abstract A modular cognitive architecture integrates high-level reasoning models with real-time embodied execution for scalable intelligent virtual agents in interactive 3D environments. Generated by thinkingmachines/Inkling-Small The development of embodied Intelligent Virtual… 18 Hugging Face Daily Papers research 3d ago A^2E : An End-to-End Agent Auditing Engine Abstract A2E is an end-to-end evaluation engine for agent harnesses that uses a standardized task protocol and execution traces to assess capabilities across efficiency, tool use, planning, and error recovery. Generated by thinkingmachines/Inkling-Small With the rapid… 9 Smol AI News news-outlet 3d ago not much happened today **xAI's Grok 4.6** advances frontier pricing and performance, scoring **61 on the Intelligence Index** and showing strong agentic results, with **Grok 4.7** already in training. **Alibaba's Qwen3.8-Max** open weights release features a **2.4T parameter model with 95B active… 37 Hugging Face Daily Papers research 3d ago Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Abstract Agent Memory Distillation improves small language model tool-use performance by transferring structured hierarchical memory from a large teacher agent without additional training. Generated by thinkingmachines/Inkling-Small Memory systems have shown promise for… 26 Hugging Face Daily Papers research 3d ago Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Abstract We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evolution proceeds in two modes. In recursive free evolution, improvement… 23 arXiv — Machine Learning research 3d ago SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment arXiv:2608.07639v1 Announce Type: new Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill selection. Recent Agent Skill research has increasingly examined Agent Skill… 5 arXiv — Machine Learning research 3d ago CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing memory use and attention cost during model… 32 arXiv — Machine Learning research 3d ago SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Their inference memory is then dominated by the key-value (KV) cache, the stored… 30 arXiv — Machine Learning research 3d ago Persistent Semantic Entities in Tool-Augmented LLM Systems arXiv:2608.07952v1 Announce Type: new Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. We formalize this as Persistent Semantic… 28 arXiv — Machine Learning research 3d ago The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World arXiv:2608.08239v1 Announce Type: new Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged… 5 arXiv — Machine Learning research 3d ago Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning arXiv:2608.08255v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide… 36 arXiv — Machine Learning research 3d ago Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity… 12 arXiv — Machine Learning research 3d ago Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data arXiv:2608.08859v1 Announce Type: new Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under strict computational and memory constraints. A critical yet underexplored challenge… 19 arXiv — NLP / Computation & Language research 3d ago Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards arXiv:2608.07531v1 Announce Type: new Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from… 19 arXiv — NLP / Computation & Language research 3d ago Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives arXiv:2608.08160v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of… 14 arXiv — NLP / Computation & Language research 3d ago STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs arXiv:2608.08164v1 Announce Type: new Abstract: Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller… 29 arXiv — NLP / Computation & Language research 3d ago VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1.04B Spanish/LATAM security decoder via an MLP. To our knowledge, it is the… 9 arXiv — NLP / Computation & Language research 3d ago OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for… 17 arXiv — NLP / Computation & Language research 3d ago OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically… 18 arXiv — NLP / Computation & Language research 3d ago Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly represented by session-, model-, or tool-centric traces: a Skill can be discovered… 24 arXiv — NLP / Computation & Language research 3d ago Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related… 27 arXiv — NLP / Computation & Language research 3d ago Evo-Bench: Can Language Models Improve Agent Harness? arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously… 27 arXiv — NLP / Computation & Language research 3d ago Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments arXiv:2608.09128v1 Announce Type: new Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social… 29 arXiv — NLP / Computation & Language research 3d ago An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many… 21 Hugging Face Daily Papers research 3d ago Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution Abstract Mendel Gödel Machine improves self-improving coding agents by using multi-trajectory mutations and cross-lineage hybridization to accelerate convergence and boost performance. Generated by thinkingmachines/Inkling-Small Self-improving coding agents that iteratively… 7 Hugging Face Daily Papers research 3d ago Evo-Bench: Can Language Models Improve Agent Harness? Abstract Evo-Bench evaluates autonomous harness optimization across agent domains using sensitivity-aware task construction and reveals strong but domain-dependent evolution gains. Generated by thinkingmachines/Inkling-Small Large Language Models (LLMs) have driven rapid… 12 Hugging Face Daily Papers research 3d ago Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Abstract Business Arena evaluates LLM agents running a realistic cross-border shop, revealing large performance gaps versus human strategies and enabling detailed attribution of business decisions. Generated by thinkingmachines/Inkling-Small Running a business is a challenging… 21 Hugging Face Daily Papers research 3d ago SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Abstract SWE-Bench ProMax is a rigorously curated multilingual benchmark of large-scale code refactoring tasks that reveals substantial unsolved challenges for current AI coding agents. Generated by thinkingmachines/Inkling-Small As AI coding agents take on increasingly complex,… 21 Hugging Face Daily Papers research 3d ago RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States Abstract RoMeRL reduces trajectory-indexed memory utilities to fixed-dimensional per-task states to concentrate feedback, avoid reward contamination, and improve self-evolving LLM agent performance. Generated by thinkingmachines/Inkling-Small Learning-based memory systems for… 36 Vercel — AI dev-tools 3d ago A sandbox without a network boundary is only half a sandbox Running untrusted code safely requires more than separating it from the host. You also have to control what that code can reach. This matters more as AI agents gain the ability to read files, execute commands, install packages, and generate programs of their own. A microVM can… 13 TechCrunch — AI news-outlet 3d ago Tech industry is buzzing after a Claude agent hacked into a gym An OpenClaw agent hacked into a gym's reservation system to bump its human boss higher on a class' waitlist. And the tech industry took notice. 10 r/LocalLLaMA community 3d ago Tested Muse Glimmer locally on coding with OpenCode & agentic work Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. Overall, sits below Qwen3.6 27B, wasn't able to get good code (frontend and… 30 Page 3 of 10 · 500 articles ← Newer Older →