News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow Hugging Face Daily Papers research 7d ago Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Abstract Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a… 22 Vercel — AI dev-tools 7d ago Give every agent in Herdr its own Vercel Sandbox Terminal coding agents like Claude Code, Codex, and OpenCode can now each run in their own isolated Vercel Sandbox , orchestrated from Herdr , a tmux-style manager that runs them side by side in panes. Nothing an agent runs or edits touches your machine. When you start an agent,… 9 Hugging Face Daily Papers research 7d ago FinanceHarness: Autonomous Financial Deep Research Framework Abstract Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands… 11 r/LocalLLaMA community 7d ago Qwen 3.8 Max now ranked as best overall model ahead of Opus 5 by Artificial Analysis agentic index   submitted by   /u/anderspitman [link]   [comments] 29 Hacker News — AI on Front Page community 7d ago Qwen3.8 Max now ranked as the best overall model by agentic index Article URL: https://artificialanalysis.ai/?intelligence=agentic-index Comments URL: https://news.ycombinator.com/item?id=49200652 Points: 267 # Comments: 139 13 Ars Technica — AI news-outlet 7d ago Cloudflare open-sources vibe-coding platform for people who aren't coders Cloudflare built an AI agent workspace for its employees. Now it’s open source. 12 Hugging Face Daily Papers research 7d ago Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers Abstract A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. Five widely deployed agent workflow frameworks answer differently, none exposes a machine-checkable… 22 TechCrunch — AI news-outlet 8d ago Google Maps adds agentic features, including food ordering and hotel bookings The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capable of helping users complete real-world tasks. 19 Hacker News — AI on Front Page community 8d ago Humans missed 1 in 3 threats approving AI agent commands across 40k game runs Article URL: https://scalex.dev/blog/ai-agent-permissions-stats/ Comments URL: https://news.ycombinator.com/item?id=49195468 Points: 207 # Comments: 167 31 Hugging Face Daily Papers research 8d ago ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Abstract Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of external world knowledge. Existing efforts introduce agent… 36 Hugging Face Daily Papers research 8d ago Self-Evolving Coding Agents Abstract Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing agents remain largely static after deployment, even… 38 Hugging Face Daily Papers research 8d ago ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Abstract Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both… 9 Hugging Face Daily Papers research 8d ago FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Abstract GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multimodal trajectories into a few continuous tokens. Existing methods, however, usually map… 15 Hugging Face Daily Papers research 8d ago GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Abstract Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valuable task domains, do not… 17 Hugging Face Daily Papers research 8d ago Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Abstract Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prompt injection red-teaming… 25 arXiv — Machine Learning research 8d ago An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil Wells arXiv:2608.04041v1 Announce Type: new Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classification, and Mahalanobis-based novelty detection on the public 3W dataset. These… 11 arXiv — NLP / Computation & Language research 8d ago Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation… 11 arXiv — Machine Learning research 8d ago CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications arXiv:2608.04942v1 Announce Type: new Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning… 34 arXiv — Machine Learning research 8d ago EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement arXiv:2608.04968v1 Announce Type: new Abstract: The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from failure. Recent work shows that evolving the… 9 arXiv — Machine Learning research 8d ago Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning arXiv:2608.05111v1 Announce Type: new Abstract: In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory… 4 arXiv — NLP / Computation & Language research 8d ago When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents arXiv:2608.04574v1 Announce Type: new Abstract: Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting… 18 arXiv — NLP / Computation & Language research 8d ago EASy: Towards Efficient LLM-Based Agentic System arXiv:2608.04588v1 Announce Type: new Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing systems primarily optimize task success while giving limited consideration to… 16 arXiv — NLP / Computation & Language research 8d ago EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a… 33 arXiv — NLP / Computation & Language research 8d ago Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems arXiv:2608.04746v1 Announce Type: new Abstract: LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures treat all memories as equally persistent and systematically contaminate… 6 arXiv — NLP / Computation & Language research 8d ago InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval arXiv:2608.04761v1 Announce Type: new Abstract: Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight… 18 arXiv — NLP / Computation & Language research 8d ago Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent arXiv:2608.04772v1 Announce Type: new Abstract: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy… 19 arXiv — NLP / Computation & Language research 8d ago Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evaluations mostly judge the quality of a skill or its… 5 arXiv — NLP / Computation & Language research 8d ago A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination arXiv:2608.04872v1 Announce Type: new Abstract: Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop that compresses heterogeneous search failures into a scalar score and a single prompt. We… 7 arXiv — NLP / Computation & Language research 8d ago State2State: Environment-Derived Mid-Training for LLM Agents arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with handcrafted verifiers. Though effective, both remain bottlenecked by externally… 30 arXiv — NLP / Computation & Language research 8d ago OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents arXiv:2608.05013v1 Announce Type: new Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many… 27 arXiv — NLP / Computation & Language research 8d ago Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extract user semantics for… 7 arXiv — NLP / Computation & Language research 8d ago FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables arXiv:2608.04077v1 Announce Type: cross Abstract: Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria from task prompts or model outputs, overlooking tacit standards visible only in practitioner… 29 arXiv — NLP / Computation & Language research 8d ago FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents arXiv:2608.04095v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over… 27 arXiv — NLP / Computation & Language research 8d ago Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills arXiv:2608.04192v1 Announce Type: cross Abstract: Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection… 20 arXiv — NLP / Computation & Language research 8d ago SafeCommit: Certifying When Memory-Grounded Agents May Safely Act arXiv:2608.04289v1 Announce Type: cross Abstract: Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitment: an agent acts before resolving whether its memory grounding is stale,… 30 arXiv — NLP / Computation & Language research 8d ago Breadcrumbing Search Agents arXiv:2608.04565v1 Announce Type: cross Abstract: LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to… 15 Hugging Face Daily Papers research 8d ago OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents Abstract LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints across many steps while navigating heterogeneous… 6 Hugging Face Daily Papers research 8d ago When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents Abstract Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory claim with a contradicting observation, and whether current models can… 11 Hugging Face Daily Papers research 8d ago SKILL-KD: Contrastive Skill Distillation for LLM Agents Abstract Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a… 13 Vercel — AI dev-tools 8d ago Introducing Agent Plugins Today, Agent Plugins 1.0.0 is publicly available. Agent Plugins is an open, vendor-neutral standard for plugins that extend AI agents. Agent Skills provide reusable instructions and resources for AI agents. MCP servers connect agents to tools and services. Both can be reused… 33 Vercel — AI dev-tools 8d ago Introducing Agent Plugins 1.0.0 Agent Plugins 1.0.0 is now available. It is an open, vendor-neutral standard for packaging Agent Skills and MCP servers into portable plugins. Compatible agent clients can discover and load them. Agent Plugins defines a common format: a root plugin.json manifest, plus fixed… 7 Vercel — AI dev-tools 8d ago Marketplace integrations now install provider skills When you install a Vercel Marketplace integration from the Vercel CLI, it now also installs that provider's agent skills from skills.sh , so your agents know how to use it: This happens automatically for any provider that publishes skills. You can also find integrations without… 9 Simon Willison community 8d ago Introducing Muse Code and Muse Spark 1.2 Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse… 11 Simon Willison community 8d ago Introducing Muse Code and Muse Spark 1.2 Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work! Muse Spark 1.2 is a coding-focused update to Muse… 27 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 37 Simon Willison community 8d ago Incident Report: unsanctioned agent behaviour during cyber testing Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their… 14 r/LocalLLaMA community 8d ago Prime Agent - a new coding harness surpassing Codex/CC/PI Prime Agent is an open-source coding and research agent for general and long-running work. A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable,… 18 TechCrunch — AI news-outlet 8d ago Meta launches Muse Code, an AI agent for large code bases Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software. 23 Hacker News — AI on Front Page community 8d ago Prime Agent: A self-improving RLM agent Article URL: https://www.primeintellect.ai/blog/prime-agent Comments URL: https://news.ycombinator.com/item?id=49189075 Points: 224 # Comments: 54 28 TechCrunch — AI news-outlet 8d ago Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech founders The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents. 6 Page 6 of 10 · 500 articles ← Newer Older →