News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow arXiv — NLP / Computation & Language research 1d ago FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents arXiv:2608.11683v1 Announce Type: cross Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that… 25 arXiv — NLP / Computation & Language research 1d ago The Sleeping Agent: What Gist-Based Context Compression Loses and Why arXiv:2608.11775v1 Announce Type: cross Abstract: Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon language model agents, yet its effect on different types of memory retrieval is poorly… 38 arXiv — NLP / Computation & Language research 1d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents arXiv:2608.11878v1 Announce Type: cross Abstract: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused… 25 arXiv — NLP / Computation & Language research 1d ago Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation arXiv:2608.12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has… 5 arXiv — NLP / Computation & Language research 1d ago AVA-Encoder: Towards Agent-Native Video Representation Learning arXiv:2608.12313v1 Announce Type: cross Abstract: Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both… 26 Hugging Face Daily Papers research 1d ago MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Abstract Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only… 21 Hugging Face Daily Papers research 1d ago Agent Safety Should Be a Runtime Contract Abstract Agent safety should be enforced at runtime through preventive controls and verifiable evidence rather than relying solely on training-time alignment methods. Generated by thinkingmachines/Inkling-Small The dominant paradigm treats AI safety as a property to be instilled… 24 Hugging Face Daily Papers research 1d ago ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Abstract ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment. Generated by thinkingmachines/Inkling-Small Large language model (LLM) agents… 12 Vercel — AI dev-tools 1d ago Use ACP-compatible harnesses with the AI SDK harness layer The AI SDK harness layer now supports any Agent Client Protocol (ACP)-compatible harness with HarnessAgent through the new @ai-sdk/harness-acp package. Previously, every harness adapter wrapped one specific runtime (Claude Code, Codex, Pi, Deep Agents, OpenCode).… 23 Hugging Face official-blog 1d ago What We Learned by Reproducing 2,200 papers from ICML Back to Articles a]:hidden"> What We Learned by Reproducing 2,200 papers from ICML Published August 13, 2026 Update on GitHub Upvote 3 Abubakar Abid abidlabs Back in July, we ran a hackathon where more than 1,200 community members brought their own coding agents and tried to… 30 Vercel — AI dev-tools 1d ago Exa joins the Vercel Agent Marketplace Exa is now available on the Vercel Agent Marketplace as a native integration. Exa's neural search engine delivers high-quality, relevant results to ground AI in fresh, current information. Add Exa to your Vercel app in seconds to power search, research agents, and context-aware… 33 Vercel — AI dev-tools 1d ago Gemini 3.7 Flash now available on AI Gateway for 50% off Gemini 3.7 Flash from Google is now available on AI Gateway for 50% off till December 31st, 2026. Gemini 3.7 Flash improves on prior Flash models at software engineering and agentic work. It resolves issues more reliably and spends less time stuck in failed agent loops, which… 12 Vercel — AI dev-tools 1d ago GLM 5.2 free for eve agents through August 27 via Blackbox on AI Gateway GLM 5.2 , the open-weights coding model from Z.ai with a 1M-token context window, is free for eve agents through August 27, served by Blackbox AI on AI Gateway . New eve agents come with GLM 5.2 as their default model. Use npx eve@latest init my-agent to get started. Existing… 35 Hugging Face Daily Papers research 1d ago DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Abstract DSAgentBench evaluates autonomous agents on complete, multi-tool data-science workflows in real computing environments and reveals major performance gaps. Generated by thinkingmachines/Inkling-Small Real-world data science involves long-horizon workflows that span data… 33 Hugging Face Daily Papers research 1d ago SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Abstract SkillZip compresses self-evolving agent skills by finding a minimal faithful structural explanation that shares repeated rules and procedures while preserving rare exceptions, without requiring evaluation rollouts. Generated by thinkingmachines/Inkling-Small… 26 MIT Technology Review — AI news-outlet 1d ago Scaling AI agents with trustworthy data Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment… 34 r/LocalLLaMA community 1d ago According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1 In OCP APAC 2026, Tai AMD SVP of compute and enterprise AI said agents don't cut GPU demand but they just pile on a whole extra layer of orchestration, retrieval, and tool-calling work that runs on CPUs instead And the usual 1:4 CPU-to-GPU ratio could move toward 1:2 or even 1:1… 23 Hugging Face Daily Papers research 1d ago InSight-doc: Agentic Visual Perception for Long-Document Understanding Abstract InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations. Generated by thinkingmachines/Inkling-Small Long-document understanding often requires reasoning over many visually rich… 27 Hugging Face Daily Papers research 2d ago 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Abstract A new photorealistic urban benchmark reveals large performance gaps for embodied agents in city-scale navigation and spatial reasoning. Generated by thinkingmachines/Inkling-Small We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of… 8 Hugging Face Daily Papers research 2d ago Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Abstract The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models. Generated by thinkingmachines/Inkling-Small The rapid advancement of Large Language… 16 OpenAI official-blog 2d ago From assistance to execution: How enterprises put AI to work OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption. 30 Hugging Face Daily Papers research 2d ago Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Abstract Pruning strategies applied at different pipeline stages reduce token usage and latency in long-horizon research agents, with early pruning yielding the greatest efficiency gains. Generated by thinkingmachines/Inkling-Small Long-horizon research agents solve open-ended… 10 Hugging Face Daily Papers research 2d ago ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Abstract Combodied Agents integrate digital and embodied tools into a closed-loop framework that models individual human-state trajectories over time to provide proportionate, consent-aware support. Generated by thinkingmachines/Inkling-Small After an older adult misses a… 30 arXiv — Machine Learning research 2d ago Sheaf-Based Federated Representation Learning arXiv:2608.10016v1 Announce Type: new Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model architectures, latent dimensionalities, and local learning… 6 arXiv — Machine Learning research 2d ago DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents arXiv:2608.10037v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Existing studies mainly focus on improving the… 38 arXiv — Machine Learning research 2d ago FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows arXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into explicit execution structures. However,… 20 arXiv — Machine Learning research 2d ago UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs arXiv:2608.10042v1 Announce Type: new Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark… 29 arXiv — Machine Learning research 2d ago MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods… 34 arXiv — Machine Learning research 2d ago Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks arXiv:2608.10357v1 Announce Type: new Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy… 20 arXiv — Machine Learning research 2d ago TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times.… 16 arXiv — Machine Learning research 2d ago Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique arXiv:2608.10430v1 Announce Type: new Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection… 14 arXiv — NLP / Computation & Language research 2d ago Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth… 30 arXiv — NLP / Computation & Language research 2d ago Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization arXiv:2608.10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total… 25 arXiv — Machine Learning research 2d ago SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning arXiv:2608.09967v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to interpret. We introduce SPOT (Sampling Policy Observation Tree), a novel… 35 arXiv — NLP / Computation & Language research 2d ago LLM Agents Factory: Retrieval of Domain-Specific LLM Agents arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by the computational cost and instability associated with the… 24 arXiv — NLP / Computation & Language research 2d ago Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems arXiv:2608.10216v1 Announce Type: new Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, semantic caches, drift guards, and answer grader gates deploy to answer the question:… 32 arXiv — NLP / Computation & Language research 2d ago Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic… 36 arXiv — NLP / Computation & Language research 2d ago Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks.… 21 arXiv — NLP / Computation & Language research 2d ago Most biomedical publications show signs of LLM-assisted writing arXiv:2608.10715v1 Announce Type: new Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about… 25 arXiv — NLP / Computation & Language research 2d ago Mitigating Context Interference for Reliable and Efficient Search Agents arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts of multi-turn search agents are lengthy and… 38 arXiv — NLP / Computation & Language research 2d ago VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? arXiv:2608.10875v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs… 20 arXiv — NLP / Computation & Language research 2d ago What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such a probe measures, in a construction chosen to make the question sharp: a ring… 5 arXiv — NLP / Computation & Language research 2d ago Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product:… 38 arXiv — NLP / Computation & Language research 2d ago TRIBE: Predicting Team Performance via Communication Behavior Ensembles arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics… 30 arXiv — NLP / Computation & Language research 2d ago OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that… 31 arXiv — NLP / Computation & Language research 2d ago Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems arXiv:2608.10218v1 Announce Type: cross Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate… 9 arXiv — NLP / Computation & Language research 2d ago DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? arXiv:2608.10366v1 Announce Type: cross Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and… 19 arXiv — NLP / Computation & Language research 2d ago Evaluating Rational Contracting in Natural Language arXiv:2608.10475v1 Announce Type: cross Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements… 37 arXiv — NLP / Computation & Language research 2d ago InSight-doc: Agentic Visual Perception for Long-Document Understanding arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual… 17 arXiv — NLP / Computation & Language research 2d ago The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal Interfaces arXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirely through text, while the motion channel beside it, the one peripheral vision… 28 Page 2 of 10 · 500 articles ← Newer Older →