News / #agents Tag Agents + tool use 500 articles archived under #agents · RSS Sign in to follow r/LocalLLaMA community 6d ago Ngram and world knowledge - why are we just building a coding model? This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that… 9 r/LocalLLaMA community 6d ago XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B https://preview.redd.it/ybwqqbrst0rh1.png?width=811&format=png&auto=webp&s=1e40c5b53e304caf2a10efa2baa6f998b9a9a0bb MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data. Made by Xiaomi… 38 arXiv — Machine Learning research 6d ago Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents arXiv:2609.22120v1 Announce Type: new Abstract: Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution.… 23 arXiv — Machine Learning research 6d ago StepKV: Step-Aware KV Cache Compression for LLM Agents arXiv:2609.22158v1 Announce Type: new Abstract: Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this… 30 arXiv — Machine Learning research 6d ago Beyond Task Completion: Training Capable and Safe Computer-Use Agents arXiv:2609.22178v1 Announce Type: new Abstract: Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must… 32 arXiv — Machine Learning research 6d ago A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents arXiv:2609.22194v1 Announce Type: new Abstract: Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT)… 18 arXiv — Machine Learning research 6d ago Toollery: Scaling LLM Agents to Thousands of Skills and Tools arXiv:2609.22218v1 Announce Type: new Abstract: As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer… 30 arXiv — Machine Learning research 6d ago Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark arXiv:2609.22222v1 Announce Type: new Abstract: Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This study asks whether authoritative metadata and execution feedback improve the… 24 arXiv — Machine Learning research 6d ago CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents arXiv:2609.22247v1 Announce Type: new Abstract: Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of… 18 arXiv — Machine Learning research 6d ago Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting arXiv:2609.22251v1 Announce Type: new Abstract: Forecasting karst aquifer dynamics is difficult because recharge responses are nonlinear, event-driven, and governed by strongly heterogeneous flow paths. This study develops and evaluates a deployment-aware framework for… 7 arXiv — Machine Learning research 6d ago RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks arXiv:2609.22258v1 Announce Type: new Abstract: Large language model-driven remote sensing (RS) agents offer a promising approach to automating geospatial analysis. However, lightweight RS agents based on compact language models struggle with multi-step interactive tasks due to… 29 arXiv — Machine Learning research 6d ago Personalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy Gradients arXiv:2609.22833v1 Announce Type: new Abstract: We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for… 13 arXiv — Machine Learning research 6d ago Testing the Construct Validity of a Functional Valence Axis in LLM Agents arXiv:2609.22850v1 Announce Type: new Abstract: Contrastive activation directions are often interpreted from what they decode or how strongly they steer behavior. But what evidence is sufficient to identify the construct represented by such a direction, rather than a correlated… 24 arXiv — Machine Learning research 6d ago Towards Full Pipeline FP8 Reinforcement Learning for LLMs arXiv:2609.22870v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8… 19 arXiv — NLP / Computation & Language research 6d ago Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents arXiv:2609.22090v1 Announce Type: new Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM… 17 arXiv — NLP / Computation & Language research 6d ago DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation arXiv:2609.22104v1 Announce Type: new Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottleneck from idea generation to idea evaluation. Existing evaluators mainly rely on… 11 arXiv — NLP / Computation & Language research 6d ago Beyond the Text: Verifying That Agent-Written Papers Are Backed by Their Artifacts arXiv:2609.22111v1 Announce Type: new Abstract: Large language model agents are increasingly capable of conducting research autonomously, producing research documents alongside the code and experiments that ostensibly support them. Yet whether the reported findings are… 6 arXiv — NLP / Computation & Language research 6d ago An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents arXiv:2609.22114v1 Announce Type: new Abstract: Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression preserves task-solving quality. These two facts do not imply the third one… 16 arXiv — NLP / Computation & Language research 6d ago PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations arXiv:2609.22200v1 Announce Type: new Abstract: LLM assistants and agentic systems log long multi-turn conversations. AI providers often scan these conversations for Personally Identifiable Information (PII) and mask the PII before storing or processing conversation data. Yet… 4 arXiv — NLP / Computation & Language research 6d ago Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research arXiv:2609.22209v1 Announce Type: new Abstract: Empirical legal research often relies on turning research questions into structured data extracted from large collections of rulings and judgments. Designing the extraction schema and then extracting the data remain a manual,… 26 arXiv — NLP / Computation & Language research 6d ago EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy arXiv:2609.22223v1 Announce Type: new Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and… 34 arXiv — NLP / Computation & Language research 6d ago BizSage: A Self-Evolving Multi-Agent Framework for Business Research with Efficient Knowledge Retrieval arXiv:2609.22235v1 Announce Type: new Abstract: While multi-agent systems based on large language models (LLMs) have shown promise in automating the progressive workflow of academic research, extending them to economics and business research, where specialized domain knowledge… 18 arXiv — NLP / Computation & Language research 6d ago Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations arXiv:2609.22255v1 Announce Type: new Abstract: Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a… 30 arXiv — NLP / Computation & Language research 6d ago Which Part of the Context Layer Does the Work? Separating Semantic Content from Retrieval Scaffolding in Text-to-SQL Agents arXiv:2609.22259v1 Announce Type: new Abstract: Context layers, curated documentation that an analytics agent fetches at query time, produce large accuracy gains on text-to-SQL benchmarks. A with/without comparison cannot say which part of the layer does the work: the semantic… 6 arXiv — NLP / Computation & Language research 6d ago An Iterative LangGraph Agent for Text-to-SQL: Natural Language Access to the Chicago Crime Database arXiv:2609.22917v1 Announce Type: new Abstract: Non-technical stakeholders frequently cannot write the SQL needed to extract insights from operational databases. We built and evaluated a Text-to-SQL agent that closes this gap end to end: a six-node LangGraph StateGraph checks… 17 arXiv — NLP / Computation & Language research 6d ago Automatic multimodal UX improvement recommendations from LLM agent user simulations arXiv:2609.22971v1 Announce Type: new Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However,… 22 arXiv — NLP / Computation & Language research 6d ago Bridging Static and Agentic RAG for Taiwanese Historical Question Answering arXiv:2609.23056v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed… 7 The Information — AI news-outlet 6d ago Meta’s Muse Tries (and Fails) to Disrupt Amazon Everyone loves a brawl between two big tech firms—even other tech CEOs. Amazon’s block on Meta Platforms’ new Muse personal agent accessing its shopping site, which came to light Sunday night, prompted Palo Alto Networks CEO Nikesh Arora to post on X on Monday that the episode… 15 Vercel — AI dev-tools 6d ago Claude Opus 5.5 now available on AI Gateway Claude Opus 5.5 from Anthropic is now available on AI Gateway . It is a step-change improvement over Opus 5 , with its biggest gains in agentic coding, long-running agent tasks, and knowledge work. Opus 5.5 is also a better collaborator over long runs. It reports back in plain… 5 r/LocalLLaMA community 6d ago 50+ Hours and 100M+ Tokens Later, Open Source Autonomous Agent is GETTING CLOSER at Solving an Open Math problem This experiment is live, you can inspect all the internal reasoning, memories, attempts here: https://artificium-covering-experiment.gr.bio/ The problem that the agent is trying to solve is a covering design problem: https://en.wikipedia.org/wiki/Covering_design Known as… 19 NVIDIA Developer Blog official-blog 6d ago How to Evaluate AI Agents From Tool Calls to Task Completion When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and... 33 TechCrunch — AI news-outlet 6d ago Meta’s Muse is outpacing ChatGPT’s early mobile launch Meta’s new AI agent Muse has racked up more downloads and daily active users in the U.S. and Canada than ChatGPT did over the same period after its mobile debut, according to new estimates from Appfigures. 33 TechCrunch — AI news-outlet 6d ago Meta’s AI agent has been blocked from using Amazon.com Muse isn't welcome as a shopper anymore. 12 Vercel — AI dev-tools 6d ago Vercel Connect now supports Microsoft Teams Vercel Connect now includes a managed connector for Microsoft Teams . Creating one gives your organization a Teams bot that your apps and agents run. People can @mention it in channels or message it directly, and your code receives the message and replies as the bot. As a Vercel… 33 r/LocalLLaMA community 6d ago A better coder for the small-GPU/small-RAM crowd! I’ve been working on making small models more capable at agentic coding and work, because most people in the world don’t have the sort of hardware needed to run 3.8-27B, or even 35B-A3B or 9B dense, and I want to extend local agentic coding capability to less privileged users.… 21 The Information — AI news-outlet 6d ago OpenAI Develops Features to Counter Grok Bot, Mulls Response to Meta’s Muse OpenAI pioneered AI that can take over web browsers and other applications on people’s behalf, and its ChatGPT app brought generative AI to consumers. Now that rivals like SpaceX and Meta are pushing the boundaries by launching personal AI agents that handle multistep tasks,… 33 r/LocalLLaMA community 6d ago M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents - MacStories   submitted by   /u/themixtergames [link]   [comments] 17 The Information — AI news-outlet 6d ago Amazon Blocks Meta’s Muse Agent Amazon blocked Meta’s Muse personalized AI agent from accessing its shopping site, Geekwire reported, saying that outside applications “should operate openly and respect service provider decisions about whether or not to participate.” The block appeared to even prevent Muse from… 14 arXiv — NLP / Computation & Language research 7d ago Scaling Discovery through Test-Time Communication arXiv:2609.21032v1 Announce Type: cross Abstract: Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time… 25 arXiv — Machine Learning research 7d ago SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs? arXiv:2609.21190v1 Announce Type: new Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and… 17 arXiv — Machine Learning research 7d ago OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems arXiv:2609.21527v1 Announce Type: new Abstract: Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score… 26 arXiv — Machine Learning research 7d ago MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems arXiv:2609.21533v1 Announce Type: new Abstract: LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs… 15 arXiv — Machine Learning research 7d ago GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills arXiv:2609.21749v1 Announce Type: new Abstract: Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing… 19 arXiv — Machine Learning research 7d ago Benchmarking World Models for Continual Learning on Compositional Tasks arXiv:2609.22055v1 Announce Type: new Abstract: A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge… 10 arXiv — NLP / Computation & Language research 7d ago Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces arXiv:2609.20844v1 Announce Type: new Abstract: Deepresearch (DR) agents interact with real-world web environments through multi-turn search and visit, causing their contexts to grow rapidly over time. We observe that, even after DR Agentic Reinforcement Learning (DR-RL), 61.6%… 11 arXiv — NLP / Computation & Language research 7d ago CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop arXiv:2609.21154v1 Announce Type: new Abstract: Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer… 34 arXiv — NLP / Computation & Language research 7d ago When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success arXiv:2609.21187v1 Announce Type: new Abstract: Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this… 31 arXiv — NLP / Computation & Language research 7d ago From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers arXiv:2609.21349v1 Announce Type: new Abstract: Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging. Existing in-context learning-based methods fail to capture how individuals react under… 28 arXiv — NLP / Computation & Language research 7d ago ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL arXiv:2609.21378v1 Announce Type: new Abstract: Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard… 34 arXiv — NLP / Computation & Language research 7d ago Do Personality-Tuned LLMs Make Better Social Agents? arXiv:2609.21857v1 Announce Type: new Abstract: LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent… 18 Page 4 of 10 · 500 articles ← Newer Older →