Hugging Face Daily Papers · · 5 min read

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

homepage: <a href=\"https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v\" rel=\"nofollow\">https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v</a></p>\n<p>code: <a href=\"https://github.com/KhanCold/merchantbench\" rel=\"nofollow\">https://github.com/KhanCold/merchantbench</a></p>\n<p>paper: <a href=\"https://arxiv.org/abs/2607.28956\" rel=\"nofollow\">https://arxiv.org/abs/2607.28956</a></p>\n","updatedAt":"2026-08-04T13:26:43.085Z","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.566243052482605},"editors":["KhanCold"],"editorAvatarUrls":["/avatars/2925913661e65e9f5731e2eee0489d5d.svg"],"reactions":[],"isReport":false}},{"id":"6a729b97e6aa4dc6927cdc8c","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-08-05T02:10:31.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"**What happens when an LLM agent is asked to run an online store for an entire year?**\n\nWe introduce **MerchantBench**, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in **98,843 real product records**, the environment gives agents access to **26 merchant tools** for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.\n\nAcross **48 year-long runs** covering **8 LLMs and 2 agent frameworks**, the best-performing configuration achieved only **27.3% of the mean final net assets attained by human participants**. Our analysis shows that **long-running does not necessarily mean long-horizon**: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.","html":"<p><strong>What happens when an LLM agent is asked to run an online store for an entire year?</strong></p>\n<p>We introduce <strong>MerchantBench</strong>, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in <strong>98,843 real product records</strong>, the environment gives agents access to <strong>26 merchant tools</strong> for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.</p>\n<p>Across <strong>48 year-long runs</strong> covering <strong>8 LLMs and 2 agent frameworks</strong>, the best-performing configuration achieved only <strong>27.3% of the mean final net assets attained by human participants</strong>. Our analysis shows that <strong>long-running does not necessarily mean long-horizon</strong>: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.</p>\n","updatedAt":"2026-08-05T02:10:31.608Z","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8864301443099976},"editors":["KhanCold"],"editorAvatarUrls":["/avatars/2925913661e65e9f5731e2eee0489d5d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28956","authors":[{"_id":"6a71e8351a375f948521bf4b","user":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user","name":"KhanCold"},"name":"Qiming Shi","status":"claimed_verified","statusLastChangedAt":"2026-08-04T16:45:04.644Z","hidden":false},{"_id":"6a71e8351a375f948521bf4c","name":"Yulong Tao","hidden":false},{"_id":"6a71e8351a375f948521bf4d","name":"Linbo Jin","hidden":false},{"_id":"6a71e8351a375f948521bf4e","name":"Zhaolu Kang","hidden":false},{"_id":"6a71e8351a375f948521bf4f","name":"Yibo Dou","hidden":false},{"_id":"6a71e8351a375f948521bf50","name":"Jiawen Zhu","hidden":false},{"_id":"6a71e8351a375f948521bf51","name":"Tianjun Pan","hidden":false},{"_id":"6a71e8351a375f948521bf52","name":"Shaokang Fu","hidden":false},{"_id":"6a71e8351a375f948521bf53","name":"Chengyu Wang","hidden":false},{"_id":"6a71e8351a375f948521bf54","name":"Siyue Li","hidden":false},{"_id":"6a71e8351a375f948521bf55","name":"Yaping Cheng","hidden":false},{"_id":"6a71e8351a375f948521bf56","name":"Di Weng","hidden":false},{"_id":"6a71e8351a375f948521bf57","name":"Chengfu Huo","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/0HQbpnB9wBO7Woe1ysxNP.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/yxAmDUotaRRggi61e8I_u.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/X4y_LdPMnfHrNLMpPiixO.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/mDzCQJG_MbJ7viVkOW-w4.png"],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations","submittedOnDailyBy":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user","name":"KhanCold"},"summary":"Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\\% of the mean final net assets achieved by human participants.","upvotes":43,"discussionId":"6a71e8351a375f948521bf58","projectPage":"https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v","githubRepo":"https://github.com/KhanCold/merchantbench","githubRepoAddedBy":"user","githubStars":16,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user"},{"_id":"64834b399b352597e41816ac","avatarUrl":"/avatars/63d9d123bffa90f43186a0bdc4455cbd.svg","isPro":false,"fullname":"Shaobai Jiang","user":"shaobaij","type":"user"},{"_id":"6a729de5d848a5998e470e9a","avatarUrl":"/avatars/03f94812d710b601a58d23fc5fbef2a4.svg","isPro":false,"fullname":"mamo","user":"userAgent-123","type":"user"},{"_id":"6a729ed1f6a9f2b9f8716584","avatarUrl":"/avatars/aca2a49ca6d0e9d00289d69eedf0396f.svg","isPro":false,"fullname":"alex","user":"alexburg","type":"user"},{"_id":"6a72a05e718367ef671ff455","avatarUrl":"/avatars/5c8f6ce82c0d597fc7f7c84e24d935b0.svg","isPro":false,"fullname":"john","user":"johnhack","type":"user"},{"_id":"6a72a1ad2a0d8080bce77e01","avatarUrl":"/avatars/54a4f1b6d4079956ce968761751bedfa.svg","isPro":false,"fullname":"bay","user":"v-bright","type":"user"},{"_id":"68a2a128b92f887b0c601d23","avatarUrl":"/avatars/f6c30645f4390e8ad476a80e68f12f1f.svg","isPro":false,"fullname":"Jiawen Zhu","user":"andone07","type":"user"},{"_id":"6a72a81bb75f91c148606ce6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/bUGL_pQCS2P9UdPzppUAD.jpeg","isPro":false,"fullname":"yibo dou","user":"dyblove","type":"user"},{"_id":"693f65a14030d063892cfe87","avatarUrl":"/avatars/246b0c9d95362926887cbfdb8802c20a.svg","isPro":false,"fullname":"FelixChristian","user":"FelixChristian","type":"user"},{"_id":"689ddc625a0892937e632ce5","avatarUrl":"/avatars/14624c6509def0444142d9d76b07b4da.svg","isPro":false,"fullname":"Shaokang Fu","user":"Pomore","type":"user"},{"_id":"6698d159b2ebada9f4a86f2d","avatarUrl":"/avatars/870594d58a8f0f6fb520c5227754d2de.svg","isPro":false,"fullname":"tianjun pan","user":"blazzer","type":"user"},{"_id":"6a6a8229a5b9c4c08badf665","avatarUrl":"/avatars/d8d4deed04213f6dfbe4b504c2e0c1ec.svg","isPro":false,"fullname":"Timothy Garcia","user":"timothy-garcia","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28956.md","query":{}}">
Papers
arxiv:2607.28956

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Published on Jul 31
· Submitted by
Qiming Shi
on Aug 5
#3 Paper of the day
Authors:

Abstract

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

Community

Paper author Paper submitter about 6 hours ago

What happens when an LLM agent is asked to run an online store for an entire year?

We introduce MerchantBench, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in 98,843 real product records, the environment gives agents access to 26 merchant tools for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.

Across 48 year-long runs covering 8 LLMs and 2 agent frameworks, the best-performing configuration achieved only 27.3% of the mean final net assets attained by human participants. Our analysis shows that long-running does not necessarily mean long-horizon: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28956
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.28956 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.28956 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28956 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers