homepage: <a href=\"https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v\" rel=\"nofollow\">https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v</a></p>\n<p>code: <a href=\"https://github.com/KhanCold/merchantbench\" rel=\"nofollow\">https://github.com/KhanCold/merchantbench</a></p>\n<p>paper: <a href=\"https://arxiv.org/abs/2607.28956\" rel=\"nofollow\">https://arxiv.org/abs/2607.28956</a></p>\n","updatedAt":"2026-08-04T13:26:43.085Z","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.566243052482605},"editors":["KhanCold"],"editorAvatarUrls":["/avatars/2925913661e65e9f5731e2eee0489d5d.svg"],"reactions":[],"isReport":false}},{"id":"6a729b97e6aa4dc6927cdc8c","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-08-05T02:10:31.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"**What happens when an LLM agent is asked to run an online store for an entire year?**\n\nWe introduce **MerchantBench**, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in **98,843 real product records**, the environment gives agents access to **26 merchant tools** for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.\n\nAcross **48 year-long runs** covering **8 LLMs and 2 agent frameworks**, the best-performing configuration achieved only **27.3% of the mean final net assets attained by human participants**. Our analysis shows that **long-running does not necessarily mean long-horizon**: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.","html":"<p><strong>What happens when an LLM agent is asked to run an online store for an entire year?</strong></p>\n<p>We introduce <strong>MerchantBench</strong>, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in <strong>98,843 real product records</strong>, the environment gives agents access to <strong>26 merchant tools</strong> for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.</p>\n<p>Across <strong>48 year-long runs</strong> covering <strong>8 LLMs and 2 agent frameworks</strong>, the best-performing configuration achieved only <strong>27.3% of the mean final net assets attained by human participants</strong>. Our analysis shows that <strong>long-running does not necessarily mean long-horizon</strong>: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.</p>\n","updatedAt":"2026-08-05T02:10:31.608Z","author":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","fullname":"Qiming Shi","name":"KhanCold","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8864301443099976},"editors":["KhanCold"],"editorAvatarUrls":["/avatars/2925913661e65e9f5731e2eee0489d5d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28956","authors":[{"_id":"6a71e8351a375f948521bf4b","user":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user","name":"KhanCold"},"name":"Qiming Shi","status":"claimed_verified","statusLastChangedAt":"2026-08-04T16:45:04.644Z","hidden":false},{"_id":"6a71e8351a375f948521bf4c","name":"Yulong Tao","hidden":false},{"_id":"6a71e8351a375f948521bf4d","name":"Linbo Jin","hidden":false},{"_id":"6a71e8351a375f948521bf4e","name":"Zhaolu Kang","hidden":false},{"_id":"6a71e8351a375f948521bf4f","name":"Yibo Dou","hidden":false},{"_id":"6a71e8351a375f948521bf50","name":"Jiawen Zhu","hidden":false},{"_id":"6a71e8351a375f948521bf51","name":"Tianjun Pan","hidden":false},{"_id":"6a71e8351a375f948521bf52","name":"Shaokang Fu","hidden":false},{"_id":"6a71e8351a375f948521bf53","name":"Chengyu Wang","hidden":false},{"_id":"6a71e8351a375f948521bf54","name":"Siyue Li","hidden":false},{"_id":"6a71e8351a375f948521bf55","name":"Yaping Cheng","hidden":false},{"_id":"6a71e8351a375f948521bf56","name":"Di Weng","hidden":false},{"_id":"6a71e8351a375f948521bf57","name":"Chengfu Huo","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/0HQbpnB9wBO7Woe1ysxNP.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/yxAmDUotaRRggi61e8I_u.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/X4y_LdPMnfHrNLMpPiixO.png","https://cdn-uploads.huggingface.co/production/uploads/693f82bc28d6c949c6ef4d1b/mDzCQJG_MbJ7viVkOW-w4.png"],"publishedAt":"2026-07-31T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations","submittedOnDailyBy":{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user","name":"KhanCold"},"summary":"Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\\% of the mean final net assets achieved by human participants.","upvotes":43,"discussionId":"6a71e8351a375f948521bf58","projectPage":"https://air.1688.com/kapp/next1688/merchantbench/?spm=defwork.home.0.0.7f8c530dSdRT9v","githubRepo":"https://github.com/KhanCold/merchantbench","githubRepoAddedBy":"user","githubStars":16,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"693f82bc28d6c949c6ef4d1b","avatarUrl":"/avatars/2925913661e65e9f5731e2eee0489d5d.svg","isPro":false,"fullname":"Qiming Shi","user":"KhanCold","type":"user"},{"_id":"64834b399b352597e41816ac","avatarUrl":"/avatars/63d9d123bffa90f43186a0bdc4455cbd.svg","isPro":false,"fullname":"Shaobai Jiang","user":"shaobaij","type":"user"},{"_id":"6a729de5d848a5998e470e9a","avatarUrl":"/avatars/03f94812d710b601a58d23fc5fbef2a4.svg","isPro":false,"fullname":"mamo","user":"userAgent-123","type":"user"},{"_id":"6a729ed1f6a9f2b9f8716584","avatarUrl":"/avatars/aca2a49ca6d0e9d00289d69eedf0396f.svg","isPro":false,"fullname":"alex","user":"alexburg","type":"user"},{"_id":"6a72a05e718367ef671ff455","avatarUrl":"/avatars/5c8f6ce82c0d597fc7f7c84e24d935b0.svg","isPro":false,"fullname":"john","user":"johnhack","type":"user"},{"_id":"6a72a1ad2a0d8080bce77e01","avatarUrl":"/avatars/54a4f1b6d4079956ce968761751bedfa.svg","isPro":false,"fullname":"bay","user":"v-bright","type":"user"},{"_id":"68a2a128b92f887b0c601d23","avatarUrl":"/avatars/f6c30645f4390e8ad476a80e68f12f1f.svg","isPro":false,"fullname":"Jiawen Zhu","user":"andone07","type":"user"},{"_id":"6a72a81bb75f91c148606ce6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/bUGL_pQCS2P9UdPzppUAD.jpeg","isPro":false,"fullname":"yibo dou","user":"dyblove","type":"user"},{"_id":"693f65a14030d063892cfe87","avatarUrl":"/avatars/246b0c9d95362926887cbfdb8802c20a.svg","isPro":false,"fullname":"FelixChristian","user":"FelixChristian","type":"user"},{"_id":"689ddc625a0892937e632ce5","avatarUrl":"/avatars/14624c6509def0444142d9d76b07b4da.svg","isPro":false,"fullname":"Shaokang Fu","user":"Pomore","type":"user"},{"_id":"6698d159b2ebada9f4a86f2d","avatarUrl":"/avatars/870594d58a8f0f6fb520c5227754d2de.svg","isPro":false,"fullname":"tianjun pan","user":"blazzer","type":"user"},{"_id":"6a6a8229a5b9c4c08badf665","avatarUrl":"/avatars/d8d4deed04213f6dfbe4b504c2e0c1ec.svg","isPro":false,"fullname":"Timothy Garcia","user":"timothy-garcia","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28956.md","query":{}}">
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Abstract
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.
Community
What happens when an LLM agent is asked to run an online store for an entire year?
We introduce MerchantBench, a 365-day, order-level simulation for evaluating the long-term coherence of LLM agents in e-commerce operations. Grounded in 98,843 real product records, the environment gives agents access to 26 merchant tools for product sourcing, pricing, order tracking, and cash-flow management, while exposing them to supplier disruptions and delayed outcomes such as refunds, negative reviews, and penalties.
Across 48 year-long runs covering 8 LLMs and 2 agent frameworks, the best-performing configuration achieved only 27.3% of the mean final net assets attained by human participants. Our analysis shows that long-running does not necessarily mean long-horizon: agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.28956 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.28956 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.28956 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.