This paper records the pre-alpha systems behind DXAP: 3,505 user-funded Base vaults and a 500-599-agent Hyperliquid fleet with 231,638 finalized turns. P&L varied widely across agents and over time. The useful result was how the system around the model determined whether an agent carried out what its user wanted. We used those findings to rebuild the current DXAP Alpha around persistent strategy automation, new data and tools, and learning with users. Aggregate figure data and discussion notes are linked in the GitHub repository.</p>\n","updatedAt":"2026-09-09T03:21:34.983Z","author":{"_id":"66bfa2d3eef032a019f21421","avatarUrl":"/avatars/0128f1f0bfa137102c3028f5da5d999c.svg","fullname":"Poof","name":"poofuse","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9195404648780823},"editors":["poofuse"],"editorAvatarUrls":["/avatars/0128f1f0bfa137102c3028f5da5d999c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.05663","authors":[{"_id":"6aa0c9cfd0174964227bec79","name":"T. J. Barton","hidden":false},{"_id":"6aa0c9cfd0174964227bec7a","name":"Chris Constantakis","hidden":false},{"_id":"6aa0c9cfd0174964227bec7b","name":"Patti Hauseman","hidden":false},{"_id":"6aa0c9cfd0174964227bec7c","name":"Annie Mous","hidden":false},{"_id":"6aa0c9cfd0174964227bec7d","name":"Alaska Hoffman","hidden":false},{"_id":"6aa0c9cfd0174964227bec7e","name":"Brian Bergeron","hidden":false},{"_id":"6aa0c9cfd0174964227bec7f","name":"Hunter Goodreau","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets","submittedOnDailyBy":{"_id":"66bfa2d3eef032a019f21421","avatarUrl":"/avatars/0128f1f0bfa137102c3028f5da5d999c.svg","isPro":true,"fullname":"Poof","user":"poofuse","type":"user","name":"poofuse"},"summary":"We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.","upvotes":16,"discussionId":"6aa0c9cfd0174964227bec80","projectPage":"https://www.dxrg.ai/blogs/continuous-record-paper","githubRepo":"https://github.com/ProjectDXAI/continuous-record-llm-trading-agents","githubRepoAddedBy":"user","ai_summary":"Autonomous language-model trading agents across production systems show behavior driven by interface design rather than strategy, exhibit volatility-blind sizing, fail to capture favorable price excursions, and display no directional edge, with frontier model decision quality statistically indistinguishable across families.","ai_keywords":["language-model agents","autonomous agents","multi-tool turns","model invocations","frontier models","model families","decision quality","choice stability"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"69f3893e036ea39f06f6839d","name":"DXRG","fullname":"DXRG AI Inc","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66bfa2d3eef032a019f21421/KWq9O-UgjsruN4B4cMGvg.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"66bfa2d3eef032a019f21421","avatarUrl":"/avatars/0128f1f0bfa137102c3028f5da5d999c.svg","isPro":true,"fullname":"Poof","user":"poofuse","type":"user"},{"_id":"6a6c7a3d3574d63d5305cd74","avatarUrl":"/avatars/78a0b7bd8c1b9996c19966b58df44488.svg","isPro":false,"fullname":"Richard Wilson","user":"Cobalt-Richard4","type":"user"},{"_id":"6a6c89729155f13f6e9253d6","avatarUrl":"/avatars/f3eae096bba4b812eaa06291ea108e11.svg","isPro":false,"fullname":"Elizabeth Taylor","user":"elizabethTaylor","type":"user"},{"_id":"6a6de7f10f8a360db52be308","avatarUrl":"/avatars/ff0b0c8b58aa8e165cf223cf59acc6dc.svg","isPro":false,"fullname":"Jennifer Miller","user":"Sable-Jennifer","type":"user"},{"_id":"6a701f5cbc4016943249bb3d","avatarUrl":"/avatars/6d0c43704f79ba786ce1f6010c3bc370.svg","isPro":false,"fullname":"Steven Brown","user":"Steven-Brown","type":"user"},{"_id":"6a9aedf53500c8f814fe7699","avatarUrl":"/avatars/1f241cf3d15b45d867a2321655b5c6b9.svg","isPro":false,"fullname":"이영희","user":"cedarDawnJ","type":"user"},{"_id":"6a9b517c1cf5e5287880c5bd","avatarUrl":"/avatars/5756f524438ab97d016579218ad3c215.svg","isPro":false,"fullname":"吉田 太一","user":"xkimura271","type":"user"},{"_id":"6aa0f3fa7217d82009109680","avatarUrl":"/avatars/bed1303deb6326c9df82f3863ce83e8d.svg","isPro":false,"fullname":"David Johnson","user":"daltontony","type":"user"},{"_id":"6aa0fc6c899c3077e3fcf09f","avatarUrl":"/avatars/c7ac56dab929483fe57667bdc5048438.svg","isPro":false,"fullname":"王萍","user":"Nimbus-Vault","type":"user"},{"_id":"6a6a9417a569279e390e1343","avatarUrl":"/avatars/4866974d3fab4ba081b094cd0819de75.svg","isPro":false,"fullname":"Mark Davis","user":"silverArcL","type":"user"},{"_id":"6a6c8887e7a7d1e67347457c","avatarUrl":"/avatars/822e32e5b1c5242266d8d4d0b4eeb00d.svg","isPro":false,"fullname":"Mary Perez","user":"ZenithTrail","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69f3893e036ea39f06f6839d","name":"DXRG","fullname":"DXRG AI Inc","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66bfa2d3eef032a019f21421/KWq9O-UgjsruN4B4cMGvg.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.05663.md","query":{}}">
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
Published on Sep 4
· Submitted by Poof on Sep 9 Abstract
Autonomous language-model trading agents across production systems show behavior driven by interface design rather than strategy, exhibit volatility-blind sizing, fail to capture favorable price excursions, and display no directional edge, with frontier model decision quality statistically indistinguishable across families.
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.
Community
This paper records the pre-alpha systems behind DXAP: 3,505 user-funded Base vaults and a 500-599-agent Hyperliquid fleet with 231,638 finalized turns. P&L varied widely across agents and over time. The useful result was how the system around the model determined whether an agent carried out what its user wanted. We used those findings to rebuild the current DXAP Alpha around persistent strategy automation, new data and tools, and learning with users. Aggregate figure data and discussion notes are linked in the GitHub repository.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.05663 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.05663 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.05663 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.