Parametric tool retrieval lets LLMs retrieve APIs by generating virtual tokens — but existing training objectives (ToolGen) catastrophically destroy the model's tool knowledge in the process, and their constrained beam-search decoding is too slow for production. TRACE introduces a two-stage curriculum that resolves this dissociation: Stage 1 seeds tool knowledge via multi-format memorization (LoRA), and Stage 2 trains the model to emit a reasoning trace before committing to tool tokens, grounded in domain-expert-curated business rules. On a combined enterprise catalog of 8,283 tools (HR + Finance), TRACE achieves <del>86% recall under single-beam greedy decoding at production latency (</del>1.9s vs ~19s for beam search, a 200× throughput gain), while improving MCQ and QA probing accuracy by +7.6 pp and +4.5 pp over Stage 1 — the first parametric retrieval system that is both knowledge-preserving and deployable at scale.</p>\n","updatedAt":"2026-07-28T13:22:59.151Z","author":{"_id":"637859f98f288aba3d01f588","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637859f98f288aba3d01f588/eP8YNMOtxTvxH-rYEYn4f.png","fullname":"Ashutosh Hathidara","name":"ashutosh1919","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f7a0ad07b08a2b3c3ac94e/lIw9a3y-z_5RGc7i64YPu.png","fullname":"SAP","name":"SAP","type":"org","isHf":false,"details":"Tabular AI, Table Representation Learning, Agents and Knowledge Graphs","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9026891589164734},"editors":["ashutosh1919"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/637859f98f288aba3d01f588/eP8YNMOtxTvxH-rYEYn4f.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.22639","authors":[{"_id":"6a68ac8e81e96813d6f78a99","name":"Sai Shruthi Sistla","hidden":false},{"_id":"6a68ac8e81e96813d6f78a9a","user":{"_id":"637859f98f288aba3d01f588","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637859f98f288aba3d01f588/eP8YNMOtxTvxH-rYEYn4f.png","isPro":false,"fullname":"Ashutosh Hathidara","user":"ashutosh1919","type":"user","name":"ashutosh1919"},"name":"Ashutosh Hathidara","status":"claimed_verified","statusLastChangedAt":"2026-07-28T16:45:04.682Z","hidden":false},{"_id":"6a68ac8e81e96813d6f78a9b","name":"Christopher Toukmaji","hidden":false},{"_id":"6a68ac8e81e96813d6f78a9c","name":"Mayank Shrivastava","hidden":false},{"_id":"6a68ac8e81e96813d6f78a9d","name":"Karthikeyan Asokkumar","hidden":false}],"publishedAt":"2026-06-22T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs","submittedOnDailyBy":{"_id":"637859f98f288aba3d01f588","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637859f98f288aba3d01f588/eP8YNMOtxTvxH-rYEYn4f.png","isPro":false,"fullname":"Ashutosh Hathidara","user":"ashutosh1919","type":"user","name":"ashutosh1919"},"summary":"Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources -- RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts -- both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B -- compared to embedding baseline performance of ~27% & ~52% -- both with single-beam greedy decoding, making it directly deployable at production latency.","upvotes":2,"discussionId":"6a68ac8e81e96813d6f78a9e","organization":{"_id":"6152dcdfecf3ca6ab820e328","name":"SAP","fullname":"SAP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f7a0ad07b08a2b3c3ac94e/lIw9a3y-z_5RGc7i64YPu.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"637859f98f288aba3d01f588","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637859f98f288aba3d01f588/eP8YNMOtxTvxH-rYEYn4f.png","isPro":false,"fullname":"Ashutosh Hathidara","user":"ashutosh1919","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6152dcdfecf3ca6ab820e328","name":"SAP","fullname":"SAP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f7a0ad07b08a2b3c3ac94e/lIw9a3y-z_5RGc7i64YPu.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.22639.md","query":{}}">
TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
Abstract
Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources -- RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts -- both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B -- compared to embedding baseline performance of ~27% & ~52% -- both with single-beam greedy decoding, making it directly deployable at production latency.
Community
Parametric tool retrieval lets LLMs retrieve APIs by generating virtual tokens — but existing training objectives (ToolGen) catastrophically destroy the model's tool knowledge in the process, and their constrained beam-search decoding is too slow for production. TRACE introduces a two-stage curriculum that resolves this dissociation: Stage 1 seeds tool knowledge via multi-format memorization (LoRA), and Stage 2 trains the model to emit a reasoning trace before committing to tool tokens, grounded in domain-expert-curated business rules. On a combined enterprise catalog of 8,283 tools (HR + Finance), TRACE achieves 86% recall under single-beam greedy decoding at production latency (1.9s vs ~19s for beam search, a 200× throughput gain), while improving MCQ and QA probing accuracy by +7.6 pp and +4.5 pp over Stage 1 — the first parametric retrieval system that is both knowledge-preserving and deployable at scale.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.22639 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.22639 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.22639 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.