Hugging Face Daily Papers · · 6 min read

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We consolidated traffic from 200+ internal apps onto one self-hosted model, closing the gaps production error analysis showed: instruction following, function calling, and our internal task mix. Instead of one joint objective we train a GRPO expert per axis and merge with two-stage SLERP — each axis hacks its reward differently (semantic collapse, over-calling, verbosity hacking). Non-reasoning mode beats a ~7× larger baseline on our Arena (69.6 vs 65.8) and now serves 50% of platform traffic, 116M requests/month.</p>\n","updatedAt":"2026-09-02T10:14:26.908Z","author":{"_id":"62609d224e6e4b84475eb8d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62609d224e6e4b84475eb8d9/4mHSdgtmIc88oqqzLLhJE.jpeg","fullname":"Alex Medvedev","name":"kenkaneki","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674ea07d320a043daeb2d98b/IwSCMolFY4Otk7sFXzWhi.jpeg","fullname":"T-Tech","name":"t-tech","type":"org","isHf":false,"details":"Scientific research; Natural language processing: speech analytics, search engines, dialogue systems; A family of LLMs; Speech technologies; Fraud prevention technologies; Computer vision; Recommender systems; Time series analysis","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8804317712783813},"editors":["kenkaneki"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/62609d224e6e4b84475eb8d9/4mHSdgtmIc88oqqzLLhJE.jpeg"],"reactions":[],"isReport":false}},{"id":"6a98126c4dbd314b16e51fdc","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-09-02T12:11:24.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Consolidating the whole corporate request mix onto one self-hosted model only pays off if the retraining loop closes the gap faster than the apps drift. In my experience the request mix shifts weekly — new tools, new prompts, new failure modes — and if your post-training cycle is a monthly batch job, you're always chasing last quarter's traffic. The number I'd want isn't coverage at snapshot time, it's the half-life of that coverage.\n\nAnd I'd want to see the judge setup before trusting the quality numbers. \"Calibrated LLM judges\" is where these pipelines usually leak — if the judge was tuned on the same traffic you're optimizing for, you're measuring how well the model mimics the judge, not how well it serves the request. Show me the judge's disagreement rate with human raters on the hard tail, not the aggregate score.","html":"<p>Consolidating the whole corporate request mix onto one self-hosted model only pays off if the retraining loop closes the gap faster than the apps drift. In my experience the request mix shifts weekly — new tools, new prompts, new failure modes — and if your post-training cycle is a monthly batch job, you're always chasing last quarter's traffic. The number I'd want isn't coverage at snapshot time, it's the half-life of that coverage.</p>\n<p>And I'd want to see the judge setup before trusting the quality numbers. \"Calibrated LLM judges\" is where these pipelines usually leak — if the judge was tuned on the same traffic you're optimizing for, you're measuring how well the model mimics the judge, not how well it serves the request. Show me the judge's disagreement rate with human raters on the hard tail, not the aggregate score.</p>\n","updatedAt":"2026-09-02T12:11:24.914Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9563876390457153},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.01572","authors":[{"_id":"6a97f078d4c640d9b759ba46","name":"Olga Tsymboi","hidden":false},{"_id":"6a97f078d4c640d9b759ba47","name":"Dmitrii Stoianov","hidden":false},{"_id":"6a97f078d4c640d9b759ba48","name":"Ramil Latypov","hidden":false},{"_id":"6a97f078d4c640d9b759ba49","user":{"_id":"64fb054ebb362cbf2fe53159","avatarUrl":"/avatars/936c37a77d46d0ea579d2f8a9aea9284.svg","isPro":false,"fullname":"Danil Taranets","user":"taranetsdan","type":"user","name":"taranetsdan"},"name":"Danil Taranets","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:30.434Z","hidden":false},{"_id":"6a97f078d4c640d9b759ba4a","name":"Daniil Dryabin","hidden":false},{"_id":"6a97f078d4c640d9b759ba4b","name":"Mikhail Gashkov","hidden":false},{"_id":"6a97f078d4c640d9b759ba4c","name":"Viktor Zelenkovskiy","hidden":false},{"_id":"6a97f078d4c640d9b759ba4d","name":"Aleksandr Fida","hidden":false},{"_id":"6a97f078d4c640d9b759ba4e","name":"Gleb Alektorov","hidden":false},{"_id":"6a97f078d4c640d9b759ba4f","user":{"_id":"6432343103d81fa4d26dac50","avatarUrl":"/avatars/5c4690b517360636ab7cd5a5bdbe6200.svg","isPro":false,"fullname":"Nikita Gulyakov","user":"elvispresniy","type":"user","name":"elvispresniy"},"name":"Nikita Gulyakov","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:28.425Z","hidden":false},{"_id":"6a97f078d4c640d9b759ba50","user":{"_id":"65f94be15df5183c9a9658fb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f94be15df5183c9a9658fb/s6SR9tQozzjKyddxXgAN4.jpeg","isPro":false,"fullname":"Arthur Babkin","user":"archeee","type":"user","name":"archeee"},"name":"Arthur Babkin","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:26.651Z","hidden":false},{"_id":"6a97f078d4c640d9b759ba51","user":{"_id":"62609d224e6e4b84475eb8d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62609d224e6e4b84475eb8d9/4mHSdgtmIc88oqqzLLhJE.jpeg","isPro":false,"fullname":"Alex Medvedev","user":"kenkaneki","type":"user","name":"kenkaneki"},"name":"Aleksandr Medvedev","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:24.859Z","hidden":false},{"_id":"6a97f078d4c640d9b759ba52","name":"Pavel Gein","hidden":false},{"_id":"6a97f078d4c640d9b759ba53","user":{"_id":"63f26358be95ed4c9a9b0583","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f26358be95ed4c9a9b0583/W7-iCNB74C3wnYJe19N8f.jpeg","isPro":false,"fullname":"Anatoly Potapov","user":"AnatoliiPotapov","type":"user","name":"AnatoliiPotapov"},"name":"Anatolii Potapov","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:22.722Z","hidden":false}],"publishedAt":"2026-09-01T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix","submittedOnDailyBy":{"_id":"62609d224e6e4b84475eb8d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62609d224e6e4b84475eb8d9/4mHSdgtmIc88oqqzLLhJE.jpeg","isPro":false,"fullname":"Alex Medvedev","user":"kenkaneki","type":"user","name":"kenkaneki"},"summary":"Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimising all objectives jointly, which introduces cross-domain reward interference, we train a separate GRPO expert per axis and merge them via two-stage SLERP. Each expert's reward exposes a distinct failure mode, namely semantic collapse, over-calling, and verbosity hacking, each requiring a domain-specific fix. In non-reasoning mode the recipe surpasses a {sim}7times larger by total parameters baseline on the in-house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function-calling with 0.79 to 0.77, while lifting general dialogue benchmarks. The model absorbs 50% of platform traffic, 116M requests per month, at a fraction of the serving cost.","upvotes":28,"discussionId":"6a97f078d4c640d9b759ba54","projectPage":"https://huggingface.co/t-tech/T-pro-it-2.1","ai_summary":"A smaller self-hosted LLM trained with separate GRPO experts merged via SLERP outperforms a much larger baseline on instruction following, function-calling, and internal tasks while serving half of platform traffic at lower cost.","ai_keywords":["GRPO","SLERP","reward interference","semantic collapse","over-calling","verbosity hacking","deterministic verifiers","calibrated LLM judges","instruction following","function-calling"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"675861e944dbb69c2673c71c","name":"t-tech","fullname":"T-Tech","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/674ea07d320a043daeb2d98b/IwSCMolFY4Otk7sFXzWhi.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"677af99cdeb62f25f447ef08","avatarUrl":"/avatars/6b8b6fe8058b2136d35ccdad7c904c42.svg","isPro":false,"fullname":"Gleb Alektorov","user":"GlebAlektorov","type":"user"},{"_id":"62609d224e6e4b84475eb8d9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62609d224e6e4b84475eb8d9/4mHSdgtmIc88oqqzLLhJE.jpeg","isPro":false,"fullname":"Alex Medvedev","user":"kenkaneki","type":"user"},{"_id":"658bc20cfdf2279d4721f218","avatarUrl":"/avatars/5f1cb94373fbbbcfed9b848c5ebdd1ad.svg","isPro":false,"fullname":"Mikhail Gashkov","user":"MikeGashkov","type":"user"},{"_id":"64fb054ebb362cbf2fe53159","avatarUrl":"/avatars/936c37a77d46d0ea579d2f8a9aea9284.svg","isPro":false,"fullname":"Danil Taranets","user":"taranetsdan","type":"user"},{"_id":"64f4c8739ee58d48e8507e0e","avatarUrl":"/avatars/4be540dfb4a949f37cba2d3c3729fbde.svg","isPro":false,"fullname":"Dmitrii Stoianov","user":"heylimon","type":"user"},{"_id":"636142a6c12a09b8a31465ff","avatarUrl":"/avatars/8d67272b5bfcf70ac55473bd19f274d8.svg","isPro":false,"fullname":"Uliana","user":"ulianavin","type":"user"},{"_id":"6432343103d81fa4d26dac50","avatarUrl":"/avatars/5c4690b517360636ab7cd5a5bdbe6200.svg","isPro":false,"fullname":"Nikita Gulyakov","user":"elvispresniy","type":"user"},{"_id":"65f94be15df5183c9a9658fb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f94be15df5183c9a9658fb/s6SR9tQozzjKyddxXgAN4.jpeg","isPro":false,"fullname":"Arthur Babkin","user":"archeee","type":"user"},{"_id":"6780dcd6acf8d824c03864da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6PeN6OXbSq0M-L4OxFTrn.png","isPro":false,"fullname":"Ramil Latypov","user":"kylecr4ne","type":"user"},{"_id":"646f8807cc049b49e35e06dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646f8807cc049b49e35e06dc/UmPgkpEcqkn8n-OADa_nz.png","isPro":false,"fullname":"Arseny Ivanov","user":"ArsenyIvanov","type":"user"},{"_id":"64da40fad8d7d6d6d9aad5b2","avatarUrl":"/avatars/6bc3c6eb6cbcb8922d649635d22727da.svg","isPro":false,"fullname":"Ivan Listopadov","user":"Ivan1008","type":"user"},{"_id":"67912afe5144f1bb7d8c97fd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/bklmplMqoURksOEWOaIj6.png","isPro":false,"fullname":"Nikita","user":"Elysium-engine","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"675861e944dbb69c2673c71c","name":"t-tech","fullname":"T-Tech","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/674ea07d320a043daeb2d98b/IwSCMolFY4Otk7sFXzWhi.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.01572.md","query":{}}">
Papers
arxiv:2609.01572

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Published on Sep 1
· Submitted by
Alex Medvedev
on Sep 2

Abstract

A smaller self-hosted LLM trained with separate GRPO experts merged via SLERP outperforms a much larger baseline on instruction following, function-calling, and internal tasks while serving half of platform traffic at lower cost.

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimising all objectives jointly, which introduces cross-domain reward interference, we train a separate GRPO expert per axis and merge them via two-stage SLERP. Each expert's reward exposes a distinct failure mode, namely semantic collapse, over-calling, and verbosity hacking, each requiring a domain-specific fix. In non-reasoning mode the recipe surpasses a {sim}7times larger by total parameters baseline on the in-house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function-calling with 0.79 to 0.77, while lifting general dialogue benchmarks. The model absorbs 50% of platform traffic, 116M requests per month, at a fraction of the serving cost.

Community

Paper author Paper submitter about 5 hours ago

We consolidated traffic from 200+ internal apps onto one self-hosted model, closing the gaps production error analysis showed: instruction following, function calling, and our internal task mix. Instead of one joint objective we train a GRPO expert per axis and merge with two-stage SLERP — each axis hacks its reward differently (semantic collapse, over-calling, verbosity hacking). Non-reasoning mode beats a ~7× larger baseline on our Arena (69.6 vs 65.8) and now serves 50% of platform traffic, 116M requests/month.

Consolidating the whole corporate request mix onto one self-hosted model only pays off if the retraining loop closes the gap faster than the apps drift. In my experience the request mix shifts weekly — new tools, new prompts, new failure modes — and if your post-training cycle is a monthly batch job, you're always chasing last quarter's traffic. The number I'd want isn't coverage at snapshot time, it's the half-life of that coverage.

And I'd want to see the judge setup before trusting the quality numbers. "Calibrated LLM judges" is where these pipelines usually leak — if the judge was tuned on the same traffic you're optimizing for, you're measuring how well the model mimics the judge, not how well it serves the request. Show me the judge's disagreement rate with human raters on the hard tail, not the aggregate score.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.01572
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.01572 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.01572 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.01572 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers