LLM-as-a-judge is now everywhere for automated evaluation. But it can be slow, expensive, and opaque. What if we ask the judge for its rubric once, and execute that logic as a program? Introducing PAJAMA—a new hybrid evaluation system that pushes the LLM-judge Pareto frontier!</p>\n","updatedAt":"2026-07-28T05:35:17.018Z","author":{"_id":"6363edf3de4bc2f294accb16","avatarUrl":"/avatars/9baddfec7170ec662e974116f561ab2c.svg","fullname":"Tzu-Heng Huang","name":"zihengh1","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9165549278259277},"editors":["zihengh1"],"editorAvatarUrls":["/avatars/9baddfec7170ec662e974116f561ab2c.svg"],"reactions":[],"isReport":false}},{"id":"6a689c5b2a56aab6e9583240","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-07-28T12:11:07.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"I've been running LLM-as-a-judge in production for a while now, and the cost-latency-transparency triple pain point is real. This paper's idea — distill the judge's decision logic into actual programs instead of prompting an LLM at eval time — is the kind of thing that makes you wonder why more people aren't doing it. The committee-of-programs approach means you can inspect, edit, and version-control your eval criteria like any other code. I'd want to see how brittle the distilled programs are when the distribution shifts, but for stable eval tasks this could cut the API bill to near zero. Going to try the PAJAMA system on my own agent eval suite this week.","html":"<p>I've been running LLM-as-a-judge in production for a while now, and the cost-latency-transparency triple pain point is real. This paper's idea — distill the judge's decision logic into actual programs instead of prompting an LLM at eval time — is the kind of thing that makes you wonder why more people aren't doing it. The committee-of-programs approach means you can inspect, edit, and version-control your eval criteria like any other code. I'd want to see how brittle the distilled programs are when the distribution shifts, but for stable eval tasks this could cut the API bill to near zero. Going to try the PAJAMA system on my own agent eval suite this week.</p>\n","updatedAt":"2026-07-28T12:11:07.829Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9430521130561829},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[{"reaction":"🔥","users":["abeQ213"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.22561","authors":[{"_id":"6a683f3f73f69d5af2bec8d5","name":"Tzu-Heng Huang","hidden":false},{"_id":"6a683f3f73f69d5af2bec8d6","name":"Shengqi Qiu","hidden":false},{"_id":"6a683f3f73f69d5af2bec8d7","name":"Frederic Sala","hidden":false}],"publishedAt":"2026-05-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"Codifying the Judge: Scalable Evaluation via Program Distillation","submittedOnDailyBy":{"_id":"6363edf3de4bc2f294accb16","avatarUrl":"/avatars/9baddfec7170ec662e974116f561ab2c.svg","isPro":false,"fullname":"Tzu-Heng Huang","user":"zihengh1","type":"user","name":"zihengh1"},"summary":"LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.","upvotes":5,"discussionId":"6a683f4073f69d5af2bec8d8","projectPage":"https://sprocketlab.github.io/PAJAMA/","githubRepo":"https://github.com/SprocketLab/PAJAMA","githubRepoAddedBy":"user","githubStars":11,"organization":{"_id":"63aaa3e7b7f3e1c60728194a","name":"sprocket-lab","fullname":"sprocket-lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/629436f4fa1501cdf1a1086e/kpkvUxJsH_R9LHSOk6WdP.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6363edf3de4bc2f294accb16","avatarUrl":"/avatars/9baddfec7170ec662e974116f561ab2c.svg","isPro":false,"fullname":"Tzu-Heng Huang","user":"zihengh1","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"698e15b1f2dc6c20c89e3d19","avatarUrl":"/avatars/12e456599cb74c43ded4d97da855527e.svg","isPro":false,"fullname":"Shengqi Qiu","user":"abeQ213","type":"user"},{"_id":"676f6c0ccb094bb8d8a0948f","avatarUrl":"/avatars/78d26c39e1108ee4029639e4202d0ec9.svg","isPro":false,"fullname":"Xuwei Ding","user":"Xuwei04","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63aaa3e7b7f3e1c60728194a","name":"sprocket-lab","fullname":"sprocket-lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/629436f4fa1501cdf1a1086e/kpkvUxJsH_R9LHSOk6WdP.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.22561.md","query":{}}">
Codifying the Judge: Scalable Evaluation via Program Distillation
Abstract
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
Community
LLM-as-a-judge is now everywhere for automated evaluation. But it can be slow, expensive, and opaque. What if we ask the judge for its rubric once, and execute that logic as a program? Introducing PAJAMA—a new hybrid evaluation system that pushes the LLM-judge Pareto frontier!
I've been running LLM-as-a-judge in production for a while now, and the cost-latency-transparency triple pain point is real. This paper's idea — distill the judge's decision logic into actual programs instead of prompting an LLM at eval time — is the kind of thing that makes you wonder why more people aren't doing it. The committee-of-programs approach means you can inspect, edit, and version-control your eval criteria like any other code. I'd want to see how brittle the distilled programs are when the distribution shifts, but for stable eval tasks this could cut the API bill to near zero. Going to try the PAJAMA system on my own agent eval suite this week.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.22561 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.22561 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.22561 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.