How well can LLMs evaluate educational videos? We introduce EduPanel, a multi-agent evaluation framework that produces more reliable, interpretable, and pedagogically grounded judgments through specialized agents and extensive human validation!</p>\n","updatedAt":"2026-07-22T05:40:34.300Z","author":{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","fullname":"董家愷","name":"Snooow1029","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8716928362846375},"editors":["Snooow1029"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18529","authors":[{"_id":"6a6057157e7f152167e47229","name":"Jia-Kai Dong","hidden":false},{"_id":"6a6057157e7f152167e4722a","name":"Yi-Cheng Lin","hidden":false},{"_id":"6a6057157e7f152167e4722b","name":"Hung-yi Lee","hidden":false}],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration","submittedOnDailyBy":{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","isPro":false,"fullname":"董家愷","user":"Snooow1029","type":"user","name":"Snooow1029"},"summary":"Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.","upvotes":2,"discussionId":"6a6057157e7f152167e4722c","githubRepo":"https://github.com/snooow1029/edupanel","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"673248e121823ee4ea594099","name":"nationaltaiwan","fullname":"台灣大學","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67324880c1f20c742be144b8/CE1UiOtpMeC8pmtdGP4Nn.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","isPro":false,"fullname":"董家愷","user":"Snooow1029","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"673248e121823ee4ea594099","name":"nationaltaiwan","fullname":"台灣大學","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67324880c1f20c742be144b8/CE1UiOtpMeC8pmtdGP4Nn.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18529.md","query":{}}">
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Published on Jul 20
· Submitted by 董家愷 on Jul 22 Abstract
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.
Community
How well can LLMs evaluate educational videos? We introduce EduPanel, a multi-agent evaluation framework that produces more reliable, interpretable, and pedagogically grounded judgments through specialized agents and extensive human validation!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.18529 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.18529 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.18529 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.