Hugging Face Daily Papers · · 3 min read

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

How well can LLMs evaluate educational videos? We introduce EduPanel, a multi-agent evaluation framework that produces more reliable, interpretable, and pedagogically grounded judgments through specialized agents and extensive human validation!</p>\n","updatedAt":"2026-07-22T05:40:34.300Z","author":{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","fullname":"董家愷","name":"Snooow1029","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8716928362846375},"editors":["Snooow1029"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18529","authors":[{"_id":"6a6057157e7f152167e47229","name":"Jia-Kai Dong","hidden":false},{"_id":"6a6057157e7f152167e4722a","name":"Yi-Cheng Lin","hidden":false},{"_id":"6a6057157e7f152167e4722b","name":"Hung-yi Lee","hidden":false}],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration","submittedOnDailyBy":{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","isPro":false,"fullname":"董家愷","user":"Snooow1029","type":"user","name":"Snooow1029"},"summary":"Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.","upvotes":2,"discussionId":"6a6057157e7f152167e4722c","githubRepo":"https://github.com/snooow1029/edupanel","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"673248e121823ee4ea594099","name":"nationaltaiwan","fullname":"台灣大學","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67324880c1f20c742be144b8/CE1UiOtpMeC8pmtdGP4Nn.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67ecfe11534d68f5a87834c2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZnmCQzju-Okic9XdvJEvN.png","isPro":false,"fullname":"董家愷","user":"Snooow1029","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"673248e121823ee4ea594099","name":"nationaltaiwan","fullname":"台灣大學","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67324880c1f20c742be144b8/CE1UiOtpMeC8pmtdGP4Nn.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18529.md","query":{}}">
Papers
arxiv:2607.18529

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Published on Jul 20
· Submitted by
董家愷
on Jul 22
Authors:
,

Abstract

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable assessments for different aspects of teaching quality. Across expert studies, architecture ablations, and learner-persona analyses, EduPanel achieves reliability comparable to a median human expert. In expert evaluation, its feedback improves scoring accuracy (MAE 0.87 to 0.73), while experts remain able to detect unreliable outputs (AUC = 0.77) instead of accepting them blindly. These results suggest that EduPanel can serve as effective assistants for educational evaluation rather than replacements for human experts.

Community

Paper submitter about 3 hours ago

How well can LLMs evaluate educational videos? We introduce EduPanel, a multi-agent evaluation framework that produces more reliable, interpretable, and pedagogically grounded judgments through specialized agents and extensive human validation!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.18529
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.18529 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.18529 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.18529 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers