Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail<br>silently on real-time detectors, whose heterogeneous operators and detection-specific components<br>impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a<br>structure-aware framework that formulates adapter placement as an auditable constraint-planning<br>problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns<br>operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface,<br>and deployment predicates, records a reason code for each excluded module, and either emits a<br>budgeted target-module plan or returns Refuse before training. Under the official VOC07+12<br>trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95<br>on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On<br>RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic<br>threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A<br>controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent,<br>although training takes 1.72 times longer. Within the evaluated detector families, placement<br>policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error<br>with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal<br>on unseen detector architectures remains an open validation problem.</p>\n","updatedAt":"2026-08-10T03:27:03.207Z","author":{"_id":"64faed2e5ca946a010857aec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64faed2e5ca946a010857aec/eR8Hx0Dyy-DrPd1rmww8_.png","fullname":"Xu Lin","name":"gatilin","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":17,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8135710954666138},"editors":["gatilin"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64faed2e5ca946a010857aec/eR8Hx0Dyy-DrPd1rmww8_.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07051","authors":[{"_id":"6a7944ca8e9301703eaa5eb1","name":"Xu Lin","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb2","name":"WenJie Nie","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb3","name":"Jinlong Peng","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb4","name":"Weifu Fu","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb5","name":"YueXiao Ma","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb6","name":"Xiawu Zheng","hidden":false},{"_id":"6a7944ca8e9301703eaa5eb7","name":"Yong Liu","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family","submittedOnDailyBy":{"_id":"64faed2e5ca946a010857aec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64faed2e5ca946a010857aec/eR8Hx0Dyy-DrPd1rmww8_.png","isPro":true,"fullname":"Xu Lin","user":"gatilin","type":"user","name":"gatilin"},"summary":"Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specific components impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a structure-aware framework that formulates adapter placement as an auditable constraint-planning problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface, and deployment predicates, records a reason code for each excluded module, and either emits a budgeted target-module plan or returns Refuse before training. Under the official VOC07+12 trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95 on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent, although training takes 1.72 times longer. Within the evaluated detector families, placement policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal on unseen detector architectures remains an open validation problem. Project Page: github.com/Tencent/YOLO-Master","upvotes":9,"discussionId":"6a7944ca8e9301703eaa5eb8","organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6537ad2189cdab24b8338be1","avatarUrl":"/avatars/4248cc9c4a5097653f6def9258b0f942.svg","isPro":false,"fullname":"weifu","user":"fuweifu","type":"user"},{"_id":"671a2c80d057ae117be351db","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/UgUEKNN6Tr1lp0IBvTv9F.png","isPro":false,"fullname":"Jinlong","user":"pjl1995","type":"user"},{"_id":"668366cd95989c5c95c8df29","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/KuRGU_akSi8YVczPKJbX6.jpeg","isPro":false,"fullname":"xuefeng","user":"real-xuefeng","type":"user"},{"_id":"65905af887944e494e37e09a","avatarUrl":"/avatars/8f079ef18d06f506b94edca7d49e4c26.svg","isPro":false,"fullname":"Hert4","user":"beyoru","type":"user"},{"_id":"679a338e76dc38ae6348949b","avatarUrl":"/avatars/30b9fec1adbc28871d6d72176aa5f5ea.svg","isPro":false,"fullname":"sweethxchen","user":"sweethx","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6407e5294edf9f5c4fd32228","avatarUrl":"/avatars/8e2d55460e9fe9c426eb552baf4b2cb0.svg","isPro":false,"fullname":"Stoney Kang","user":"sikang99","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"619f9755da83161f25840698","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/619f9755da83161f25840698/FM421pE1mz5v1YhrxA8ZA.jpeg","isPro":false,"fullname":"Muhammad Umair","user":"umair894","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"66543b6e420092799d2f625c","name":"tencent","fullname":"Tencent","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07051.md","query":{}}">
YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
Abstract
Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specific components impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a structure-aware framework that formulates adapter placement as an auditable constraint-planning problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface, and deployment predicates, records a reason code for each excluded module, and either emits a budgeted target-module plan or returns Refuse before training. Under the official VOC07+12 trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95 on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent, although training takes 1.72 times longer. Within the evaluated detector families, placement policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal on unseen detector architectures remains an open validation problem. Project Page: github.com/Tencent/YOLO-Master
Community
Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail
silently on real-time detectors, whose heterogeneous operators and detection-specific components
impose placement constraints absent from regular Transformer stacks. We propose YOLO-PEFT, a
structure-aware framework that formulates adapter placement as an auditable constraint-planning
problem. Given a detector graph, a PEFT request, and a resource budget, YOLO-PEFT assigns
operator and semantic roles, evaluates explicit operator-validity, detector-semantic, graph-interface,
and deployment predicates, records a reason code for each excluded module, and either emits a
budgeted target-module plan or returns Refuse before training. Under the official VOC07+12
trainval-to-VOC07 test protocol, planner-selected RS-LoRA reaches 0.7138 and 0.7307 mAP50-95
on YOLO11s and YOLO12s, respectively, compared with 0.6428 and 0.6662 for Full-SFT. On
RT-DETR-L, all seven evaluated LoRA-family configurations cross the predefined catastrophic
threshold, supporting a calibrated Refuse-to-Full-SFT decision within the evaluated coverage. A
controlled YOLO11 audit further shows that LoRA reduces peak training memory by 43.9 percent,
although training takes 1.72 times longer. Within the evaluated detector families, placement
policies, and calibration coverage, YOLO-PEFT replaces manual target-module trial and error
with explicit, inspectable planning while preserving verified train-save-merge-export paths; refusal
on unseen detector architectures remains an open validation problem.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.07051 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.07051 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.07051 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.