Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment
Abstract:Defect grading of power transmission equipment (DGPTE) is crucial to the stability of electric energy transmission. Although existing machine learning methods exhibit strong capabilities in defect detection, they are plagued by difficulties in integrating expert experience and facing class imbalance in more refined defect grading field. To address this issue, this paper introduces a novel defect grading framework based on multimodal large language model (MLLM). Specifically, this approach maximizes the commercial MLLMs' potential of DGPTE through in-context learning and obtains the state-of-te-art (SOTA) model. By sending a secondary request to this model, a small number of chain of thought-based question-answer pairs (Q\&As) are generated, which effectively reduces the cost of manual annotation. In this way, these high-quality interpretable Q\&As are used to train Qwen3-VL-8B via Low-Rank Adaption-based supervised fine-tuning (SFT). Experimental results on three DGPTE tasks demonstrate that fine-tuning only the language model layer yields the SOTA performance. Furthermore, multi-task joint fine-tuning verifies the feasibility of handling multiple grading tasks within only a single lightweight MLLM.
| Comments: | 9pages, 6figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.28822 [cs.CL] |
| (or arXiv:2605.28822v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.28822
arXiv-issued DOI via DataCite
|
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs
May 29
-
A Modular Architecture for Typologically Controlled Lexicon Generation
May 29
-
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
May 29
-
From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale
May 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.