My frozen-encoder decision heads were reading 3 tokens per label: fixing the option budget took banking77 from 67.5% to 76.0% [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| I distil the closed decisions an app asks an LLM for (pick a label, yes/no, a score) into small heads on a frozen 400M ModernBERT encoder, the open Laya checkpoint. Per decision I train the head Laya ships (type embedding, two transformer layers, a scorer over option markers), fit a temperature on a holdout, pick the operating threshold from the coverage curve at a target agreement, and fall back to the teacher below it. Three things came out of the last week of measuring it. 1. The model was reading about 3 tokens of each label. Laya gives all the options together a 192-token budget. With 77 labels each one gets cut to 4 tokens including its marker, so
For reference, the hosted Jev model scores 76.4% on the 500-row jevbench sample, and a fully fine-tuned ModernBERT-base reaches ~94%, so the frozen encoder is the ceiling here, not the head. 2. A frozen encoder learns what the text states, not arithmetic. A payment-risk rule over amount, hour and country reached 0.42 holdout agreement. The same rule written as a sentence about the customer ("first transfer to this payee, larger than usual") reached 0.94, same rows, same head. 3. Agreement with the teacher is not accuracy, so I added a gold check. Caching the encoder output once per decision makes 24 epochs over 8,000 rows an 18-minute job on a laptop RTX 5060. Code, data scripts and every number: github.com/bladedevoff/stuntd (Apache-2.0). Browser demo: huggingface.co/spaces/pollix/stuntd Next I want to unfreeze the top encoder layers while keeping the cache (cache the output of layer 26 of 28), to see how much of the gap to a full fine-tune that closes. Happy to hear if someone has measured that trade-off. [link] [comments] |
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.