r/MachineLearning · · 2 min read

My frozen-encoder decision heads were reading 3 tokens per label: fixing the option budget took banking77 from 67.5% to 76.0% [P]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

My frozen-encoder decision heads were reading 3 tokens per label: fixing the option budget took banking77 from 67.5% to 76.0% [P]

I distil the closed decisions an app asks an LLM for (pick a label, yes/no, a score) into small heads on a frozen 400M ModernBERT encoder, the open Laya checkpoint. Per decision I train the head Laya ships (type embedding, two transformer layers, a scorer over option markers), fit a temperature on a holdout, pick the operating threshold from the coverage curve at a target agreement, and fall back to the teacher below it. Three things came out of the last week of measuring it.

1. The model was reading about 3 tokens of each label. Laya gives all the options together a 192-token budget. With 77 labels each one gets cut to 4 tokens including its marker, so declined_card_payment, declined_cash_withdrawal and declined_transfer all became declined_. A Reddit commenter ran the ModernBERT tokenizer over the 77 names: with underscores they average 6.48 tokens and 18 collide at a 3-token cut; with spaces 3.74 tokens and 7 collide. Sizing the option window per decision from its labels and showing labels with spaces moved banking77 on the full official test split (3,080 rows):

accuracy
zero-shot Laya 38.2% (jevbench sample)
trained head, 192-token budget 67.5%
trained head, window sized to the labels 75.2%
same, 48 epochs (picked on the holdout) 76.0% (±1.5)

For reference, the hosted Jev model scores 76.4% on the 500-row jevbench sample, and a fully fine-tuned ModernBERT-base reaches ~94%, so the frozen encoder is the ceiling here, not the head.

2. A frozen encoder learns what the text states, not arithmetic. A payment-risk rule over amount, hour and country reached 0.42 holdout agreement. The same rule written as a sentence about the customer ("first transfer to this payee, larger than usual") reached 0.94, same rows, same head.

3. Agreement with the teacher is not accuracy, so I added a gold check. report --gold scores the head and the teacher against rows a person verified, and splits rows the head may have trained on from new ones. On a banking demo: 100% on rows the store had seen versus 93.5% on new rows, which is why the split matters. The threshold picked on the holdout held on fresh data: 99.1% promised, 99.0% served, 55% of requests answered locally.

Caching the encoder output once per decision makes 24 epochs over 8,000 rows an 18-minute job on a laptop RTX 5060. Code, data scripts and every number: github.com/bladedevoff/stuntd (Apache-2.0). Browser demo: huggingface.co/spaces/pollix/stuntd

Next I want to unfreeze the top encoder layers while keeping the cache (cache the output of layer 26 of 28), to see how much of the gap to a full fine-tune that closes. Happy to hear if someone has measured that trade-off.

submitted by /u/Inevitable-Log5414
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning