Why do people keep fine-tuning on summarized/censored SOTA CoT traces?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Am I missing something? It seems like some people think distillation is magic and will raise the quality of output above what the base model is actually capable of. It's especially weird to me to see all these Fable fine-tunes, because as far as I understand it, they miss the fact that the reasoning traces you get from Anthropic's models are completely different from the actual chain of thought the model outputs, which makes it pretty much a guarantee that the result will be worse than before.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.