Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| Hey r/MachineLearning! I’ve been working on a new optimizer called Tauon (turns out there is already "teon" but well if you have better idea, - i will gladly accept it! Anyway the core idea of optimizer is about polynomials and orthogonalization just like muon, the whole difference is that i managed to lower total number of steps (first through spectral filtering down to 3 steps then through coeff scheduling down to 2) + reduced matrix size (through dct-2). And I wanted to share some initial benchmark results... Benchmark Setup: Trained a GPT-Mini (d_model=512, 6 Layers) on TinyShakespeare against Muon and AdamW.
Results:
And yeah i know that its hilariously tiny benchmark but well i have only 2 hours left on my kaggle free T4 so i really couldnt more + i hope someone would be able test it on a bigger setup! Links & Code:
Would love to get your thoughts on the optimizer! If you have any ideas, suggestions - please tell me. Cheers, everyone! [link] [comments] |
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.