Improved TPS of Gemma 4 31B : the journey and also creating custom patches with VLLM fork
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM and apply my own patch to break into making a configuration work as per my idea. I wrote a full article so that even beginners can understand and improve TPS for any model, this will serve as a mental model for anyone who start with TPS optimisations.
Also I strongly believe being just a mecha pilot wont be enough for inference optimisations. I had to go deep dive into the codebase and architecture for the solution to work.
P.S: Wrote in my own words.
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.