r/LocalLLaMA · · 1 min read

Improved TPS of Gemma 4 31B : the journey and also creating custom patches with VLLM fork

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM and apply my own patch to break into making a configuration work as per my idea. I wrote a full article so that even beginners can understand and improve TPS for any model, this will serve as a mental model for anyone who start with TPS optimisations.

Also I strongly believe being just a mecha pilot wont be enough for inference optimisations. I had to go deep dive into the codebase and architecture for the solution to work.

P.S: Wrote in my own words.

https://medium.com/@abhijithneilabraham/learning-inference-how-to-host-and-improve-the-token-speed-of-an-llm-cff5623ab505?postPublishedType=initial

submitted by /u/metalvendetta
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA