Android Studios native Gemma 4 runs on llama.cpp
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I'm not sure how many people care about Android Studio, but I think it's cool that Google uses llama.cpp. My guess is that it is Vulkan and the QAT versions of Gemma 4. It supports multi-GPU and 31B has a max. context length of 128k. It uses 34 GB VRAM when fully loaded. I don't see an option to change the context length or show PP/TG speed. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.