Hermes on Android (Graphene OS)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL
I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gateway running on my laptop as the backend, with Llama.cpp and Qwen 3.6 35b. Paired my mobile 5070ti with an RTX 3090 eGPU on TB5 for about 34GB of VRAM.
2500 tps prompt processing
80-150 tps generation w/mtp
Incredible setup imo. Light and fast. 100% local. Just wanted to show off whats possible and how cool the Hermes GUI can look on your phone. Im using the Razer Core X V2. Unfortunately I could not get it working in Linux but it works great in Windows 11. I previously had a full tower desktop with 2x 3090s and for some reason this thing is faster paired with the 5070ti over TB5.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.