r/LocalLLaMA · · 1 min read

POCKET-35B agentic model on cpu 59 t/s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

POCKET-35B agentic model on cpu 59 t/s

https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF

A 35B model that runs on your PC with no GPU — and on your phone. Just stock llama.cpp. No fork, no CUDA, no cloud.

The POCKET lineup — pick by your device

Repo File Size Runs on Best for Korean PPL*
POCKET-35B-GGUF Q4_K_M 21 GB PC / server (32 GB RAM) top quality 5.79
POCKET-35B-GGUF Q2_K 13 GB mini-PC, no GPU daily driver 6.49
POCKET-35B-GGUF IQ1_M 8.2 GB 16 GB RAM box smallest full model 9.69
POCKET-KR-GGUF IQ2_M 5.1 GB Android 8 GB+ 🇰🇷 Korean phone 7.95
POCKET-KR-MLX 2-bit 5.1 GB 🍎 iPhone / iPad / Mac 🇰🇷 Korean, Apple-native 7.95
POCKET-EN-GGUF iPhone-mix 5.3 GB 🍎 iPhone (PocketPal) 🌍 English phone
POCKET-EN-GGUF PC-mix 6.8 GB PC / Android 🌍 English, best quality
submitted by /u/Powerful_Evening5495
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA