r/LocalLLaMA · · 1 min read

I have just moved from MacBook M5 pro 48 GB to RTX3090

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s

I think I can finally get rid of my Claude subscription, this is good enough for me. I am a software dev and I can get what I need from this set up and be more productive. I am using the linux machine serving the model over an Open AI endpoint and using a custom build desktop app with pi behind it all.

Regarding speeds average 20t/s on Mac with llama.cpp with mtp Unsloth Q6, Q4 on Mac didnt make much difference in speed for me.

Linux running https://github.com/syv-ai/qwen38-27b-rtx3090 which is vLLM and a Q4 model I believe. Tool calls definitely fail more but qwen3.8 seems smart enough to fix and correct itself

submitted by /u/gutard
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA