Image Processing model and Audio Processing model on 32GB VRAM and 64GB RAM?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I have been playing around with LLMs on a dual 5060ti (Windows) rig, and now want to change things up.
I built a separate dual 5070ti (Debian) rig and now have that running Qwen 3.6 27b UD Q6 MTP @ 100k context, without any multimodal capacity. That's solid for the text processing and generation stuff I want to do right now (which is mainly IT support tickets triage and PowerShell scripting via OpenCode as the "harness").
This now leaves my dual 5060ti rig ready for repurposing. I'd been having trouble with "terminal loops" in Qwen and Gemma. My plan is to flatten the OS and start again with Debian as that's been solid.
I'm thinking the 5060ti rig could augment the text processing of the 5070ti rig, and handle things like deciphering screenshots and other images in tickets. I'm wondering what models out there excel at that in the sub 32GB VRAM space, and if I can have enough VRAM spare to run something alongside it that could process Audio (I'm thinking voicemail and voice note transcriptions mainly). It would be a bonus if it could generate audio, but not essential right now.
I'd prefer to keep to GGUFs so I can use llama.cpp on the 5060ti rig to keep things uniform between the rigs, if possible.
These two rigs are dedicated to the AI models running on them. I'll handle the orchestration myself (largely via n8n) from another machine.
What are your recommendations for models to consider, please?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.