r/LocalLLaMA · · 12 min read

Cloud-AI Cold-Turkey: Real Dev Work with Local AI (Ornith 1.5 35b-a3b and Qwen 3.8-27b; 8GB VRAM vs 32GB VRAM)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Cloud-AI Cold-Turkey: Real Dev Work with Local AI (Ornith 1.5 35b-a3b and Qwen 3.8-27b; 8GB VRAM vs 32GB VRAM)

Spent several weeks on an 'as-much-Local-AI-as-possible' regime and have been - mostly - impressed.

Yes, Qwen 3.8 27b is the (rightful) star of the show (obligatory one-shot-mario-build-reddit-comment here!). But don't underestimate what models with more modest hardware requirements can do.

With Qwen forsaking us 30b-a3b-enjoyers in their latest releases, I figured I'd share my experiences with Ornith 1.5 35b-a3b - which offers that size and plays well agentically. I picked it over the similarly sized "Qwen 3.6 35b-a3b" MoE model because that Qwen MoE has no thinking levels (beyond 'On' or 'Off') and hence tends to underthink and thus undercook its answers. I had some luck queuing a 'double-check your results' follow-ups with it; but more thinking off-the-bat would be first prize.

Ornith (and a few other similar models) addresses this through additional training that leaves if feeling like the equivalent of a High reasoning mode for the Qwen MoE (I don't know how much additional knowledge it's acquired - but the fact that it seems to work harder legitimately improves the results in my experience).

Harness

Started with Pi; found that whilst Qwen 3.8 27b was fine within it, Ornith battled with file-writes/edits frequently failing/retrying.

Swapped over to OpenCode (full OpenCode-via-llama model config is below) which resolved this at the expense of a higher base context (about 11k tokens in my own setup - which includes a few optional plugins)

I already use OpenCode across the board for all my cloud AI uses (mainly OpenAI/GLM) so I've been happy with this consolidation move overall; swapping to a cloud model mid-chat when things get complicated works well for me; this is where OpenCode excels.

Hardware and Performance

30tps [100k-140k-context] up to 40tps [<100k-context] TG; 300-400 tps PP.

Running on a NVIDIA RTX 5060 Mobile (8GB VRAM); Intel Core Ultra 9 275HX; 32GB DDR5 RAM.

This took some tweaking which, again, I'll detail later.

And yes, this pales in comparison with the 5090 server setup; but for the more limited hardware, I'm actually happy with the results and find it reasonably snappy in use - with a few OpenCode usage-optimizations (again, will detail later!)

First Test: Web OS

Ran Ornith through a Bijan-Bowen inspired web-OS 'get-a-feel-for-the-model' test: basics worked out the box but required two additional turns to resolve minor bugs (maximize/minimize buttons didn't work; snake-game instant-died upon game-start).

It's here: https://jsfiddle.net/db89uwpj/

Not mind-blowing but good enough - and we can see some similarities to Qwen 3.8 27b from Bijan's Qwen video; both models clearly share some of the same lineage : https://youtu.be/6kjXzTVmT58?t=442

Second: Network Troubleshooting

Experienced several random router-drop-outs on my home network over the span of several days - since the log contains credentials that I didn't want to clean for online submission (thanks for nothing, Asus!) I dumped it straight into Ornith - fully locally - and within 10 minutes it had identified:

  1. The crash was localized to the WLAN sub-system only; not the whole router.
  2. WPS was a likely culprit since the WLAN-drop-out seemed to be preceded by a WPS error.

Since the logs are rolled-over about twice a day, it offered to create a monitoring system to autonomously log onto the router every few hours (since telnet is disabled), grab/merge the logs into a single consolidated log per-day, and monitor (and trigger a console alert) for when the scenario recurred.

Sounded ambitious - but indeed, it successfully built an app to authenticate via the login page, navigate to system-log page, capture the log-control textbox contents, and snapshot it to disk, merging it with the current day's logs - and it then set a scheduled-job to run it hourly.

It worked - one shot.

But it didn't need to: I disabled WPS as per its original suggestion and the issue simply hasn't recurred since. Ornith's diagnosis appears to have been spot-on.

Promising?

Third: Some Real Work

And... here's where things got complicated.

I'm involved in several large-to-medium-scale software systems that tend to have many sub-systems that interact. The documentation thereof is OK but certainly not exhaustive.

Getting either Ornith 1.5 35b-a3b or Qwen 3.8 27b to do solid planning on changes to sections of these systems - especially when they interact with other dependencies - has been a total crapshoot.

If the features were constrained to a maximum of 2 or 3 files, both models did surprisingly well with both planning and implementation.

If the features went beyond that but were typical 'modify-DAL-then-Business-Layer-then-Frontend' type changes, Qwen usually handled well (with Ornith trailing - doing just OK here).

But the minute changes required, say, a method signature change that would have 5 or 6 calls (across as many source files) require updates, both models simply made mistakes that indicated misunderstanding of what the code did; regardless of the system I tried it in. It would compile just fine - but be buggy and often non-functional.

So in the case of both models, the only option was to get a larger model (GLM 5.3 or - my preferred planner - GPT 5.6 Sol-High) to do the planning and then use Ornith or Qwen as an implementation model.

For production code I also found it was important to manually review the diffs and get Sol-High to review; it mostly over-engineered edge-cases (which I'd ignore) but sometimes, post-implement, it'd plug significant gaps.

For simpler changes, Qwen was usually superior but sometimes overcomplicated/overthink'ed (I use Qwen 3.8 27b on its xhigh reasoning level). So occasionally, Ornith produced the better solution.

In more than one instance when changes were contained to just a few code files, even Sol found itself impressed with Ornith:

Sol-High's Ornith ASP.Net Code Review

'Damned impressive' is not how Sol usually reviews other models' code... good showing from Ornith; I laughed watching Sol do a web-search to verify if Ornith's solution was actually feasible (it's been flawless in prod since!).

I even had it correct a few bugs in Astra's code on a test game back when Astra launched (I bench new models on games!), which it again handled well, getting a nod from Astra:

Astra 'Noticing' Ornith's Bug-Fix

And another Bijen-Bowen inspired test (a 3D Subway) by Ornith:

Ornith's Subway-Scene Attempt. It's... meh.

...this one falls well short of Qwen 3.8 27b's output (refer to Bijan's video again: https://youtu.be/6kjXzTVmT58?t=1043 - this is an incredible showing from Qwen) but was nonetheless respectable for the model size.

It's not all bad: in production work, Sol's reviews sometimes preferred Ornith's output to Qwen in a few instances:

Ornith 1.5 35b-a3b beats out Qwen 3.8 27b. Usually it's the opposite, though.

...but as was often the case in my use, neither was production-ready; and in most cases Qwen won out.

In short

Having shipped a few thousand lines of code from each model, for contained tasks - especially when I have a decent understanding of what needs to change and what basic coding steps to take - I'm honestly happy using Ornith whenever away from my Qwen server. It punches above its weight(s) (and I imagine similar models like Tiel would do equally as well) whilst maintaining acceptable speed on my stand-alone laptop; especially when given decently-detailed prompting as guidance.

But even with the dense-model Qwen on standby, I can't work efficiently without having a larger model like GLM 5.3 or GPT-Sol/Astra handling the planning and review.

Hence, neither local model is a substitute for my cloud subs in a professional dev setting quite yet; though Qwen 3.8 27b feels a lot like Luna-High on implementation; whilst Ornith feels somewhere between -Medium and -Low.

Very, very impressive showings for both, taking into account their respective sizes and hardware compatibility: Qwen 3.8 27b is the most capable local model I've ever run; and Ornith delivers far more than 25% of its capability in just 25% of it's VRAM footprint.

Performance Tricks

OpenCode

  1. For boosting performance, I vibe-coded a /title plugin that manually titles new OpenCode sessions by looking for a /title keyword on the first message; this saves a full round-trip to llama.cpp for generating a title from the prompt (and also avoids a possible cache-flush if I start a /New session on the same project).
  2. I also always send a starter-prompt (like a full-stop) to OpenCode to get the model loaded (with current-project context/agent.md) whilst I type the full actual prompt out, so that lead-time gets reduced (because by the time the real prompt is typed, prompt-processing on the project context - thanks to that first starter-message - is done and the model is ready to process the actual prompt without starting from scratch).
  3. I also vibe-coded a current-speed-indicator for OpenCode that displays Prompt-Processing Speed (and percentage) as well as Token-Generation-Speed so I can see what performance looks like in realtime.

General

  1. Swapping my external displays to the Intel iGPU (by using a 2-HDMI USB-C hub) leaves the entire Nvidia GPU's 8GB of VRAM available to the model (only iGPU VRAM gets consumed by Windows); actual Ornith use was about 7GB VRAM and 25GB system RAM.
  2. Strangely, running my laptop in Balanced (rather than Performance) mode gave me better sustained speeds (and quieter fans); presumably Performance introduced throttling.

The Technical Details - OpenCode Config

I use OpenCode with two llama.cpp servers: Ornith 1.5 35B-A3B on my local 8 GB GPU, and Qwen3.8 27B on a separate RTX 5090 machine. These are the current model entries from my OpenCode config - both servers use llama.cpp b10622. Add both provider entries under provider in opencode.json; replace the placeholders with your own values.

Ornith

Model download: Ornith 1.5 35B-A3B repaired MTP GGUF - file ornith15-trained-head.gguf (17.44 GB). This is the APEX-MTP Compact mixed quant built from mudler's APEX release, with a continued-trained MTP head spliced into the same GGUF. The repaired head is already inside the file; so no separate MTP model download is needed.

{ "ornith-local": { "npm": "@ai-sdk/openai-compatible", "name": "Ornith 1.5 35B A3B local llama.cpp", "options": { "baseURL": "http://<LOCAL_LOOPBACK_ADDRESS>:1235/v1", "apiKey": "", "timeout": false, "headerTimeout": false }, "models": { "ornith15-35b-a3b": { "id": "ornith-ai/ornith-1.5-35b-a3b", "name": "Ornith 1.5 35B A3B repaired MTP 144K", "reasoning": true, "tool_call": true, "temperature": true, "interleaved": "reasoning_content", "modalities": { "input": ["text"], "output": ["text"] }, "limit": { "context": 147456, "output": 16384 }, "options": { "temperature": 0.6, "top_p": 0.95, "top_k": 20, "min_p": 0.0, "presence_penalty": 0.0, "repeat_penalty": 1.0, "chat_template_kwargs": { "enable_thinking": true, "preserve_thinking": true }, "parallel_tool_calls": false, "thinking_budget_tokens": 8192 }, "variants": { "fast": { "chat_template_kwargs": { "enable_thinking": false, "preserve_thinking": false }, "parallel_tool_calls": false, "thinking_budget_tokens": 0 } } } } } } 

Qwen

Model download: Unsloth Qwen3.8 27B GGUF - file Qwen3.8-27B-UD-Q5_K_M.gguf (19.8 GB / 18.4 GiB). My server installation pinned revision 313447f257f7ebde0b968e4778feef774546ed81. The server runs on the RTX 5090 machine; OpenCode connects over my LAN.

{ "qwen5090": { "npm": "@ai-sdk/openai-compatible", "name": "Qwen3.8 27B on RTX 5090", "options": { "baseURL": "http://<REMOTE_SERVER_ADDRESS>:11435/v1", "apiKey": "", "timeout": false, "headerTimeout": false }, "models": { "qwen38-27b": { "id": "qwen/qwen3.8-27b", "name": "Qwen3.8 27B RTX 5090", "reasoning": true, "tool_call": true, "temperature": true, "interleaved": "reasoning_content", "modalities": { "input": ["text"], "output": ["text"] }, "limit": { "context": 122880, "output": 16384 }, "options": { "chat_template_kwargs": { "enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "medium" }, "parallel_tool_calls": true, "thinking_budget_tokens": 8192 }, "variants": { "low": { "chat_template_kwargs": { "enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "low" }, "parallel_tool_calls": true, "thinking_budget_tokens": 2048 }, "medium": { "chat_template_kwargs": { "enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "medium" }, "parallel_tool_calls": true, "thinking_budget_tokens": 8192 }, "xhigh": { "chat_template_kwargs": { "enable_thinking": true, "preserve_thinking": true, "reasoning_effort": "xhigh" }, "parallel_tool_calls": true, "thinking_budget_tokens": 12288 } } } } } } 

The Technical Details - Llama.cpp Config

Runtime: llama.cpp b10622 release. Ornith uses the Windows CUDA 12.4 build; Qwen uses the Windows CUDA 13.3 build. The code below is an excerpt of my PowerShell script:

Ornith

My launcher starts launch-model.ps1 -Target ornith-opencode -ServerOnly, which calls start-ornith.ps1. The currently running first-attempt MTP instance has 155648 context, 41 GPU layers, and the MoE expert placement shown below. This is therefore it's effective llama-server configuration. I run it from the folder containing llama-server.exe after replacing the placeholders. The expert placement is specific to my 8 GB GPU but should work on others with the same capacity.

$model = '<PATH_TO_ornith15-trained-head.gguf>' $cpuExpertPlacement = @( 'blk\.7\.ffn_(gate|up|gate_up|down).*=CPU' 8..40 | ForEach-Object { 'blk\.{0}\.ffn_(up|down|gate_up|gate)_(ch|)exps=CPU' -f $_ } ) -join ',' $serverArgs = @( '--model', $model, '--alias', 'ornith-ai/ornith-1.5-35b-a3b', '--host', '<LOCAL_LOOPBACK_ADDRESS>', '--port', '1235', '--api-key', '<YOUR_ORNITH_API_KEY>', '--ctx-size', '155648', '--parallel', '1', '--cache-prompt', '--cache-ram', '0', '--no-cache-idle-slots', '--slots', '--batch-size', '2048', '--ubatch-size', '512', '--threads', '20', '--threads-batch', '24', '--flash-attn', 'on', '--cache-type-k', 'q8_0', '--cache-type-v', 'q8_0', '--gpu-layers', '41', '--override-tensor', $cpuExpertPlacement, '--fit', 'off', '--fit-target', '192', '--load-mode', 'mmap', '--jinja', '--reasoning', 'auto', '--reasoning-format', 'deepseek', '--reasoning-budget', '8192', '--reasoning-preserve', '--temp', '0.6', '--top-p', '0.95', '--top-k', '20', '--min-p', '0.0', '--repeat-penalty', '1.0', '--presence-penalty', '0.0', '--spec-type', 'draft-mtp', '--spec-draft-n-max', '2', '--spec-draft-type-k', 'q4_0', '--spec-draft-type-v', 'q4_0', '--spec-draft-threads', '20', '--spec-draft-threads-batch', '24', '--metrics', '--log-colors', 'off', '--log-timestamps' ) & .\llama-server.exe @serverArgs 

Main K/V cache is Q8_0; the MTP draft K/V cache is Q4_0. --spec-draft-n-max 2 uses the GGUF's repaired embedded MTP head; there is no separate draft model. OpenCode advertises 147456 tokens even though the server allocates 155648.

Qwen

This runs on the RTX 5090 server via start-server.cmd -> start-server.ps1. Its current host-config.json sets a 131072-token server context, one slot, and the Unsloth UD-Q5_K_M file. The server uses all GPU layers with --fit off; it does not use the smaller local Qwen profile. Usage is just under 26GB of VRAM with no context loaded.

$serverArgs = @( '--model', '<PATH_TO_Qwen3.8-27B-UD-Q5_K_M.gguf>', '--alias', 'qwen/qwen3.8-27b', '--host', '<YOUR_SERVER_BIND_ADDRESS>', '--port', '11435', '--api-key-file', '<PATH_TO_YOUR_API_KEY_FILE>', '--ctx-size', '131072', '--parallel', '1', '--cache-prompt', '--slots', '--batch-size', '2048', '--ubatch-size', '512', '--threads', '16', '--threads-batch', '16', '--flash-attn', 'on', '--cache-type-k', 'q8_0', '--cache-type-v', 'q8_0', '--gpu-layers', 'all', '--split-mode', 'none', '--main-gpu', '0', '--fit', 'off', '--load-mode', 'none', '--no-mmproj', '--jinja', '--reasoning', 'auto', '--reasoning-format', 'deepseek', '--reasoning-effort', 'default', '--reasoning-budget', '-1', '--no-reasoning-preserve', '--temp', '1.0', '--top-p', '0.95', '--top-k', '20', '--min-p', '0.0', '--repeat-penalty', '1.0', '--presence-penalty', '0.0', '--spec-type', 'draft-mtp', '--spec-draft-n-max', '2', '--spec-draft-threads', '16', '--spec-draft-threads-batch', '16', '--spec-draft-ngl', 'all', '--spec-draft-type-k', 'q8_0', '--spec-draft-type-v', 'q8_0', '--cache-ram', '0', '--no-cache-idle-slots', '--metrics', '--no-webui', '--cors-origins', '<YOUR_LOOPBACK_ORIGIN>', '--timeout', '7200', '--log-file', '<PATH_TO_SERVER_LOG>', '--log-colors', 'off', '--log-timestamps' ) & .\llama-server.exe @serverArgs 

The Qwen server's reasoning budget is unlimited (-1). The 8192-token value in the default OpenCode entry is a client-side thinking setting. OpenCode advertises 122880 tokens and allows 16384 output tokens; the server allocates 131072 tokens. I use the text-only GGUF and explicitly disable mmproj.

submitted by /u/Sensitive_Song4219
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA