Cloud-AI Cold-Turkey: Real Dev Work with Local AI (Ornith 1.5 35b-a3b and Qwen 3.8-27b; 8GB VRAM vs 32GB VRAM)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Spent several weeks on an 'as-much-Local-AI-as-possible' regime and have been - mostly - impressed. Yes, Qwen 3.8 27b is the (rightful) star of the show (obligatory one-shot-mario-build-reddit-comment here!). But don't underestimate what models with more modest hardware requirements can do. With Qwen forsaking us 30b-a3b-enjoyers in their latest releases, I figured I'd share my experiences with Ornith 1.5 35b-a3b - which offers that size and plays well agentically. I picked it over the similarly sized "Qwen 3.6 35b-a3b" MoE model because that Qwen MoE has no thinking levels (beyond 'On' or 'Off') and hence tends to underthink and thus undercook its answers. I had some luck queuing a 'double-check your results' follow-ups with it; but more thinking off-the-bat would be first prize. Ornith (and a few other similar models) addresses this through additional training that leaves if feeling like the equivalent of a High reasoning mode for the Qwen MoE (I don't know how much additional knowledge it's acquired - but the fact that it seems to work harder legitimately improves the results in my experience). HarnessStarted with Pi; found that whilst Qwen 3.8 27b was fine within it, Ornith battled with file-writes/edits frequently failing/retrying. Swapped over to OpenCode (full OpenCode-via-llama model config is below) which resolved this at the expense of a higher base context (about 11k tokens in my own setup - which includes a few optional plugins) I already use OpenCode across the board for all my cloud AI uses (mainly OpenAI/GLM) so I've been happy with this consolidation move overall; swapping to a cloud model mid-chat when things get complicated works well for me; this is where OpenCode excels. Hardware and Performance30tps [100k-140k-context] up to 40tps [<100k-context] TG; 300-400 tps PP. Running on a NVIDIA RTX 5060 Mobile (8GB VRAM); Intel Core Ultra 9 275HX; 32GB DDR5 RAM. This took some tweaking which, again, I'll detail later. And yes, this pales in comparison with the 5090 server setup; but for the more limited hardware, I'm actually happy with the results and find it reasonably snappy in use - with a few OpenCode usage-optimizations (again, will detail later!) First Test: Web OSRan Ornith through a Bijan-Bowen inspired web-OS 'get-a-feel-for-the-model' test: basics worked out the box but required two additional turns to resolve minor bugs (maximize/minimize buttons didn't work; snake-game instant-died upon game-start). It's here: https://jsfiddle.net/db89uwpj/ Not mind-blowing but good enough - and we can see some similarities to Qwen 3.8 27b from Bijan's Qwen video; both models clearly share some of the same lineage : https://youtu.be/6kjXzTVmT58?t=442 Second: Network TroubleshootingExperienced several random router-drop-outs on my home network over the span of several days - since the log contains credentials that I didn't want to clean for online submission (thanks for nothing, Asus!) I dumped it straight into Ornith - fully locally - and within 10 minutes it had identified:
Since the logs are rolled-over about twice a day, it offered to create a monitoring system to autonomously log onto the router every few hours (since telnet is disabled), grab/merge the logs into a single consolidated log per-day, and monitor (and trigger a console alert) for when the scenario recurred. Sounded ambitious - but indeed, it successfully built an app to authenticate via the login page, navigate to system-log page, capture the log-control textbox contents, and snapshot it to disk, merging it with the current day's logs - and it then set a scheduled-job to run it hourly. It worked - one shot. But it didn't need to: I disabled WPS as per its original suggestion and the issue simply hasn't recurred since. Ornith's diagnosis appears to have been spot-on. Promising? Third: Some Real WorkAnd... here's where things got complicated. I'm involved in several large-to-medium-scale software systems that tend to have many sub-systems that interact. The documentation thereof is OK but certainly not exhaustive. Getting either Ornith 1.5 35b-a3b or Qwen 3.8 27b to do solid planning on changes to sections of these systems - especially when they interact with other dependencies - has been a total crapshoot. If the features were constrained to a maximum of 2 or 3 files, both models did surprisingly well with both planning and implementation. If the features went beyond that but were typical 'modify-DAL-then-Business-Layer-then-Frontend' type changes, Qwen usually handled well (with Ornith trailing - doing just OK here). But the minute changes required, say, a method signature change that would have 5 or 6 calls (across as many source files) require updates, both models simply made mistakes that indicated misunderstanding of what the code did; regardless of the system I tried it in. It would compile just fine - but be buggy and often non-functional. So in the case of both models, the only option was to get a larger model (GLM 5.3 or - my preferred planner - GPT 5.6 Sol-High) to do the planning and then use Ornith or Qwen as an implementation model. For production code I also found it was important to manually review the diffs and get Sol-High to review; it mostly over-engineered edge-cases (which I'd ignore) but sometimes, post-implement, it'd plug significant gaps. For simpler changes, Qwen was usually superior but sometimes overcomplicated/overthink'ed (I use Qwen 3.8 27b on its xhigh reasoning level). So occasionally, Ornith produced the better solution. In more than one instance when changes were contained to just a few code files, even Sol found itself impressed with Ornith: Sol-High's Ornith ASP.Net Code Review 'Damned impressive' is not how Sol usually reviews other models' code... good showing from Ornith; I laughed watching Sol do a web-search to verify if Ornith's solution was actually feasible (it's been flawless in prod since!). I even had it correct a few bugs in Astra's code on a test game back when Astra launched (I bench new models on games!), which it again handled well, getting a nod from Astra: Astra 'Noticing' Ornith's Bug-Fix And another Bijen-Bowen inspired test (a 3D Subway) by Ornith: Ornith's Subway-Scene Attempt. It's... meh. ...this one falls well short of Qwen 3.8 27b's output (refer to Bijan's video again: https://youtu.be/6kjXzTVmT58?t=1043 - this is an incredible showing from Qwen) but was nonetheless respectable for the model size. It's not all bad: in production work, Sol's reviews sometimes preferred Ornith's output to Qwen in a few instances: Ornith 1.5 35b-a3b beats out Qwen 3.8 27b. Usually it's the opposite, though. ...but as was often the case in my use, neither was production-ready; and in most cases Qwen won out. In shortHaving shipped a few thousand lines of code from each model, for contained tasks - especially when I have a decent understanding of what needs to change and what basic coding steps to take - I'm honestly happy using Ornith whenever away from my Qwen server. It punches above its weight(s) (and I imagine similar models like Tiel would do equally as well) whilst maintaining acceptable speed on my stand-alone laptop; especially when given decently-detailed prompting as guidance. But even with the dense-model Qwen on standby, I can't work efficiently without having a larger model like GLM 5.3 or GPT-Sol/Astra handling the planning and review. Hence, neither local model is a substitute for my cloud subs in a professional dev setting quite yet; though Qwen 3.8 27b feels a lot like Luna-High on implementation; whilst Ornith feels somewhere between -Medium and -Low. Very, very impressive showings for both, taking into account their respective sizes and hardware compatibility: Qwen 3.8 27b is the most capable local model I've ever run; and Ornith delivers far more than 25% of its capability in just 25% of it's VRAM footprint. Performance TricksOpenCode
General
The Technical Details - OpenCode ConfigI use OpenCode with two llama.cpp servers: Ornith 1.5 35B-A3B on my local 8 GB GPU, and Qwen3.8 27B on a separate RTX 5090 machine. These are the current model entries from my OpenCode config - both servers use llama.cpp b10622. Add both provider entries under Ornith Model download: Ornith 1.5 35B-A3B repaired MTP GGUF - file Qwen Model download: Unsloth Qwen3.8 27B GGUF - file The Technical Details - Llama.cpp ConfigRuntime: llama.cpp b10622 release. Ornith uses the Windows CUDA 12.4 build; Qwen uses the Windows CUDA 13.3 build. The code below is an excerpt of my PowerShell script: Ornith My launcher starts Main K/V cache is Q8_0; the MTP draft K/V cache is Q4_0. Qwen This runs on the RTX 5090 server via The Qwen server's reasoning budget is unlimited ( [link] [comments] |
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.