r/LocalLLaMA · · 1 min read

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings:

  • explicit effort level on every role (unset ones defaulted to xhigh)

  • thinking_token_budget of 7500

  • maxTokens 8k → 32k (file writes were getting cut off)

  • tool output over 10 KB goes to a file

  • max 4 subagents, appendOnlyContext on

More info in the post, happy to answer any questions.

submitted by /u/bolts98
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA