r/LocalLLaMA · · 1 min read

Solved: LLM inference on Windows was 2–3x slower when the server window wasn't focused

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

The fix: run the server detached/headless instead of keeping it attached to a console window.

RTX 5090, ~27B NVFP4 model via ninfer:

  • Terminal focused: 130–200 tok/s
  • Terminal unfocused: 50–60 tok/s
  • Click the terminal → immediately back to 130+ tok/s

At first I thought GPU throttling, but the GPU wasn't the problem.

During the slow runs:

  • SM clock was actually higher: 2550–2600 MHz vs 2100–2200 MHz
  • Power limit was the same: 400 W
  • No PCIe power-state drop
  • decode-host stayed around 390–494 µs

The big difference was wait:

wait tok/s
Unfocused 38–41 ms 50–60
Focused 16–17 ms 130–220

So the GPU wasn't getting slower. The serving loop was just getting delayed on the CPU side.

Running the server detached (docker run -d / headless) fixed it: wait stays around 17 ms and throughput stays around 130–200 tok/s, regardless of which window is focused.

I also reproduced the same thing with a native Windows build, so this isn't WSL2-specific.

Seems to be some kind of Windows foreground/background CPU scheduling behavior.

Has anyone else seen this with local LLM inference on Windows 11?

submitted by /u/koloved
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA