Solved: LLM inference on Windows was 2–3x slower when the server window wasn't focused
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
The fix: run the server detached/headless instead of keeping it attached to a console window.
RTX 5090, ~27B NVFP4 model via ninfer:
- Terminal focused: 130–200 tok/s
- Terminal unfocused: 50–60 tok/s
- Click the terminal → immediately back to 130+ tok/s
At first I thought GPU throttling, but the GPU wasn't the problem.
During the slow runs:
- SM clock was actually higher: 2550–2600 MHz vs 2100–2200 MHz
- Power limit was the same: 400 W
- No PCIe power-state drop
decode-hoststayed around 390–494 µs
The big difference was wait:
wait | tok/s | |
|---|---|---|
| Unfocused | 38–41 ms | 50–60 |
| Focused | 16–17 ms | 130–220 |
So the GPU wasn't getting slower. The serving loop was just getting delayed on the CPU side.
Running the server detached (docker run -d / headless) fixed it: wait stays around 17 ms and throughput stays around 130–200 tok/s, regardless of which window is focused.
I also reproduced the same thing with a native Windows build, so this isn't WSL2-specific.
Seems to be some kind of Windows foreground/background CPU scheduling behavior.
Has anyone else seen this with local LLM inference on Windows 11?
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.