We built an open-source GPU profiler you point an AI agent at, instead of reading traces yourself
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
We've been tuning vLLM/SGLang/llama.cpp setups for a long time and got tired of the profiling part: nsys trace, open the GUI, squint, change a flag, repeat. The profilers assume a human is looking at the timeline. These days the thing doing our tuning is usually an agent, and it can't look at a timeline.
So: https://github.com/graphsignal/graphsignal
It's a sidecar profiler. You wrap whatever you're running:
graphsignal-run vllm serve <model> --port 8000
graphsignal-run sglang serve --model-path <model> --port 8000
graphsignal-run ./llama-server -m model.gguf
graphsignal-run python whatever.py
and it serves everything it measures at http://127.0.0.1:18259/signals as one JSON: time per kernel / per CUDA graph / per memcpy and sync, NVML stuff (util, VRAM, power, clocks, throttling, XID errors), the engine's Prometheus metrics if it has them, and any tracebacks it catches in the console output.
Then you give the agent the SKILL.md from the repo and ask for the outcome ("figure out why my GPU sits at 40% during decode and fix it") and it does the loop: run, load, read, change flag, rerun. Also works fine with curl and jq if you'd rather be the agent yourself.
A few details:
- No code changes, no imports. Works with any CUDA or ROCm process, not just the big engines.
- CUDA-graph decode (vLLM, SGLang, llama.cpp) hides all the kernels inside the graph replay. --cuda-graph-trace node breaks them out by kernel name.
- Once a kernel is named and you want to know which part of it, there's a probe header (one file, lock-free) you or the agent drop into the code. Probe values show up in /signals next to everything else.
- Runs locally, binds to 127.0.0.1, uploads nothing unless you give it an API key. No root.
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.