r/LocalLLaMA · · 1 min read

We built an open-source GPU profiler you point an AI agent at, instead of reading traces yourself

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

We've been tuning vLLM/SGLang/llama.cpp setups for a long time and got tired of the profiling part: nsys trace, open the GUI, squint, change a flag, repeat. The profilers assume a human is looking at the timeline. These days the thing doing our tuning is usually an agent, and it can't look at a timeline.

So: https://github.com/graphsignal/graphsignal

It's a sidecar profiler. You wrap whatever you're running:

graphsignal-run vllm serve <model> --port 8000
graphsignal-run sglang serve --model-path <model> --port 8000
graphsignal-run ./llama-server -m model.gguf
graphsignal-run python whatever.py

and it serves everything it measures at http://127.0.0.1:18259/signals as one JSON: time per kernel / per CUDA graph / per memcpy and sync, NVML stuff (util, VRAM, power, clocks, throttling, XID errors), the engine's Prometheus metrics if it has them, and any tracebacks it catches in the console output.

Then you give the agent the SKILL.md from the repo and ask for the outcome ("figure out why my GPU sits at 40% during decode and fix it") and it does the loop: run, load, read, change flag, rerun. Also works fine with curl and jq if you'd rather be the agent yourself.

A few details:

- No code changes, no imports. Works with any CUDA or ROCm process, not just the big engines.
- CUDA-graph decode (vLLM, SGLang, llama.cpp) hides all the kernels inside the graph replay. --cuda-graph-trace node breaks them out by kernel name.
- Once a kernel is named and you want to know which part of it, there's a probe header (one file, lock-free) you or the agent drop into the code. Probe values show up in /signals next to everything else.
- Runs locally, binds to 127.0.0.1, uploads nothing unless you give it an API key. No root.

submitted by /u/l0g1cs
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA