{"repo":"bryanvine/vllmstat","free":true,"listed":false,"github":"https://github.com/bryanvine/vllmstat","clone":"git clone https://github.com/bryanvine/vllmstat.git","description":"nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)","language":"Python","stars":11,"topics":["gpu","llm","monitoring","nvtop","observability","textual","tui","vllm"],"license":"Apache-2.0","category":"analytics","readme_excerpt":"vllmstat nvtop for vLLM — a zero-infrastructure interactive terminal dashboard for vLLM serving performance. --- Why vllmstat? The standard observability stack for vLLM is Prometheus + Grafana: powerful, but heavyweight. You need a running Prometheus instance, a Grafana server, a dashboard JSON import, and a browser tab — all just to see whether your inference server is busy. vllmstat replaces that for day-to-day monitoring. One command, no infrastructure. It scrapes the vLLM server's built-in /metrics endpoint directly and renders everything in your terminal, refreshing every second. There is one other terminal tool ( vllm-top on PyPI), but it is a basic watch -style metrics printer: no interactivity, no GPU panel, no latency percentiles, no speculative-decoding acceptance, no KV-compression ratio. vllmstat fills that gap — it is closer to nvtop than to watch . --- Install Or with pipx (isolated install, globally available): Or run it ephemerally without installing: --- Usage Point it at your vLLM server and it starts immediately: Key bindings Key Action ----- -------- q Quit p Pause / resume polling g Toggle GPU panel / column on/off r Reset the SESSION averages (of the selected instance) t Toggle the TEE request-feed panel ↑ / ↓ (or k / j ) Fleet overview: move the selection Enter Fleet overview: open the selected instance's dashboard Esc Drill-in: return to the fleet overview + / = Halve the refresh interval (faster) - Double the refresh interval (slower) Flags Flag Defau","default_branch":null,"files":null,"tree":[],"storefront":"/r/bryanvine","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/bryanvine/vllmstat/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}