{"repo":"caffe-in/inferlens","free":true,"listed":false,"github":"https://github.com/caffe-in/inferlens","clone":"git clone https://github.com/caffe-in/inferlens.git","description":"A lightweight observability and benchmarking toolkit for self-hosted LLM inference services.","language":"Go","stars":10,"topics":[],"license":"Apache-2.0","category":"self-hosted-apps","readme_excerpt":"inferlens A lightweight observability and debugging CLI for self-hosted LLM inference services. v0.0.2 Goal InferLens ping probes one inference request and prints a readable client timeline plus mode-specific diagnostics. inferlens ping is equivalent to inferlens ping serve and remains focused on the local vLLM serve loop. v0.0.2 also adds ping api for user-provided OpenAI-compatible streaming APIs and ping offline for local vLLM offline inference through the bundled Python helper. Requirements - Go 1.23+ - For ping serve : a vLLM OpenAI-compatible server, usually http://localhost:8000 - For ping offline : a Python environment with vllm installed Quick Start Start vLLM locally: Build the CLI: Run the default serve probe: This is the same as: Ping Modes ping serve Use this for a local or self-hosted vLLM server. It sends one streaming chat completion request and reads vLLM /metrics before and after the probe. If /metrics is unavailable, the probe can still succeed and the report marks server metrics as unavailable. ping api Use this for a user-provided OpenAI-compatible streaming API. It measures client-side streaming behavior only and does not inspect vLLM server metrics. OPENAI API KEY is optional. When it is empty, InferLens sends no Authorization header. API mode requires streaming chat completions in v0.0.2. ping offline Use this for one local vLLM offline inference. InferLens runs scripts/vllm offline probe.py internally and reports model load/generation timing. There is","default_branch":null,"files":null,"tree":[],"storefront":"/r/caffe-in","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/caffe-in/inferlens/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}