{"repo":"hec-ovi/vllm-qwen","free":true,"listed":false,"github":"https://github.com/hec-ovi/vllm-qwen","clone":"git clone https://github.com/hec-ovi/vllm-qwen.git","description":"vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.","language":"Python","stars":17,"topics":["amd","docker","gfx1151","llm-serving","openai-compatible","qwen","qwen3","rocm","ryzen-ai","strix-halo"],"license":"Unlicense","category":"deployment-docker-iac","readme_excerpt":"vllm-qwen Qwen3.6-27B (BF16) served over OpenAI-compatible HTTP on AMD Strix Halo. --- What this is A thin Docker Compose wrapper that runs Qwen/Qwen3.6-27B (BF16) behind an OpenAI-compatible HTTP API on an AMD Ryzen AI Max+ 395 \"Strix Halo\" ( gfx1151 , RDNA 3.5, 128 GB UMA). Serves /v1/completions , /v1/chat/completions , /v1/responses , and vision inputs through the same endpoint. Native 256K context. vLLM is built from source against a TheRock nightly ROCm SDK with a small patch set for Strix Halo. There is no prebuilt image path, consumer AMD GPUs aren't in AMD's mainstream ROCm support matrix yet, so source is the only clean route. Want 75% faster decode and working tool calls instead? See the sibling repo llama-qwen : same hardware, Qwen 3.6-27B Q8 0 via llama.cpp. Decode 7.5 t/s vs this repo's 4.3 t/s, tool calling verified clean (vLLM currently has three open upstream parser bugs). Trade-off: no vision, no /v1/responses . Pick that one for agentic / coding / chat speed, this one for vision or structured reasoning output. --- Stack Layer Version --- --- Host OS Ubuntu 26.04 (container base) ROCm TheRock 7.13.0a20260424 (S3 nightly; resolves to latest at build time) PyTorch 2.10.0+rocm7.12.0rc1 (AMD gfx1151 prerelease wheels) Triton 3.6.0+rocm7.12.0rc1 vLLM 0.19.2rc1 upstream HEAD, built from source, 12 local patches Model Qwen/Qwen3.6-27B (BF16, official) --- Hardware Tested on: Ryzen AI Max+ 395 / 128 GB UMA (Radeon 8060S iGPU, gfx1151 ). Kernel ≥ 6.18. Docker with /d","default_branch":null,"files":null,"tree":[],"storefront":"/r/hec-ovi","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/hec-ovi/vllm-qwen/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}