{"repo":"MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM","free":true,"listed":false,"github":"https://github.com/MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM","clone":"git clone https://github.com/MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM.git","description":"Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference","language":"Jinja","stars":60,"topics":[],"license":null,"category":"machine-learning","readme_excerpt":"PLS USE THE NEW REPO - https://github.com/MiaAI-Lab/Qwen3.6-27B-NVFP4-DFlash-DGX-Spark Qwen3.6-27B (NVFP4) DFlash for DGX Spark A production-ready vLLM deployment wrapper for Qwen3.6-27B — an NVFP4-quantized dense model ( 27B params) with DFlash speculative decoding . This repo bundles a ready-to-run Docker container, a custom chat template, and start/stop scripts so you can spin up a fully OpenAI-compatible inference server in minutes. --- ✨ Key Features Feature Details --- --- Model Qwen3.6-27B-NVFP4 — NVFP4 quantised dense ( 27B params) Vision Supports image input (multimodal) Quantization ModelOpt ( --quantization modelopt ) Inference Engine vLLM (nightly aarch64) with Flash Attention backend Speculative Decoding DFlash, draft model z-lab/Qwen3.6-27B-DFlash , 10 speculative tokens Context Window Up to 262 144 tokens (256K) OpenAI-Compatible API /v1/chat/completions , /v1/completions , /v1/models Served Model Name qwen36-27b-nvidia-nvfp4-dflash Tool Use Qwen3-coder tool-call parser, auto tool choice enabled Thinking/Reasoning CoT / chain-of-thought with block support (configurable) Reasoning Parser Qwen3-specific parser via --reasoning-parser qwen3 Streaming Full SSE streaming support Prefix Caching Enabled via --enable-prefix-caching Chunked Prefill Enabled via --enable-chunked-prefill KV Cache bfloat16 ( --kv-cache-dtype bfloat16 ) Custom Chat Template Full Jinja template with tool use and thinking support ARM64 Native vLLM nightly aarch64 image --- 📊 Performance Decode","default_branch":null,"files":null,"tree":[],"storefront":"/r/MiaAI-Lab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/MiaAI-Lab/Qwen3.6-27B-NVFP4-vLLM/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}