{"repo":"MiaAI-Lab/Qwen3.6-35B-A3B-NVFP4-vLLM","free":true,"listed":false,"github":"https://github.com/MiaAI-Lab/Qwen3.6-35B-A3B-NVFP4-vLLM","clone":"git clone https://github.com/MiaAI-Lab/Qwen3.6-35B-A3B-NVFP4-vLLM.git","description":"Self-hosted vLLM inference for Qwen3.6-35B-A3B-NVFP4","language":"Jinja","stars":32,"topics":[],"license":null,"category":"self-hosted-apps","readme_excerpt":"Qwen3.6-35B-A3B (NVFP4) — Self-Hosted Inference via vLLM A production-ready vLLM deployment wrapper for Qwen3.6-35B-A3B — an NVIDIA-moixed-expert, NVIDIA-FP4-quantized version of Qwen3.6 with a 35B dense + 3B auxiliary expert architecture. This repo bundles a ready-to-run Docker container, a custom chat template, and start/stop scripts so you can spin up a fully OpenAI-compatible inference server in minutes. --- ✨ Key Features Feature Details --- --- Model Qwen3.6-35B-A3B-NVFP4 — NVFP4 quantised MoE (35B active / 214B total params) Inference Engine vLLM (nightly) with FlashInfer attention + Marlin MoE backend Speculative Decoding MTP (Multi-Token Prediction), 3 speculative tokens Context Window Up to 262 144 tokens (256K) OpenAI-Compatible API /v1/chat/completions , /v1/completions , /v1/models Vision Support Multi-modal image input (up to 4 images per request) Tool Use Qwen3-coder tool-call parser, auto tool choice enabled Thinking/Reasoning CoT / chain-of-thought with block support (configurable) Reasoning Parser Qwen3-specific parser via --reasoning-parser qwen3 Streaming Full SSE streaming support Prefix Caching Enabled via --enable-prefix-caching Chunked Prefill Enabled via --enable-chunked-prefill Async Scheduling Enabled via --async-scheduling Custom Chat Template Full Jinja template with vision, tool use, and thinking support ARM64 Ready Self-contained GCC, Python dev deps, and triton cache --- 📋 Architecture Overview The container exposes an OpenAI-compatible REST A","default_branch":null,"files":null,"tree":[],"storefront":"/r/MiaAI-Lab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/MiaAI-Lab/Qwen3.6-35B-A3B-NVFP4-vLLM/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}