{"repo":"Anemll/dspark-vllm-gx10","free":true,"listed":false,"github":"https://github.com/Anemll/dspark-vllm-gx10","clone":"git clone https://github.com/Anemll/dspark-vllm-gx10.git","description":"Two-node DGX Spark/ASUS GX10 DeepSeek V4 Flash DSpark NVFP4 port for vLLM 0.25, with live dashboard and reproducible deployment.","language":"Python","stars":61,"topics":["asus-gx10","dashboard","deepseek-v4","dgx-spark","gb10","nvfp4","vllm"],"license":"MIT","category":"dashboards-admin","readme_excerpt":"DSpark vLLM for two DGX Spark / ASUS GX10 nodes This is a tested two-node GB10 port of the DeepSeek V4 Flash DSpark/NVFP4 serving path to vLLM 0.25.1. It bridges vLLM's DeepSeek V4 runtime to FlashInfer's native SM120/SM121 sparse-MLA kernel, adds a b12x native-MXFP4 MoE backend, and packages reproducible deployment, a live dashboard, version switching, and benchmark evidence. Validated configuration - 2 × NVIDIA DGX Spark or ASUS Ascent GX10 (GB10, SM121, ARM64) - dedicated high-speed fabric between nodes - tensor parallelism: TP=2 - DeepSeek V4 Flash DSpark model using NVFP4 DS MLA KV cache - vLLM source tag v0.25.1 ; runtime reports 0.25.2.dev0+g752a3a504.d20260714 - FlashInfer pinned to 0472b9b3f2fba11b463f8526f390297d52a8aad7 - b12x pinned to 7dc6fb8fcc6446ea093537d1657df81985fa5f43 What this port changes - adds nvfp4 ds mla as a first-class DeepSeek V4 KV-cache format throughout vLLM configuration, quantization, and cache-size accounting; - uses the tested 584-byte packed sparse-MLA token envelope for both MLA and sliding-window cache groups; - adapts vLLM's FlashInfer SM120/SM121 wrapper to split oversized 256-token SWA pages into zero-copy 64-token views while preserving compressed C128 pages; - supports TP=2's 32 query heads and pads unsupported sparse-index widths to FlashInfer's native 128/512/1024 dispatch widths with invalid-slot sentinels; - adds a modular b12x MXFP4 MoE backend with native weight preparation, caller-owned scratch, CUDA-graph-safe execution, GB1","default_branch":null,"files":null,"tree":[],"storefront":"/r/Anemll","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Anemll/dspark-vllm-gx10/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}