{"repo":"Sandermage/sndr_core_engine","free":true,"listed":false,"github":"https://github.com/Sandermage/sndr_core_engine","clone":"git clone https://github.com/Sandermage/sndr_core_engine.git","description":"SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.","language":"Python","stars":132,"topics":["cuda","llm-inference","qwen","vllm","speculative-decoding","tool-calling","rtx-3090","awq","consumer-gpu","gemma"],"license":"Apache-2.0","category":"machine-learning","readme_excerpt":"SNDR Core Engine Genesis vLLM Patches — runtime vLLM patches that run frontier-class open LLMs — Qwen3.6 (7B · 27B · 35B-A3B) and Gemma 4 (26B · 31B · DiffusionGemma) — on consumer NVIDIA GPUs with 24 GB (RTX 3090 / 4090 / 5090, RTX A5000 / A6000): 1.5× faster inference, quantized tool calling that works, MTP speculative decoding , and up to 280K-token context — no fork, no rebuild. Built on vLLM, so it serves the models the engine supports; the deep optimizations (TurboQuant KV, hybrid GDN, spec-decode, tuned kernels) are family-tuned for Qwen3.6 and Gemma 4 — the two families we validate on every pin. 🎮 Own a different card? The 24 GiB envelope is class-wide, and sndr up auto-projects VRAM for your GPU. RTX 4090 · RTX 5090 (32 GiB) · dual RTX 3090 — honest per-class gotchas (Ampere-calibrated tuning, no-NVLink P2P, idle-VRAM headroom) in the FAQ. Contents: Get running · Who is this for · Why SNDR Core · How it compares · What it is · How it works · The platform end-to-end · Headline numbers · Fleet validation · Persistent memory · Pick your path · Install & run · FAQ · Documentation map · Repository structure · Contributing Turn a consumer NVIDIA card into a production local-AI server. SNDR Core transforms the open-source vLLM engine in memory at boot — no fork, no rebuild — so frontier-class open models (Qwen3.6 up to 35B-A3B , Gemma 4 up to 31B ) run 1.5× faster than stock vLLM with up to a 280K-token served context , on hardware you can actually buy (A5000, RTX 4090 / 5","default_branch":null,"files":null,"tree":[],"storefront":"/r/Sandermage","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Sandermage/sndr_core_engine/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}