{"repo":"eburgueno/chat-rk1","free":true,"listed":false,"github":"https://github.com/eburgueno/chat-rk1","clone":"git clone https://github.com/eburgueno/chat-rk1.git","description":"Self-hosted, NPU-accelerated LLM chat on Turing RK1 running Talos Linux","language":"Shell","stars":10,"topics":[],"license":null,"category":"chat-messaging","readme_excerpt":"chat-rk1 — an NPU-accelerated LLM chat UI on a Turing RK1, on Talos Linux AI disclosure: this repository — manifests, scripts, and documentation — was largely written with AI assistance (Claude), directed and reviewed by a human. Everything performance-related was measured on real hardware (the numbers below are from a live cluster, not model output), but read with the same healthy skepticism you'd apply to any homelab writeup. \"I have a Turing RK1 (RK3588) running Talos — can I run an LLM on it, with the NPU actually doing something useful?\" Yes. This repo takes you from that question to a working, browser-based chat UI (Open WebUI + llama.cpp's llama-server ) on your Kubernetes cluster, with the RK3588's NPU accelerating the part it's genuinely good at — prefill , i.e. time-to-first-token. Everything needed is vendored here: the device-tree overlay, the (optional) patched kernel module and Talos extension recipe, the container image build, the Kubernetes manifests, and the rescue kit for when you poke device trees on real hardware. Measured end-to-end on the reference node (RK1 32 GB, Talos v1.13.4, kernel 6.18.34, NPU at 600 MHz, Qwen2.5-3B-Instruct f16, 1940-token prompt, the exact stack these manifests deploy): time to first token tokens/s while streaming --- --- --- CPU only (8 cores) 106.6 s 3.0 with NPU 60.0 s 3.0 The NPU answers 1.8× sooner on a long prompt and streams at the same speed. At the benchmark level ( llama-bench , pp512) the gap is 2.4× (49 vs 20 t/s). Th","default_branch":null,"files":null,"tree":[],"storefront":"/r/eburgueno","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/eburgueno/chat-rk1/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}