{"repo":"Peter-awe/KiwiFARS","free":true,"listed":false,"github":"https://github.com/Peter-awe/KiwiFARS","clone":"git clone https://github.com/Peter-awe/KiwiFARS.git","description":"Self-hosted LLM research infrastructure on HPC — vLLM serving, RAG literature search, and a SLURM-orchestrated 9-stage pipeline","language":"Python","stars":14,"topics":[],"license":null,"category":"self-hosted-apps","readme_excerpt":"KiwiFARS Self-hosted LLM inference and experiment-orchestration infrastructure on Narval HPC (SLURM). What This Is - Private LLM inference (Qwen3.5-35B + DeepSeek-R1-32B) via vLLM on A100-40GB - 9-stage experiment pipeline : hypothesis tracking, SLURM-orchestrated runs, validation, and automated run reports with consistency checks - RAG literature database : PDF ingestion + FAISS vector search + arXiv fetching - Daily paper briefing : Automated arXiv digest with LLM summaries via email Quick Start Directory Layout GPU Topologies Mode Main Model Judge Model Free GPUs Use Case ------ ----------- ------------- :---------: ---------- 2-GPU FP8 × 1 (32K ctx) FP8 × 1 (16K ctx) 6 Daily mode 3-GPU BF16 TP=2 (65K ctx) FP8 × 1 (16K ctx) 5 Long-context tasks 4-GPU BF16 TP=2 (65K ctx) BF16 TP=2 (32K ctx) 4 Multi-model evaluation Important Notes - A100-40GB (not 80GB) — all VRAM calculations are for 40GB cards - Compute nodes have no internet — all downloads happen on login node via setup once.sh - Completely isolated from other project directories on the cluster - SLURM account: your-account gpu","default_branch":null,"files":null,"tree":[],"storefront":"/r/Peter-awe","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Peter-awe/KiwiFARS/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}