{"owner":"swellweb","github":"https://github.com/swellweb","claimed":false,"inventory":[],"indexed":[{"repo":"swellweb/reame","github":"https://github.com/swellweb/reame","description":"CPU-first LLM inference server on llama.cpp. Runs useful models on free-tier ARM boxes; rewriting the input made it ~6x faster and more accurate than tuning the engine. MIT, benchmarks and failures included.","language":"C++","stars":107,"topics":["cpu","gguf","inference","kv-cache","llama-cpp","llm","openai-api","speculative-decoding","local-llm","self-hosted"],"license":"MIT","category":"ai-agents"}],"how_to_buy":"GET /r/swellweb/<repo> (Accept: application/json) for any listed repo here: tree, README, price and the checkout to pay (x402; rehearse first at its test twin, simulated money). Repos under 'indexed' are free: clone them from GitHub."}