{"repo":"konjoai/squish","free":true,"listed":false,"github":"https://github.com/konjoai/squish","clone":"git clone https://github.com/konjoai/squish.git","description":"⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-response time. OpenAI/Ollama-compatible. No cloud, no API keys.","language":"Python","stars":18,"topics":["apple-silicon","llm","local-ai","local-llm","macos","mlx","ollama","openai-api","quantization","qwen"],"license":null,"category":"ai-agents","readme_excerpt":"Squish The fastest way to run local LLMs on Apple Silicon. Sub-second model loads. Beats Ollama on throughput, tail latency, and full-response time. One OpenAI/Ollama-compatible daemon — no cloud, no API keys, fully offline. Read the deep dive on the engineering behind Squish's speed: Local LLMs Are Finally Fast Enough. --- Squish separates how a model's weights are stored from how they run . Store them compressed and Metal-native; map them straight into unified memory; skip the dtype-conversion pass that makes every other loader slow. The result: a model that's ready in half a second , served by a persistent daemon that out-decodes Ollama and never re-does work it's already done. --- The Numbers Measured on an Apple M3 MacBook Pro, 16 GB — thermally controlled (each engine measured from the same 50 °C baseline; validated by a first-vs-last drift check ≤ 1.7 % and live die-temperature logging, so the numbers reflect the engines, not the order they ran). Serving: Qwen2.5-7B-Instruct , Squish INT4/INT3 vs Ollama qwen2.5:7b (Q4 K M), against both Ollama 0.18.2 and 0.30.7 (0.30.7 shown; 0.18.2 within noise). Metric Ollama Squish --- ---: ---: Cold start — load + first token (1.5B) 20–30 s ≈ 0.5 s &nbsp; (54× load) Full response @ 4000-token prompt (repeated exactly)\\ 37.5 s 3.8 s &nbsp; (9.8× faster) Decode throughput @ 75 tokens 20.3 tok/s 24.0 tok/s &nbsp; (INT3) Inter-token tail (p95) @ 75 tokens 52.4 ms 42.7 ms &nbsp; (INT3) Repeat-prompt TTFT (KV cache hit) 160 ms 4–11 ms Pe","default_branch":null,"files":null,"tree":[],"storefront":"/r/konjoai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/konjoai/squish/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}