{"repo":"mrdushidush/claudette","free":true,"listed":false,"github":"https://github.com/mrdushidush/claudette","clone":"git clone https://github.com/mrdushidush/claudette.git","description":"Air-gapped AI coding agent in one Rust binary + Q56, a hidden-test benchmark for local coding models: 16 configs, 36 full runs, no LLM judge. Runs entirely on your own hardware through Ollama or LM Studio. No cloud brain. No API key. No telemetry.","language":"Rust","stars":18,"topics":["cli","llm","local-first","ollama","rust","ai-agent","local-llm","personal-assistant","telegram-bot","coding-agent"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"Claudette An air-gapped AI coding agent in one Rust binary — run it --offline and your code physically cannot leave the machine. It drives a model you run locally through Ollama or LM Studio; there is no cloud-brain code in the binary at all. It also ships Q56 — a hidden-test benchmark that measures which local model is actually worth running, with no LLM judge anywhere in the loop. --- 📊 Start here: Q56, a hidden-test benchmark for local coding models Before the agent — the measurement. Claudette ships with a 56-task coding benchmark whose grading a model cannot talk its way through: the fixture it sees carries only happy-path tests, and at grade time a verifier injects hidden reviewer tests and builds them against whatever the model actually left on disk. Pass/fail is cargo test , pytest and node . There is no LLM judge — LLM judges inflate. 16 model configurations · 36 full runs · one RTX 5060 Ti 16 GB · every held-constant recorded, not assumed. Three results that are worth your time even if you never install this: - A 7.5B model beat a 24B. gemma-4-e4b (7.5B, 4.97 GiB) scored 42/56 , above devstral-small-2-24b (40) and gpt-oss-20b (38). Three runs each, non-overlapping ranges. Size buys nothing between 7.5B and 12B — and that plateau has hard cliffs on both sides. - The error bar belongs to the model, not the benchmark. Across three identical consecutive runs, gpt-oss-20b swung 8 points (40, 38, 32) while gemma-4-e2b was bit-for-bit identical (31, 31, 31). Weakness does","default_branch":null,"files":null,"tree":[],"storefront":"/r/mrdushidush","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/mrdushidush/claudette/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}