{"repo":"GodModeAI2025/skill-forge","free":true,"listed":false,"github":"https://github.com/GodModeAI2025/skill-forge","clone":"git clone https://github.com/GodModeAI2025/skill-forge.git","description":"Autonomous AI skill improvement through iterative experimentation — inspired by Karpathy's autoresearch. An agent mutates skill instructions, evaluates against objective metrics, keeps improvements, reverts regressions. No human in the loop.","language":"Python","stars":17,"topics":["agents","ai","automation","autonomous-agents","autoresearch","claude","engineering","enterprise-ai","eval","karpathy"],"license":"MIT","category":"ai-agents","readme_excerpt":"Skill Forge v3 Autonomous improvement of AI skills and generic codebases through iterative experimentation. An AI agent modifies instructions or code, evaluates each change against objective metrics, keeps improvements, and reverts regressions. Runs fully autonomous or in guided mode where the user decides at every step. What it does Skill Forge runs an experiment loop in two modes: Skill Mode : optimizes a Claude Cowork Skill's SKILL.md against eval assertions: Generic Mode : optimizes any file against any shell command that returns a number: You point it at a skill or codebase, it finds weaknesses, fixes them, and delivers an improved version with a full experiment log. Run it overnight, wake up to a better skill. What v3 changed v2 described a loop. v3 makes the loop's decisions mean something. The changes fall into four groups. The decision is code, not prose - One decision function. decide() and the decide subcommand are the only place where two scores turn into a verdict. Three outcomes, no fourth. In v2 the cascade lived as pseudocode in SKILL.md, and its NEUTRAL branch was unreachable: over 4001 deltas it fired exactly once. - NEUTRAL reverts. A tie rolls the mutation back. A loop that keeps the new version on a zero round drifts away from its baseline without a single measurement to justify it. - near miss is a flag on NEUTRAL , not a separate outcome. It marks the band just below the keep threshold, so the next round can vary the same hypothesis instead of dropping ","default_branch":null,"files":null,"tree":[],"storefront":"/r/GodModeAI2025","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/GodModeAI2025/skill-forge/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}