{"repo":"patrick-toulme/harnessgym","free":true,"listed":false,"github":"https://github.com/patrick-toulme/harnessgym","clone":"git clone https://github.com/patrick-toulme/harnessgym.git","description":"Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay fresh sessions with it activated. Works with Codex and Claude Code.","language":"Python","stars":40,"topics":["agents","ai-agents","benchmarking","claude-code","codex","coding-agents","developer-tools","harness","llm","mcp"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"HarnessGym 📖 Documentation: harnessgym.com — the docs are built from docs/ with MkDocs Material and deployed on every push to main . HarnessGym is an open-source framework for iterative agent harness improvement. It runs a coding agent on a hard task, reflects in the same session on which reusable harness artifact would have helped most, builds that single artifact under .harnessgym/ , and starts the next iteration in a fresh session with the accumulated registry context. The package is alpha software for developers evaluating agent workflows. The core package has no third-party runtime dependencies; runner backends shell out to the agent CLI you choose. Codex and Claude Code are supported, and the deterministic fake runner works offline for smoke tests and demos. Install For a source checkout with tests and the bundled examples: Runner prerequisites: - --runner exec uses the codex CLI, configurable with --codex-bin . - --runner claude uses the claude CLI, configurable with --claude-bin . - --runner fake is deterministic and does not require an agent account. - Some examples below require additional local tooling such as NumPy, a C/C++ compiler, PyTorch, Triton, or access to a remote GPU. Quickstart Create or choose a task file in an existing workspace, then run HarnessGym against that workspace: For an offline smoke test from a source checkout, use the bundled numerical debugging demo: CLI Optimization tasks can stop on an objective score instead of only status: solved : Fo","default_branch":null,"files":null,"tree":[],"storefront":"/r/patrick-toulme","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/patrick-toulme/harnessgym/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}