{"repo":"smixs/skill-conductor","free":true,"listed":false,"github":"https://github.com/smixs/skill-conductor","clone":"git clone https://github.com/smixs/skill-conductor.git","description":"Architecture-first skill lifecycle for AI agents. BinEval binary scoring with threshold-blind, cross-family-calibrated judges, gated self-update loop, pressure testing, 10 authoring principles grounded in empirical research.","language":"Python","stars":160,"topics":["agent-skills","ai-agents","anthropic","claude-code","llm","productivity","prompt-engineering","skills","tdd","knowledge-management"],"license":"MIT","category":"ai-agents","readme_excerpt":"Skill Conductor A skill that creates, evaluates, and improves other skills. Meta-level. Architecture-first skill lifecycle: design → build → test → evaluate → package . Most skill tools jump straight to \"write SKILL.md.\" Conductor makes you choose the architecture first — because rewriting a wrong pattern costs more than writing it right. Install v3.2.0 — Evidence-based upgrade: form-matching, judge calibration, pressure testing - Principle #10: Match the form to the failure — classify the baseline failure before writing a rule; prohibitions bulletproof discipline failures but measurably backfire on shaping failures (obra/superpowers wording tests + Guardrails polarity data). Plus: no nuance clauses, exemption clauses don't scope. - Critique-before-verdict judges — all three eval agents (grader, comparator, bineval) now write the detailed evidence critique BEFORE committing to the 1/0 verdict, with a borderline few-shot example in each (Hamel Husain's judge methodology). - Threshold-blind judging — the BinEval judge no longer computes the overall score or the GATE; the orchestrator aggregates. A judge that knows the bar is biased toward it. - Automatic cross-family judge calibration — a second judge from a different model family answers the same bank; stable disagreement flags a badly worded question, not a dispute. Self-preference-bias guard on final acceptance. - Variance discipline — improvements on non-critical questions count only when they reproduce in 2 consecutive run","default_branch":null,"files":null,"tree":[],"storefront":"/r/smixs","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/smixs/skill-conductor/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}