{"repo":"TaewoooPark/Agent-Blackbox","free":true,"listed":false,"github":"https://github.com/TaewoooPark/Agent-Blackbox","clone":"git clone https://github.com/TaewoooPark/Agent-Blackbox.git","description":"Local-first flight recorder for coding agents : replay every run as a live session map, score the context bill, and write the fix back into AGENTS.md — no API key, one npx command.","language":"TypeScript","stars":72,"topics":["agent","agent-observability","agents","ai","ai-agents","claude","coding-agent","context-engineering","dashboard","developer-tools"],"license":"MIT","category":"ai-agents","readme_excerpt":"Agent-Blackbox Open your coding agent's black box. English · 한국어 · 中文 · 日本語 &nbsp; &nbsp; Agent-Blackbox is a local-first flight recorder and context-efficiency profiler for coding agents. It turns every agent run into a live, replayable operational graph — what the agent read, changed, ran, decided, delegated, blocked on, and verified — reconstructed from observed events, not from the agent's own summary. Then it scores the run on two axes — how economically it used its context window, and whether the task actually landed — judged on a yardstick that fits the task type (research / debug / ops…) and your own past runs , and tells you, concretely, how to make the next one cheaper and faster. Works with Claude Code, Codex, and OpenCode — same recorder, same map, same efficiency score. Record one host or all of them at once. \"The transcript is what the agent said. The black box is what it did — and what it cost.\" taewoopark.com — author site --- Why Agent-Blackbox You can't just ask the agent what a task cost. A 2026 study of eight frontier models on agentic coding (SWE-bench Verified) found they predict their own token usage with a correlation of just 0.39 — and systematically underestimate the real bill. Same task, same model: runs vary up to 30× in tokens. Expert difficulty ratings barely track real cost. And agentic runs already burn 1000× more tokens than ordinary coding, almost all of it input context. So don't ask — measure. Agent-Blackbox replays every run as an observed","default_branch":null,"files":null,"tree":[],"storefront":"/r/TaewoooPark","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/TaewoooPark/Agent-Blackbox/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}