{"repo":"1304674612/agentbench","free":true,"listed":false,"github":"https://github.com/1304674612/agentbench","clone":"git clone https://github.com/1304674612/agentbench.git","description":"The Regression Testing Framework for AI Agents. Replay · Evaluate · Assert · Catch Regressions — in CI. Like Jest for your AI layer.","language":"TypeScript","stars":23,"topics":["agent-testing","ai-agent","ai-testing","devtools","llm-evaluation","nextjs","prompt-engineering","regression-testing","testing-framework","typescript"],"license":null,"category":"ai-agents","readme_excerpt":"AgentBench The Regression Testing Framework for AI Agents Replay · Evaluate · Compare · Assert · Catch Regressions — in CI Quick Start · Methodology · Why · DSL · Ecosystem · Examples · vs Others · Documentation · Releases --- 🚀 Quick Start 30 秒。从安装到看到测试通过。 然后你就可以把测试文件改成你自己的 Agent。 --- 第一个测试：5 行代码 跑一次： 看到结果： --- 然后呢？ 1. 把你的 Agent 代码放到 src/agent.ts 2. 在 tests/ 里写断言 —— 22 种 matcher 任选 3. agentbench test 跑起来 4. 加一个 GitHub Actions workflow，PR 里自动拦住回归 从 0 到 CI 门禁，5 分钟。 --- 🧪 Testing Methodology Testing AI agents is fundamentally different from testing deterministic software. AgentBench is built on a layered strategy: deterministic assertions (tools, tokens, latency) run on every PR; LLM-assisted quality scores (correctness, safety, faithfulness) gate pre-release. Read the Agent Testing Pyramid for the full strategy, and the Anti-Patterns to avoid the most common testing mistakes. --- 📖 Why AgentBench? \"I tracked my time. Coding was 10%. Testing was 90%. Not because I'm slow — because there was no tool.\" — Read the full origin story AI Agents are unpredictable . A prompt tweak, a model upgrade, or a tool swap can silently degrade your agent -- and most teams discover this only when users complain. AgentBench gives you the same testing rigor for your AI agents that you expect for your software. Without AgentBench - \"I think the new prompt is better\" - Manual spot-checking -- misses regressions - No idea if GPT to Claude breaks behavior - Cannot reproduce or bisect failures - cons","default_branch":null,"files":null,"tree":[],"storefront":"/r/1304674612","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/1304674612/agentbench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}