{"repo":"Nexis-AI/NexBench","free":true,"listed":false,"github":"https://github.com/Nexis-AI/NexBench","clone":"git clone https://github.com/Nexis-AI/NexBench.git","description":"NEXBENCH measures what an autonomous agent can actually do on-chain — execute transactions, route swaps, bridge funds, manage DeFi positions, research tokens, catch drainers, reconstruct portfolios, and run treasury governance — across 214 tasks in 8 categories, each run 5 times against deterministic forked-mainnet environments.","language":"TypeScript","stars":99,"topics":["ai","ai-agents","ai-tools","benchmark","benchmark-framework","blockchain","durable-agents","llm","on-chain"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"NEXBENCH The reproducible benchmark for autonomous Web3 agents. Deterministic forked-mainnet environments · programmatic verifiers · tamper-evident submissions. Quickstart · How it works · Build an agent · Submit · Docs · Methodology --- NEXBENCH measures what an autonomous agent can actually do on-chain — execute transactions, route swaps, bridge funds, manage DeFi positions, research tokens, catch drainers, reconstruct portfolios, and run treasury governance — across 214 tasks in 8 categories , each run 5 times against deterministic forked-mainnet environments . Two design choices set it apart: - Programmatic verifiers, not LLM judges. Every task is graded by asserting on post-run chain state, balances, event logs, or gold numeric answers. Determinism in, no self-preference bias, no vibes. - Tamper-evident by construction. Every result is a hash-sealed run manifest. Scores must sit on a mathematically achievable grid, the run id is a content hash that recomputes on intake, and trace archives are Merkle-rooted. Twelve checks run on every submission. See the threat model. Status. Suite v2.1; package/harness v2.1.7. The full 214-task suite runs against the pinned reference environment pack . This repository publishes 24 public task specifications : exactly 6 runnable-local tasks include bundled deterministic environments and verifiers, while 18 metadata-only specs require the reference pack and are never executed by nexbench run . Quickstart Requires Node ≥ 20 . Install from n","default_branch":null,"files":null,"tree":[],"storefront":"/r/Nexis-AI","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Nexis-AI/NexBench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}