{"repo":"lexmount/browseruse-agent-bench","free":true,"listed":false,"github":"https://github.com/lexmount/browseruse-agent-bench","clone":"git clone https://github.com/lexmount/browseruse-agent-bench.git","description":"Real-world browser-agent benchmark: 210 tasks across 107 websites, multi-agent/multi-browser evaluation, reproducible leaderboard and result submissions.","language":"Python","stars":19,"topics":["agent","benchmark","browseruse","agent-evaluation","ai-agents","browser-agent","browser-automation","computer-use","leaderboard","llm-evaluation"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"Landing Page • Issues • Discussions • Leaderboard • Documentation • Dataset English 简体中文 Why browseruse-agent-bench browseruse-agent-bench is a reproducible evaluation framework for browser agents. LexBench-Browser is the built-in public dataset used by the default benchmark workflow. Together they make external results easy to run, compare, cite, and submit back. What you can do Why it matters ----------------- ---------------- Run LexBench-Browser: 210 public tasks across 107 real websites Test browser agents on long-tail multilingual workflows beyond toy pages Compare Agent × Model × Browser × Eval Separate agent quality from model choice, browser backend, and judge strategy Inspect leaderboard, cost, latency, token usage, and trajectories Debug failures instead of only reporting a final score Submit agents, dataset tasks, and reproducible results Turn forks and PRs into visible benchmark contributions Description browseruse-agent-bench is an all-in-one evaluation framework for AI browser agents, designed to benchmark multiple agents across multiple datasets, browser backends, and models under controlled and reproducible settings. The Python package/CLI is published as browseruse-bench and bubench . It supports both local and cloud browsers, integrates LLM-as-Judge for automated evaluation, and provides a built-in local leaderboard along with efficiency and cost metrics such as agent steps, end-to-end latency, and token usage. Supported Datasets - [x] LexBench-Browser — Br","default_branch":null,"files":null,"tree":[],"storefront":"/r/lexmount","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/lexmount/browseruse-agent-bench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}