{"repo":"InternScience/ResearchClawBench","free":true,"listed":false,"github":"https://github.com/InternScience/ResearchClawBench","clone":"git clone https://github.com/InternScience/ResearchClawBench.git","description":"🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery","language":"Jupyter Notebook","stars":242,"topics":["auto-research","openclaw","agent","ai4science","benchmark","claude","claude-code","codex","discovery","evaluation"],"license":"MIT","category":"ai-agents","readme_excerpt":"ResearchClawBench &#160; &#160; &#160; &#160; &#160; Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery Quick Start Submit Tasks How It Works Domains Leaderboard Add Your Agent --- ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers . Unlike benchmarks that test coding ability or factual recall, ResearchClawBench asks: given a curated scientific workspace and the same research goal, can an AI agent arrive at the same (or better) scientific conclusions? Overview ✨ Highlights 🔄 Two-Stage Pipeline Autonomous research + rigorous peer-review-style evaluation 🧪 40 Real-Science Tasks 10 disciplines, curated datasets from published papers 👁️ Expert-Annotated Data Tasks, rubrics (checklists) & datasets curated by domain experts 🤖 Multi-Agent Support Claude Code, Codex CLI, OpenClaw, ResearchClaw, ... & custom agents 🚀 Re-Discovery to New-Discovery 50 = match the paper, 70+ = surpass it 📋 Fine-Grained Rubric (Checklist) Per-item keywords, weights & reasoning 📡 Live Streaming UI Watch agents code, plot & write in real-time 🍃 Lightweight Dependencies Pure Flask + vanilla JS, no heavy frameworks 🎬 Demo https://github.com/user-attachments/assets/94829265-80a8-4d61-a744-3800603de6d9 💡 Why ResearchClawBench? Most AI benchmarks evaluate what models ","default_branch":null,"files":null,"tree":[],"storefront":"/r/InternScience","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/InternScience/ResearchClawBench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}