{"repo":"Elfsong/Mercury","free":true,"listed":false,"github":"https://github.com/Elfsong/Mercury","clone":"git clone https://github.com/Elfsong/Mercury.git","description":"Code Efficiency Benchmark","language":"Jupyter Notebook","stars":87,"topics":["benchmark","code-generation","llm","software-engineering"],"license":null,"category":"ai-agents","readme_excerpt":"Mercury: A Code Efficiency Benchmark for Code Large Language Models 🪐 Welcome to Mercury! Mercury is the first code efficiency benchmark designed for LLM code synthesis tasks. It consists of 1,889 programming tasks covering diverse difficulty levels, along with test case generators that produce unlimited cases for comprehensive evaluation. [March 6, 2026] Release Mercury Eval for Mercury Evaluation! [October 8, 2024] Mercury has been accepted to NeurIPS 2024 🌟 [September 20, 2024] We release a way bigger dataset Venus , which supports more languages. It also provides Memory measurement other than Time . [July 10, 2024] We are building Code Arena now for more efficient Code LLMs evaluation! [June 24, 2024] We are currently working on the Multilingual Mercury 🚧 [May 26, 2024] Mercury is now available on BigCode 🌟 Mercury Datasets Access We publish and maintain our datasets at Mercury@HF How to use Mercury Evaluation Benchmark Visualization Citation Questions? Should you have any questions regarding this paper, please feel free to email us (mingzhe@nus.edu.sg). Thank you for your attention!","default_branch":null,"files":null,"tree":[],"storefront":"/r/Elfsong","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Elfsong/Mercury/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}