{"repo":"aisa-group/InferenceBench","free":true,"listed":false,"github":"https://github.com/aisa-group/InferenceBench","clone":"git clone https://github.com/aisa-group/InferenceBench.git","description":"Benchmarking Open-Ended Inference Optimization by AI Agents","language":"Python","stars":40,"topics":["ai-evals","ai-research-automation","ai-safety","benchmarks","claude-code","codex-cli","sglang","vllm"],"license":"Apache-2.0","category":"machine-learning","readme_excerpt":"InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents This is the official repository for the paper \" InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents \". Authors: Jehyeok Yeon, Ben Rank, Maksym Andriushchenko --- Overview InferenceBench measures whether autonomous CLI agents can act as ML systems engineers in a genuinely open-ended setting. Each run gives the agent a base LLM, a single NVIDIA H100, a wall-clock budget, and a scenario-specific objective; the agent must deliver a running, OpenAI-compatible inference server that maximizes the scenario's primary metric while passing both a quality gate and an integrity gate. Unlike narrower benchmarks where the action space collapses to hyperparameter tuning over a known recipe, inference systems engineering forces real composition choices such as inference framework, attention backend, quantization format, KV-cache layout, scheduler tuning under brittle infrastructure where wrong combinations crash on launch rather than degrading gracefully. The benchmark is designed to test whether agents search an open engineering space or retrieve memorized configurations from it. Headline Result Across 15 frontier agent configurations on Mistral-7B-Instruct-v0.3 with a 2-hour budget per run, agents reliably beat a naïve PyTorch reference and often match or exceed default-configuration serving engines, but non-agent search (Random / SMAC3 / TPE) given the same 2-hour budget on vLL","default_branch":null,"files":null,"tree":[],"storefront":"/r/aisa-group","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/aisa-group/InferenceBench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}