{"repo":"QuesmaOrg/otel-bench","free":true,"listed":false,"github":"https://github.com/QuesmaOrg/otel-bench","clone":"git clone https://github.com/QuesmaOrg/otel-bench.git","description":"OpenTelemetry Benchmark - can AI trace your failed login?","language":"Shell","stars":20,"topics":["ai-agents","benchmark","opentelemetry"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"OpenTelemetry Benchmark (OTelBench) by Quesma An open-source benchmark for evaluating AI models on OpenTelemetry instrumentation tasks across multiple programming languages. Benchmark: OTelBench results Blog post: Benchmarking OpenTelemetry: Can AI trace your failed login? Quick start Requires Harbor ( uv tool install harbor ), Docker, and relevant API KEYs. By default, we use the terminus-2 agent (default for Harbor) via OpenRouter to compare models. You are free to use others, including well-known CLI AI Agents like Claude Code, Codex, or Cursor CLI. You need to clone this repo: Run a single task, for a single model: Task names allow wildcards, so if you want to run all Go tasks, it works like: Run all tasks with a few models, with 3 attempts per model-task combination: You can view trajectories (interactions between the agent and the system) via harbor view jobs . Our overview of Harbor in Migrating CompileBench to Harbor: standardizing AI agent evals. Content The OpenTelemetry dataset datasets/otel contains a set of tasks testing AI models' ability to instrument applications with OpenTelemetry across 11 programming languages. So far, it contains the following tasks: C++ : simple, advanced, distributed-context-propagation Go : http-tracing, distributed-context-propagation, workflow-tracing, microservices, grpc-fix, microservices-logs, microservices-traces, microservices-traces-simple Java : simple, advanced, distributed-context-propagation, microservices JavaScript : micro","default_branch":null,"files":null,"tree":[],"storefront":"/r/QuesmaOrg","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/QuesmaOrg/otel-bench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}