{"repo":"lechmazur/bazaar","free":true,"listed":false,"github":"https://github.com/lechmazur/bazaar","clone":"git clone https://github.com/lechmazur/bazaar.git","description":"The BAZAAR challenges LLMs to navigate the double-auction marketplace, where buyers and sellers must make strategic decisions with incomplete information. Each agent receives a private value and must decide how to quote based solely on the history of previous rounds. A realistic test of market intuition and strategic adaptation.","language":null,"stars":38,"topics":["claude","gemini","grok","llama","llm","lrm","o3","o4-mini","opus","qwen"],"license":null,"category":"ai-agents","readme_excerpt":"BAZAAR - B enchmark for A uction-based Z ero-Intelligence and A daptive A gent R esearch Evaluating LLMs in Economic Decision-Making within a Competitive Simulated Market This benchmark probes fundamental questions about AI economic reasoning: Can LLMs learn bidding strategies through experience? Do they adapt to varying market conditions? How do they balance the tension between aggressive quotes that increase trade probability and conservative quotes that preserve profit margins? By analyzing thousands of games, we reveal the economic instincts of state-of-the-art language models. The BAZAAR ( B enchmark for A uction-based Z ero-Intelligence and A daptive A gent R esearch) challenges Large Language Models to navigate the complexities of a double-auction marketplace, where buyers and sellers must make strategic decisions with incomplete information. Each agent receives a private value and must decide how to quote based solely on the history of previous rounds. No agent can see the current order book or others' private values, creating a realistic test of market intuition and strategic adaptation. ----- Visualizations & Metrics TrueSkill Leaderboard (μ ± σ) A horizontal bar chart ranks each model by the median TrueSkill μ across many passes. Ratings are derived from Conditional Surplus Alpha (CSα), which measures how much better or worse an agent performs than a simple truthful baseline under identical market conditions. CSα normalizes performance so models remain comparable a","default_branch":null,"files":null,"tree":[],"storefront":"/r/lechmazur","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/lechmazur/bazaar/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}