{"repo":"ccmdi/geobench","free":true,"listed":false,"github":"https://github.com/ccmdi/geobench","clone":"git clone https://github.com/ccmdi/geobench.git","description":"GeoGuessr benchmark for language models","language":"Python","stars":63,"topics":["benchmark","chatgpt","claude","gemini","geoguessr","llama","llm"],"license":"MIT","category":"ai-agents","readme_excerpt":"GeoBench is a benchmark for evaluating how well large language models can geolocate images, through the context of GeoGuessr. This project tests whether models can generalize beyond their primary training modalities to perform spatial reasoning tasks. Leaderboard For an in-depth explanation of the results, covering things like model behavior and reasoning, see my writeup. Installation Setup your .env based on SAMPLE.env for whichever model providers you wish to test for (e.g. ANTHROPIC API KEY must be set to test Claude). Instructions for setting up NCFA can be found here. Create a dataset Test a model Models go by their class name in models.py . Claude 3.5 Haiku goes by Claude3 5Haiku , for instance. Compare guesses Running the browser/main.py script and opening visualization.html can show you all guesses for a location made by the models.","default_branch":null,"files":null,"tree":[],"storefront":"/r/ccmdi","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/ccmdi/geobench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}