{"repo":"corralm/yc-scraper","free":true,"listed":false,"github":"https://github.com/corralm/yc-scraper","clone":"git clone https://github.com/corralm/yc-scraper.git","description":"✌️Y Combinator directory scraper","language":"Python","stars":100,"topics":["python","scrapy","selenium","webscraping","ycombinator","kaggle","dataset"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Y Combinator Directory Scraper A Python scraper for extracting company data from the Y Combinator directory, featuring an interactive web-based explorer. Features - 🚀 User-Friendly Selenium Scraping : No API keys required - just Firefox and geckodriver - 💾 30-Day URL Caching : Avoid unnecessary re-scraping - 🔄 Checkpoint/Resume : Recover from interrupted scrapes - 🎯 Flexible Batch Filtering : Select specific batches or recent N batches - 🌐 Interactive Web Explorer : Browse and filter companies with a sleek UI - 📊 Rich Dataset : Includes founder profiles with bios and social links About Y Combinator Y Combinator is a startup accelerator that has invested in over 4,000 companies with a combined valuation exceeding $600B. Notable alumni include Airbnb, Stripe, DoorDash, Coinbase, and Reddit. Requirements - Python : 3.11 or higher - Browser : Firefox - WebDriver : geckodriver Installing geckodriver macOS : Linux : Windows : Download from GitHub releases and add to PATH. Installation 1. Clone the repository 2. Create a virtual environment (recommended) 3. Install dependencies Usage Quick Start Step 1: Extract Company URLs This script: - Opens the YC directory in headless Firefox - Iterates through all batch filters (Summer 2007 - present) - Collects unique company URLs - Saves to scrapy-project/ycombinator/start urls.txt Step 2: Scrape Company Data Supported output formats : - JSON Lines ( .jl ) - Recommended for large datasets - JSON ( .json ) - Standard JSON array - CSV ( ","default_branch":null,"files":null,"tree":[],"storefront":"/r/corralm","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/corralm/yc-scraper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}