{"repo":"AmirhosseinHonardoust/Autocurator-Synthetic-Data-Benchmark","free":true,"listed":false,"github":"https://github.com/AmirhosseinHonardoust/Autocurator-Synthetic-Data-Benchmark","clone":"git clone https://github.com/AmirhosseinHonardoust/Autocurator-Synthetic-Data-Benchmark.git","description":"Autocurator is a comprehensive benchmarking toolkit for evaluating synthetic tabular data. It measures fidelity, coverage, privacy, and utility through quantitative metrics, visual reports, and PCA/correlation diagnostics. Ideal for validating VAE, GAN, Copula, or Diffusion-generated datasets.","language":"Python","stars":20,"topics":["benchmark","copula","coverage","data-privacy","data-quality","data-science","data-validation","deep-learning","diffusion","evaluation"],"license":"MIT","category":"machine-learning","readme_excerpt":"Autocurator — Synthetic Data A modular evaluation toolkit that benchmarks synthetic tabular data against real data across four axes — fidelity, coverage, privacy, and utility — with reference-validated metrics , a generator sensitivity benchmark , HTML reporting , YAML-driven configuration , and a reproducible, multi-version CI pipeline . Important: Autocurator is an evaluation and diagnostic tool , not a certification of privacy or safety. Its metrics estimate how closely synthetic data resembles real data and how much it may leak ; they are decision-support signals, not guarantees. A low membership-inference score is not a legal privacy guarantee, and high utility does not certify fitness for any downstream use. Read every number alongside the limitations below. --- Table of Contents - Project Overview - What This Project Does - What This Project Does Not Do - Key Features - System Workflow - Project Structure - Installation - Quick Start - Running a Benchmark - Metrics Explained - Metric Output - Evaluation Summary - Validation and Sensitivity - Visual Reports - Testing and CI - Code Quality - Reproducibility - Limitations - Responsible Use - Future Improvements - Tech Stack - Author - License --- Project Overview Synthetic data is often presented as if \"looks realistic\" and \"is safe to share\" were the same thing. In reality they pull in different directions: data that perfectly reproduces the original is high-fidelity but leaks privacy, while data that is aggressively pri","default_branch":null,"files":null,"tree":[],"storefront":"/r/AmirhosseinHonardoust","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AmirhosseinHonardoust/Autocurator-Synthetic-Data-Benchmark/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}