{"repo":"AmirhosseinHonardoust/Synthetic-Data-Artist","free":true,"listed":false,"github":"https://github.com/AmirhosseinHonardoust/Synthetic-Data-Artist","clone":"git clone https://github.com/AmirhosseinHonardoust/Synthetic-Data-Artist.git","description":"A professional, research-grade comparison of Gaussian Copula and Variational Autoencoder (VAE) methods for synthetic tabular data generation. Includes full evaluation pipeline with distribution overlap, correlation analysis, PCA projections, pairplots, metrics, and automated visual reports.","language":"Python","stars":21,"topics":["synthetic-data","copula","correlation-analysis","data-augmentation","data-privacy","data-science","data-visualization","deep-learning","generative-model","machine-learning"],"license":"MIT","category":"machine-learning","readme_excerpt":"Synthetic Data Artist A professional research-style Python project for generating and evaluating synthetic tabular data . The project compares a Gaussian Copula generator with a lightweight Variational Autoencoder (VAE) and evaluates the generated data using distribution, correlation, categorical similarity, boundary validity, privacy-proxy, and optional downstream machine-learning utility checks. Important: This project is a research and portfolio demo , not a certified privacy-preserving synthetic data product. It can help analyze synthetic data quality, but it does not provide formal differential privacy or guarantee that generated records are safe to release. --- Table of Contents - Project Overview - What This Project Does - What This Project Does Not Do - Features - Methods - Charts and Visual Analysis - How the Evaluation Works - Project Structure - Installation - Running the Generator - Command-Line Usage - Configuration - Generated Outputs - Evaluation - Privacy Proxy Analysis - Testing - Code Quality - Limitations - Responsible Use - Future Improvements - Tech Stack - Author - License --- Project Overview Synthetic data generation is useful when teams want to experiment, prototype, share examples, or test workflows without exposing raw sensitive datasets. However, synthetic data is often misunderstood. Generating fake-looking rows does not automatically make a dataset private, useful, or statistically realistic. This project takes a more careful approach. It does no","default_branch":null,"files":null,"tree":[],"storefront":"/r/AmirhosseinHonardoust","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/AmirhosseinHonardoust/Synthetic-Data-Artist/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}