{"repo":"UrbsLab/STREAMLINE","free":true,"listed":false,"github":"https://github.com/UrbsLab/STREAMLINE","clone":"git clone https://github.com/UrbsLab/STREAMLINE.git","description":"Simple Transparent End-To-End Automated Machine Learning Pipeline for Supervised Learning in Tabular Binary Classification Data","language":"Jupyter Notebook","stars":82,"topics":["automl-pipeline","binary-classification","data-science","data-visualization","feature-selection","imputation","machine-learning","model-application","statistical-analysis","supervised-learning"],"license":"GPL-3.0","category":"machine-learning","readme_excerpt":"STREAMLINE STREAMLINE is a parallelizable, customizable, end-to-end automated machine learning pipeline for supervised learning in structured, tabular data. It supports binary classification, multiclass classification, and regression outcomes. It includes exploratory analysis, data processing, imputation, scaling, feature learning, feature importance estimation, feature selection, multiple algorithm model training with hyperparameter optimization, ensemble modeling, comprehensive model evaluation with statistical comparisons, dataset comparison, easy replication, and PDF summary reporting. STREAMLINE is designed to make rigorous machine learning workflows easier to run, compare, reproduce, and extend. It can be run through Google Colab, a local Jupyter notebook, command-line phase runners, or configuration files that execute the full pipeline. Quick Links - Detailed documentation - Google Colab demo notebook - STREAMLINE Notebook.ipynb for local Jupyter runs - run configs/ for config-driven demo pipelines - Citation information Workflow Schematic Pipeline Overview The pipeline is organized into 11 phases: Phase Name Purpose --- --- --- P1 Data Process Load datasets, define CV partitions, run exploratory analysis, and prepare dataset-specific metadata P2 Impute and Scale Apply preprocessing, imputation, and scaling P3 Feature Learning Learn transformed features such as PCA-derived components and write feature-learning artifacts P4 Feature Importance Score features with filter-","default_branch":null,"files":null,"tree":[],"storefront":"/r/UrbsLab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/UrbsLab/STREAMLINE/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}