{"repo":"Sara12-2/High_Cost_Patient_prediction_Softec_Competition_Project","free":true,"listed":false,"github":"https://github.com/Sara12-2/High_Cost_Patient_prediction_Softec_Competition_Project","clone":"git clone https://github.com/Sara12-2/High_Cost_Patient_prediction_Softec_Competition_Project.git","description":"This project predicts high-cost healthcare patients using historical medical claims, utilization, and demographic data for a binary classification task. It uses advanced feature engineering and ensemble models to identify high-risk members with optimized F1-score for competition performance.","language":"Jupyter Notebook","stars":10,"topics":["binary-classification","feature-engineering","lightgbm","xgboost","data-cleaning","evaluation","metrics"],"license":null,"category":"analytics","readme_excerpt":"🏥 High-Cost Patient Prediction — Softec 2026 ML Competition Predicting which health-plan members will become high-cost patients ( $30,000 in medical costs) next year, using historical claims, diagnosis, procedure, and demographic data. Built for the Softtec 2026 ML Competition hosted at FAST NUCES, Lahore Campus . --- 📌 Problem Statement Healthcare payers need to identify members who are likely to incur high medical costs before those costs happen — so care management teams can intervene early with proactive outreach and reduce downstream spending. Task: Binary classification — predict HighCostLabel (1 = will cost more than $30,000 next year, 0 = otherwise) for each health plan member. Why it's hard: Only 11.18% of members in the training data are high-cost — a significant class imbalance that makes recall (catching the actually-expensive members) much harder than raw accuracy. --- 📊 Dataset The competition data is split across 5 linked tables, joined on Member Key : File Description --- --- main df Core member record — costs, visit counts, admissions, HCC risk scores, NEXT YEAR COST (training target source) drg df Diagnosis Related Group codes per claim cpt df CPT procedure codes per claim icd df ICD diagnosis codes per claim dob df Member date of birth and gender Training set: 109,562 samples (after filtering) · Test set: 12,387 samples Only records from MONTH = -1 or MONTH = 12 are kept — these snapshots were found to be the strongest predictors of next-year cost. --- ���","default_branch":null,"files":null,"tree":[],"storefront":"/r/Sara12-2","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/Sara12-2/High_Cost_Patient_prediction_Softec_Competition_Project/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}