The idea
The project explored how different machine-learning approaches could identify people at risk for diabetes and how explanations could make the resulting profiles easier to inspect.
Approach
I compared Logistic Regression, XGBoost, MLP, and LSTM models on 253,680 BRFSS records. SHAP was used to explain model outputs, and K-Means clustering helped reveal data-driven patient risk profiles.
Reported outcomes
Tools & methods
More about the project
The work compared several model families and paired prediction with interpretation. The use of SHAP and clustering supported exploration of the patterns behind the risk classifications.
What I took from it
The project connected predictive performance with interpretability: a risk score is more useful to examine when the factors and broader profiles behind it can also be explored.