Using health survey data to predict heart disease risk with transparency.

A PR-AUC jump from 0.23 to 0.51 is exactly the kind of red flag that deserves a public write-up, and this project delivers it. The author stripped out the entire cardiovascular questionnaire section, then quantified how…

4 min readMachine Learning

Most predictive modeling projects hide their messiest decisions. This one puts them on the table. The author took four cycles of NHANES data, roughly 21,500 adults, and built a coronary heart disease classifier using demographics, blood pressure, body measurements, and a lipid panel. The headline numbers are honest: ROC-AUC of 0.875 and PR-AUC of 0.239 for logistic regression, with random forest and gradient boosting landing about the same. But what makes this worth your attention is not the performance. It is the deliberate, documented handling of two traps that routinely sink applied machine learning work: target leakage and miscalibration.

The leakage story is a masterclass in intellectual honesty. NHANES asks respondents directly about other cardiovascular diagnoses, including stroke, heart attack, and angina. Including those variables sent PR-AUC from 0.23 to 0.51. That is not a feature engineering win; it is the model learning that people who say they have one heart condition tend to say they have others. The author could have quietly dropped those columns and moved on. Instead, they quantified the effect and wrote it up. That is the kind of transparency we need more of, especially in healthcare-adjacent work where inflated metrics can lead to false confidence. For anyone building similar models, this is a reminder that the gap between correlation and causation is not a bug you fix by adding more data; it is a design choice you own.

The calibration section is equally instructive. With a CHD prevalence around 4%, the class-weighted logistic regression produced raw predicted probabilities near 30% while the actual rate stayed at 4%. That is a massive disconnect, and the fix was a sigmoid recalibration fit on the development set before the test set was touched. The author also froze the decision threshold on the dev set, so the reported test metrics were not tuned after the fact. This is how you build trust in a model: by showing your work, not just your final accuracy. For readers who are new to this space, the practical takeaway is simple. If you are using a model with rare outcomes, do not trust raw probabilities. Recalibrate, validate on a held-out set, and report the precision-recall curve, not just ROC-AUC. Age alone gets 0.83 AUC here, so the added complexity of the full model buys you relatively little. That is not a failure; it is a signal about where the real signal lives.

What would we tell a reader who asked about this project? Start by reading the code and the write-up, then ask yourself what you would do differently. The author openly notes that smoking status, diabetes, and blood pressure medication use are not yet in the feature set, even though NHANES has all three. That is a concrete next step, and it is likely to move the needle more than any additional model tuning. The PPV at the chosen threshold is 0.13, meaning most positive predictions are wrong. The author states this directly, which is rare and refreshing. The open question we are watching is whether the calibration approach generalizes to other rare-event datasets, and whether the author will add those missing features in a future iteration. That is the detail to watch, because it will tell us if this is a one-off exercise or the start of a more rigorous standard for applied health data work.

From Machine Learning

https://github.com/YouCele/nhanes-chd-classification

I built a project that predicts self-reported, physician-diagnosed coronary heart disease using four cycles of NHANES data (2011-2012 to 2017-2018), about 21,500 adults after cleaning. It compares logistic regression with random forest and gradient boosting, trained on demographics, blood pressure, body measurements and a lipid panel.

Read the original at Machine Learning