applied ml
Cardiac Risk Stratification
Multi-modal cardiac MRI pipeline: U-Net segmentation, radiomics, and a calibrated classifier that explains every prediction.

about
Segments cardiac MRI with a U-Net, extracts radiomics and infarct-burden features from the masks, and feeds them to a calibrated XGBoost classifier. Calibration matters here: an uncalibrated risk score is not a risk score. Explainability runs at both levels: SHAP for the tabular features, Grad-CAM back onto the imaging, so a clinician can see which part of the heart drove the number.
multi-modal
imaging + tabular
calibrated
probability output
stack
- TensorFlow
- U-Net
- XGBoost
- SHAP
- Grad-CAM
- FastAPI
- Next.js
how it works
- 1
Segmentation
A U-Net trained on the EMIDEC dataset segments the myocardium from short-axis MRI, extended from 3 to 5 classes to include infarction and no-reflow.
- 2
Radiomics
PyRadiomics extracts shape, texture, and intensity features from the masks, plus infarct-burden features (infarct volume and share of myocardium).
- 3
Fusion
Imaging features are merged with four clinical biomarkers: age, LVEF, troponin, and NT-proBNP.
- 4
Prediction
A calibrated XGBoost classifier, tuned with nested Optuna search and evaluated under repeated stratified cross-validation, outputs Low / Moderate / High / Very High with class probabilities.
- 5
Explanation
SHAP contributions per feature, Grad-CAM heatmaps on the MRI, and a rule-based cross-check shown next to every prediction.
engineering notes
Dropped the ensemble
A stacked XGBoost + Attention-MLP ensemble was retired after permutation importance showed the MLP branch contributed zero signal. The single calibrated model matches its accuracy with far less complexity.
Traced a SHAP anomaly to the data
Troponin's SHAP value came out identical for 0.01, 0.08, and 4.5 ng/L. Traced to the booster level: not a code bug. The training distribution (mean ~70, std ~94) squeezes every value from 0 to ~25 into one narrow z-score band that no tree splits within.
Segmentation on very rare classes
Infarction reaches a Dice of 0.275 while making up only 0.176% of training pixels. No-reflow, ten times rarer, is not learnable from 100 patients, and the features built on it are documented as such.
known limits
- The training label comes from a rule on age, LVEF, troponin, and NT-proBNP, not a clinical outcome - the models imitate that rule, and accuracy should be read that way.
- Clinical-only requests fill imaging features with population medians, which can bias toward Very High Risk. The rule-based cross-check is there to catch it.