ML4HOR
ML4HOR
Machine Learning for Health Outcomes with RPrediction
Course description
ML4HOR is the advanced course of the MHDSR pathway: a practical introduction to machine learning for health outcomes research using R. It covers the full supervised workflow (from classic models to boosting and neural networks), unsupervised learning and clustering for patient phenotyping, honest tuning and model selection, evaluation and validation, and interpretability.
Agentic and AI-assisted workflows – driving a reproducible ML pipeline with a command-line coding agent, and LLM-assisted structured extraction and text classification – are covered in the companion course AIDSR.
The course emphasizes validation, data-leakage prevention, appropriate performance metrics, calibration, interpretability, and clinical relevance.
Learning objectives
By the end of ML4HOR, participants will be able to:
- frame supervised and unsupervised problems in health outcomes research;
- prepare tabular health datasets and avoid data leakage;
- apply clustering for patient phenotyping and subgroup discovery;
- fit penalized regression, tree ensembles, boosting, and multilayer perceptrons;
- tune and select models with nested cross-validation and fair comparison;
- evaluate discrimination, calibration, and clinical usefulness, and choose thresholds;
- validate models internally and externally;
- interpret models with permutation importance, PDP/ICE, and SHAP.
Schedule
| No. | Session |
|---|---|
| 1 | Machine learning fundamentals: supervised vs unsupervised, bias-variance, train/test, cross-validation, resampling |
| 2 | Data preparation: preprocessing, missing data, encoding, standardization, feature engineering, data leakage |
| 3 | Unsupervised learning and clustering: distances, k-means, hierarchical clustering, choosing K, silhouette, stability, PCA for exploration and visualization |
| 4 | Supervised models: penalized regression, decision trees, random forests, gradient boosting, multilayer perceptron, regularization, early stopping, model comparison |
| 5 | Tuning, model selection and validation: hyperparameters, grid/random search, nested cross-validation, overfitting, AUC, sensitivity/specificity, PPV/NPV, F1, calibration, Brier score, thresholds, internal/external validation |
| 6 | Interpretability and clinical usefulness: variable importance, permutation importance, PDP/ICE, SHAP, local vs global interpretation, decision thresholds in context |
Practical information
- Audience
- Researchers, PhD students, clinicians, public health professionals, data analysts
- Prerequisites
- Basic R and statistical literacy, as covered in BIOSTATR
- Format
- Live sessions, slides, notebooks, datasets, exercises
- Schedule
- Mon/Wed/Fri 18h00–20h00 (Morocco time) · 2 weeks
- Tools
- R, RStudio, Quarto
- Assessment
- Reproducible Quarto mini-project
- Next step
- AIDSR – Agentic & AI-assisted Data Science with R