BIOSTATR
Applied Biostatistics with R
BIOSTATR
Applied Biostatistics with RFoundations
Course description
BIOSTATR is the base course of the MHDSR pathway: a practical, rigorous introduction to applied biostatistics for health research using R. It takes participants from a clean R and data-manipulation workflow through inference, regression, survival analysis, and missing-data handling with multiple imputation, all with reproducible Quarto reporting.
No statistics or programming background is required. The course emphasizes estimation, interpretation, reproducibility, and the appropriate use of statistical methods in clinical research.
Learning objectives
By the end of BIOSTATR, participants will be able to:
- organize a reproducible analysis with R, Quarto, and version control;
- import, clean, reshape, and quality-check health datasets;
- summarize and visualize health data appropriately, including stratified summaries;
- interpret confidence intervals, p-values, effect sizes, power, and sample size;
- compare groups with suitable tests for quantitative and qualitative outcomes;
- fit and interpret linear and logistic regression, including confounding and interactions;
- analyze time-to-event data with Kaplan-Meier, the log-rank test, and the Cox model;
- handle missing data with multiple imputation and Rubin’s rules.
Schedule
| No. | Session |
|---|---|
| 1 | R foundations and data manipulation: environment, projects, scripts, objects and types, import/export, filtering, transformation, joins, pivots, dates, strings, recoding, quality control, reproducible reports |
| 2 | Descriptive statistics and visualization: descriptive statistics, distributions, summary tables, Table 1, graphics, stratification |
| 3 | Statistical inference and hypothesis testing: estimation, confidence intervals, tests, effect sizes, power, sample size, quantitative and qualitative comparisons |
| 4 | Linear and logistic regression: adjusted coefficients, odds ratios, confounding, interactions, diagnostics, predictions, ROC curve |
| 5 | Survival analysis: censoring, Kaplan-Meier, log-rank test, Cox model, hazard ratios, proportional-hazards assumption |
| 6 | Missing data and multiple imputation: MCAR/MAR/MNAR, complete-case analysis, single and multiple imputation, imputation model, Rubin’s rules, sensitivity analyses |
Practical information
- Audience
- Health researchers, clinicians, students; no statistics background required
- Prerequisites
- None; R/RStudio setup covered in Session 1
- Format
- Live sessions, slides, notebooks, datasets, exercises
- Schedule
- Mon/Wed/Fri 18h00–20h00 (Morocco time) · 2 weeks
- Tools
- R, RStudio, Quarto
- Assessment
- Reproducible Quarto mini-project
- Next step
- ML4HOR – Machine Learning for Health Outcomes with R