Causality · capstone

An emulated trial, end to end

One analysis, start to finish: the main causal workflow from the course, used in order: the estimand, the protocol, positivity, a Super Learner with cross-fitting, an influence-function interval, an E-value, and the paragraph you would send to the heart team.

Planning a randomized trial instead? The SAP builder writes the estimand, the covariate-adjusted primary analysis, and the operating characteristics a reviewer will ask for.

Further reading: Hernán and Robins (2016), Using big data to emulate a target trial when a randomized trial is not available, American Journal of Epidemiology 183(8):758–764; Cashin et al. (2025), the TARGET guideline for reporting target trial emulations, JAMA; ICH E9(R1); van der Laan, Polley and Hubbard (2007), Super Learner, Statistical Applications in Genetics and Molecular Biology 6(1); Chernozhukov et al. (2018), Double/debiased machine learning for treatment and structural parameters, The Econometrics Journal 21(1):C1–C68; Petersen et al. (2012), Diagnosing and responding to violations in the positivity assumption; VanderWeele and Ding (2017), Sensitivity analysis in observational research: introducing the E-value, Annals of Internal Medicine 167(4):268–274; FDA (2025), Use of Real-World Evidence To Support Regulatory Decision-Making for Medical Devices (final guidance). The registry is simulated and seeded; its truth is known to very high Monte Carlo precision (a one-million-patient average) because the data-generating mechanism is known.