← Course map

Symbols with a job

Read a symbol as an action or an object. The same letter can mean different things in different papers; the definitions below are this course's convention.

P₀
The true probability law generating the observed data. We use it to define the target and evaluate an estimator, but do not know it in real data.
P
A generic candidate law. A point in the statistical model, not necessarily the truth.
Ψ and ψ₀
Ψ is a map from a probability law to a number. ψ₀=Ψ(P₀) is its true value. These are different kinds of object.
P̂ and ψ̂
P̂ is a fitted law (or collection of fitted nuisance components). ψ̂ is an estimate of the target. A plug-in is Ψ(P̂).
Pₙ f
The empirical average n⁻¹Σᵢf(Oᵢ). The empirical distribution is a probability measure, but need not belong to a model of smooth densities.
O or Z
One whole observation, often (X,A,Y). In the mean lessons, Z is a scalar outcome. Lowercase z is a possible value, not another parameter.
p and pε
A density or probability mass function and a nearby version on a path. In pε=p(1+εh), mean-zero h preserves total mass; allowed ε must also preserve nonnegativity.
h, a score
The derivative of log pε at ε=0. It measures relative probability change, and has mean zero under regular normalized paths.
L²₀(P)
Mean-zero functions with finite second moment. Their inner product is Eₚ[fg], and squared length is Eₚ[f²].
T, tangent space
The closed linear span of scores of regular paths through P allowed by the model. A finite plane is one illustration, not the general definition.
Tη, nuisance tangent space
Allowed directions along which the target has zero first derivative. What counts as nuisance depends on the target.
D*, canonical gradient / EIF
The unique element of T representing target derivatives: derivative along h equals Eₚ[D*h]. It is mean-zero. In a restricted model these properties need membership in T to establish uniqueness.
D, ϕ, φ
Common names for a gradient or an estimator's influence function. Specify which one: all mean-zero gradients have the form D*+r with r orthogonal to T; only D* has minimum variance in the regular class.
g or π
The treatment propensity P(A=1 | X). “Positivity” asks that the treatment levels needed for the target occur wherever they are required.
m or μ; Q in some papers
The conditional outcome mean mₐ(x)=E[Y | A=a,X=x]. Q may denote a broader outcome component. It is not automatically the whole fitted law P̂.
H, clever covariate
For the ATE, H=A/g−(1−A)/(1−g). It weights residuals inside the efficient influence function. For a treated-mean target it is A/g.
R₂, remainder
What is left after a first-order expansion. For the ATE it contains products of propensity and outcome errors. A small product helps inference; a rate exactly n⁻¹⁄² is not little-o of that rate.
ATE, ATT, ATC
Average treatment effects for everyone, for those actually treated, and for those actually untreated. The selected people stay fixed across both intervention worlds; their baseline composition changes which group effects receive weight. Move the population weights →
Risk difference, risk ratio
For the same population and time horizon, subtract the intervention risks or divide them. A difference of −0.1 is −10 percentage points; a ratio of 0.5 means half the risk. The ratio needs a nonzero denominator. Compare event grids →
Sₐ(t), RMSTₐ(τ)
Survival under treatment a, and its area through horizon τ. A contrast between arms can be a causal target under stated identifying assumptions. A survival difference is a vertical gap (probability); an RMST difference is a signed area (time). Move the time horizon →

Turn these symbols into movements →