Symbols with a job
Read a symbol as an action or an object. The same letter can mean different things in different papers; the definitions below are this course's convention.
- P₀
- The true probability law generating the observed data. We use it to define the target and evaluate an estimator, but do not know it in real data.
- P
- A generic candidate law. A point in the statistical model, not necessarily the truth.
- Ψ and ψ₀
- Ψ is a map from a probability law to a number. ψ₀=Ψ(P₀) is its true value. These are different kinds of object.
- P̂ and ψ̂
- P̂ is a fitted law (or collection of fitted nuisance components). ψ̂ is an estimate of the target. A plug-in is Ψ(P̂).
- Pₙ f
- The empirical average n⁻¹Σᵢf(Oᵢ). The empirical distribution is a probability measure, but need not belong to a model of smooth densities.
- O or Z
- One whole observation, often (X,A,Y). In the mean lessons, Z is a scalar outcome. Lowercase z is a possible value, not another parameter.
- p and pε
- A density or probability mass function and a nearby version on a path. In pε=p(1+εh), mean-zero h preserves total mass; allowed ε must also preserve nonnegativity.
- h, a score
- The derivative of log pε at ε=0. It measures relative probability change, and has mean zero under regular normalized paths.
- L²₀(P)
- Mean-zero functions with finite second moment. Their inner product is Eₚ[fg], and squared length is Eₚ[f²].
- T, tangent space
- The closed linear span of scores of regular paths through P allowed by the model. A finite plane is one illustration, not the general definition.
- Tη, nuisance tangent space
- Allowed directions along which the target has zero first derivative. What counts as nuisance depends on the target.
- D*, canonical gradient / EIF
- The unique element of T representing target derivatives: derivative along h equals Eₚ[D*h]. It is mean-zero. In a restricted model these properties need membership in T to establish uniqueness.
- D, ϕ, φ
- Common names for a gradient or an estimator's influence function. Specify which one: all mean-zero gradients have the form D*+r with r orthogonal to T; only D* has minimum variance in the regular class.
- g or π
- The treatment propensity P(A=1 | X). “Positivity” asks that the treatment levels needed for the target occur wherever they are required.
- m or μ; Q in some papers
- The conditional outcome mean mₐ(x)=E[Y | A=a,X=x]. Q may denote a broader outcome component. It is not automatically the whole fitted law P̂.
- H, clever covariate
- For the ATE, H=A/g−(1−A)/(1−g). It weights residuals inside the efficient influence function. For a treated-mean target it is A/g.
- R₂, remainder
- What is left after a first-order expansion. For the ATE it contains products of propensity and outcome errors. A small product helps inference; a rate exactly n⁻¹⁄² is not little-o of that rate.
- ATE, ATT, ATC
- Average treatment effects for everyone, for those actually treated, and for those actually untreated. The selected people stay fixed across both intervention worlds; their baseline composition changes which group effects receive weight. Move the population weights →
- Risk difference, risk ratio
- For the same population and time horizon, subtract the intervention risks or divide them. A difference of −0.1 is −10 percentage points; a ratio of 0.5 means half the risk. The ratio needs a nonzero denominator. Compare event grids →
- Sₐ(t), RMSTₐ(τ)
- Survival under treatment a, and its area through horizon τ. A contrast between arms can be a causal target under stated identifying assumptions. A survival difference is a vertical gap (probability); an RMST difference is a signed area (time). Move the time horizon →