How hard does one patient pull an estimate? This lesson starts from a seesaw you can push and ends with the formula every efficient estimator in the course is built on. Each idea gets a picture first and its proper name second.
Here are the course's 100 patients, each resting on a beam at their outcome. The beam balances at their mean, written ψ (psi).
The slide is the patient's distance from the balance point, shrunk by their share of the mass. Per unit of mass, the pull is just that distance, z − ψ.
That pull has a name: the influence function, written ϕ (phi). For the mean, ϕ(z) = z − ψ.
Swap the 100 dots for a smooth curve of possible values. Drop a droplet of mass anywhere and record how far the balance point moves per unit of mass.
Sweep the droplet across the line. The recorded points fall on a straight line of slope 1 through ψ: the influence function ϕ(z) = z − ψ, traced by hand.
Why does ϕ deserve a special name? To see it, picture every possible distribution as a single point in one big space. The curve is one point.
Move μ and σ and the point slides across a flat sheet: the Normal family, where two numbers pin the curve down. Add the bump and the point lifts off the sheet, because no Normal has that shape.
A model is the set of points you allow. The Normal model is the sheet. The nonparametric model allows every density, so it is the whole region around the sheet.
The model picture is schematic: the sheet's grid is exact in μ and σ, but the height of the lift is not a distance. The bump is a direction the Normal model has no way to move in. That difference is why models of different sizes have different sets of directions (tangent spaces), which the geometry lessons draw.
From here on, p is the two-humped curve from the droplet step, not the Normal from the last step. At ε = 0 the path sits exactly on it.
Pick a direction h: a curve that says where to add mass (h above zero) and where to remove it (h below zero). Then slide ε to walk along the path.
The dashed curve is p, the solid curve is pε, and the shading shows mass moving from red to green. A path like this is also called a submodel: a one-dimensional slice of the model through p.
A finite-dimensional parametric model has finitely many independent directions, though infinitely many paths. The tangent space is the closed linear span of the scores of regular paths; in the full nonparametric model it is L²₀(P), every square-integrable function with mean zero. A linear path like this one also needs ε small enough that pε stays nonnegative; the displayed range guarantees it. All densities here live on one normalized grid on [−4, 4].
Fix one value z₀. As ε moves, the height of the curve at z₀ goes up or down. The rate at which its logarithm changes, at ε = 0, is the score at z₀.
Drag the probe and ε. The log-height panel shows a line through ε = 0 whose slope is exactly h(z₀).
Collect that slope at every z and you have a function: the score of the path. For this path it is the direction h you chose, because log(1 + εh) ≈ εh.
In symbols, h(z) = d/dε log pε(z) at ε = 0. For a Normal(μ, σ²) moved along μ, log pε(z) = const − (z − μ − ε)²/2σ², whose ε-derivative at 0 is (z − μ)/σ²: the familiar parametric score, and the "tilt" direction here up to scale.
Watch the two shaded areas being measured. Green counts up, red counts down.
They cancel exactly. The total mass is 1 at every ε, so its rate of change, ∫ h p, is zero: every score has mean zero.
In Build a canonical gradient, an optional 3D view draws each density p as the point √p on a sphere. Mean zero is the same fact as ⟨h√p, √p⟩ = EP[h] = 0: every score direction is perpendicular to √p, so paths stay on the sphere.
Now walk along the path and watch the mean ψ move. Its speed is the signed area of ϕ·h·p, where ϕ(z) = z − ψ is the same pull you traced on the seesaw.
Compare the two readout numbers: the actual move per unit ε and the area E[ϕh]. For the mean they agree exactly, because the mean is linear in p. This speed is called the pathwise derivative.
The derivation: ψ(Pε) = ∫ z p(z)(1 + εh(z)) dz = ψ(P) + ε ∫ z h p dz. Because ∫ h p = 0 you may subtract any constant from z inside the last integral, so the slope is ∫ (z − ψ) h p = EP[ϕh]. For nonlinear parameters the finite difference and E[ϕh] agree to first order in ε. The Mean Along a Path lesson takes this one line apart in slow motion.
Change h and the speed changes, but ϕ never has to be recomputed: the speed is always E[ϕh]. The table checks this for four directions against the actual move.
One function that gives the speed in every direction is what a gradient is. A parameter with such a ϕ is called pathwise differentiable, and that is the entry ticket to the whole theory.
In the nonparametric model the tangent space is every square-integrable mean-zero function, so the gradient is unique: the efficient influence function of the mean is z − μ with no alternatives. In a smaller model several gradients agree on the tangent space, and the efficient one is their projection onto it (Build a Canonical Gradient draws this).
Adding a droplet of mass ε at z₀ is itself a way to move the distribution: shrink p by (1 − ε) and put ε at z₀. Move the spike and watch the mean move by exactly ε·(z₀ − ψ).
So the seesaw pull and the pathwise speed are the same ϕ. Put mass 1/n at each of n observations and the estimate's error is, to first order, the average of ϕ over the sample.
For a discrete mass function with p(z₀) > 0, the contamination path Pε = (1 − ε)P + εδz₀ has score h(z) = 𝟙{z = z₀}/p(z₀) − 1, and E[ϕh] = ϕ(z₀). For a continuous density it is only a formal device (a Gateaux derivative), not a regular path with an L² score. For the mean, the error of the sample mean is exactly Pₙ(Z − μ). For a general asymptotically linear estimator the leading error is Pₙϕ under additional conditions; Pₙ − P is a signed empirical perturbation, not one small point mass.
Now one observation is (x₀, a₀, y₀): covariate, arm, outcome. The target is the average treatment effect, ψ = average over x of m₁(x) − m₀(x), where ma(x) is the mean outcome in arm a at x.
Add a little mass at that observation and it pulls ψ two ways. First, x₀ gets a little more weight, which moves ψ by m₁(x₀) − m₀(x₀) − ψ. Second, the regression for its own arm is pulled toward y₀, by the residual divided by the chance of that arm at x₀.
Move x₀ to where the chosen arm is rare (a control at high x, say). The second piece grows, because one rare observation stands in for many. That is where inverse-propensity weights come from, and why positivity matters.
P(A = a₀ | x₀) is g(x₀) for a treated observation and 1 − g(x₀) for a control. The second piece is the correction term in AIPW.
Where this leads: a parameter changes along any path at E[ϕ h], so ϕ is the one object that gives the first-order effect of every allowed change to the data. How ϕ also sets the best precision any estimator can reach comes later, in Build a canonical gradient and Efficiency Theory, Drawn.
Some directions tell you a lot about ψ and some tell you nothing. Along one path the problem has a single unknown, ε, so it has an ordinary variance floor.
Rotate h away from ϕ and watch that floor. It is highest when h points along ϕ, and there it equals Var(ϕ).
No estimator can beat the floor on the hardest path, so none can beat Var(ϕ) overall. That is the efficiency bound, and the hardest path is called the least favorable one.
Along a path with score h the problem is parametric in ε, so Cramér–Rao applies: no regular estimator can have asymptotic variance below (dψ/dε)²/I, with Fisher information I = E[h²]. Substituting dψ/dε = E[ϕh] gives E[ϕh]²/E[h²], which by Cauchy–Schwarz is at most E[ϕ²], with equality when h ∝ ϕ.
TMLE fluctuates along a submodel whose score is the efficient influence function for this reason: it is the direction in which the data are least informative about ψ per unit of likelihood.
What is exact, simulated and schematic: every density number is an exact sum on a fine grid, the seesaw uses the course's 100 cohort outcomes exactly, the ATE step uses a known law, and the model sheet in step 3 is schematic.
Further reading: the construction follows Tsiatis, Semiparametric Theory and Missing Data (ch. 3–4), and Schuler & van der Laan, Introduction to Modern Causal Inference (§3.1–3.3). The point-mass derivation of the ATE influence function is the Gateaux-derivative route in Hines et al. (2022), "Demystifying statistical learning based on efficient influence functions."