You have built every piece of efficiency theory already. This page lays five of them side by side, in the order they hand off to each other. Replay what you like and skip what feels familiar.
Each point in the blob 𝓜 is a whole distribution your assumptions allow. The estimand ψ sends each one to a number, such as the ATE, ψ(P) = E[m₁(X) − m₀(X)], where ma(x) is the mean outcome in arm a.
Level curves join distributions with the same value of ψ. Everything later is about how this map bends near the truth.
This is a finite-dimensional illustration: the landscape is the two-parameter Gaussian model of the last scene, and the blob is a viewing window. Nonparametric models may be infinite dimensional. Identification is what lets ψ at the truth be the causal quantity; estimation is about the shape of the map.
A good estimator's error is, to first order, the average of each patient's pull ϕ(Oi) over the sample. That is what asymptotically linear means, and it is why the errors pile up in a bell curve of width √(Var ϕ / n).
Each ball is one whole study. For the first three, the arrows show each patient pushing the ball by ϕ(Oi)/n; a last small nudge, the remainder, lands it on the estimate. Red balls are 95% intervals that missed the truth.
Simulated: 200 seeded studies for each n (AIPW with both nuisance models fitted, no cross-fitting; the plug-in is the treated-minus-control difference). The arrows use the true influence function, which only a simulation can know. Exact: the centres 2 and 2.64 and the outline curves. Schematic: the pegs.
An estimator is asymptotically linear when ψ̂ − ψ(P) = Pnϕ + oP(n−1/2) for a mean-zero ϕ. With independent observations and finite Var(ϕ), the central limit theorem gives the normal limit. Here is the simplest case: each observation contributes ϕ(Zi) = Zi − μ, and the error is their average.
Move away from P along a path with score h. The parameter changes at rate EP[ϕh], the same ϕ for every direction. Rotate the direction and watch the derivative rise and fall as the cosine of the angle to the gradient.
Regularity of an estimator requires that its limit not depend on which of these paths the truth drifts along. The Hájek–Le Cam convolution theorem then says the best regular estimator has the gradient in the tangent space as its influence function.
Put each density at the point √p, so every distribution lies on the unit sphere of L₂. A path is a curve on the sphere; its velocity ½h√p is orthogonal to √p, which is exactly EP[h] = 0. The closed linear span of the scores is the tangent space 𝒯.
In an RCT the treatment mechanism is known, so paths cannot move it and the tangent space loses those directions. For the population ATE the efficient influence function is already orthogonal to them, so knowing the propensity alone leaves the bound unchanged.
Estimators that agree on every pathwise derivative can still have different influence functions. All of them lie on one line Φ, and exactly one lies inside the tangent space 𝒯.
That one is the shortest, so it has the smallest variance: the efficient influence function, written D*. D* is the ϕ of the efficient estimator.
Augmenting an inefficient estimator, as AIPW augments IPW, is projecting its influence function onto 𝒯.
Any two gradients have the same covariance with every score, so their difference is orthogonal to 𝒯: the set of mean-zero gradients is D* + 𝒯⊥. It is a line only when that complement is one dimensional, as drawn. By Pythagoras, Var(ϕ) = ‖D*‖² + ‖ϕ − D*‖².
Fit the nuisances, get P̂, and evaluate ψ(P̂): the plug-in. Walk the straight path from P̂ to the truth P and watch ψ along it.
The slope at P̂ toward the truth cannot be computed, but its sample version, Pnϕ̂, can. Add it back and the first-order bias is gone: that is the one-step estimator, and for the ATE it is AIPW.
Halve the nuisance error: the first-order term halves and the remainder quarters. What is left is sampling noise, which shrinks like 1/√n, and a remainder that is second order in the nuisance error.
For the ATE, Ψ(P̂) − Ψ(P₀) = −P₀D*(P̂) + R₂ with R₂ = E₀[(ĝ − g₀){(m̂₁ − m₁₀)/ĝ + (m̂₀ − m₀₀)/(1 − ĝ)}], where g is the propensity score and ma the outcome regressions. With inverse propensities bounded, |R₂| is at most a constant times the product of the two L² errors: the rectangle below. It depicts that magnitude bound, not the signed remainder. Efficient inference needs R₂ = oP(n−1/2); two errors of exactly n−1/4 do not by themselves meet it. The sampling piece (Pn − P)(ϕ̂ − ϕ) must also be oP(n−1/2), which is where Donsker conditions or cross-fitting come in. The inference laboratory tests all of this with coverage.
If either error is exactly zero the rectangle has no area: double robustness. If both shrink faster than n−1/4, the product is smaller than n−1/2.
The one-step corrects the number and never moves in 𝓜. TMLE instead nudges the fitted distribution P̂ along the direction of the efficient influence function until its average over the sample is zero, then plugs in.
Play the walk and compare where each lands on the ψ line below the model.
Both remove the same first-order bias. TMLE's answer is a plug-in, so it always stays inside the parameter space, for example between 0 and 1 for a risk.
The model here is fully specified: Z follows N₂(θ, I) and Ψ(θ) = 2 + 0.9θ₁ + 0.5θ₂ + c(0.6θ₁² − 0.7θ₁θ₂ + 0.2θ₂²). Its efficient influence function is ∇Ψ(θ)·(Z − θ), and the line θ + ε∇Ψ(θ) is locally least favorable, with ε̂ = ∇Ψ(θ)·(Z̄ − θ)/‖∇Ψ(θ)‖². The point Pₙ stands for the sufficient statistic Z̄. A single line fit zeros its own score; for curved Ψ it need not zero the updated one, so the default re-aims and iterates (iterated local submodels, not a universal submodel). Equality with the one-step at c = 0 is specific to this example.
Next, the inference laboratory puts these promises to the test: the rate boundary from the remainder rectangle, cross-fitting, and what happens to interval coverage when a flexible learner is fitted and evaluated on the same patients.
Further reading: the geometry follows Schuler & van der Laan, Introduction to Modern Causal Inference, ch. 3–4; the projection picture is Tsiatis's (Semiparametric Theory and Missing Data, ch. 3–4); the remainder rectangle is the Cauchy–Schwarz bound as in Kennedy, "Semiparametric doubly robust targeted double machine learning: a review" (2022), and Chernozhukov et al. (2018), "Double/debiased machine learning".