Semiparametric efficiency, drawn

The one-step estimator, as a shape

Three acts. First the geometry: a plug-in is a point on a curve, the influence function is the tangent, and one Newton step lands you within the curvature of the truth. Then the same thing on simulated data, one operation at a time. Then repeat the experiment to measure bias, variance, and interval coverage separately. The opening curve is a schematic first-order expansion, not a fully specified statistical model.

Act 1

The curve ψ(Pε) along the path from your estimate to the truth

Put your fitted distribution P̂ at ε = 0 and the true P at ε = 1, and walk the straight line between them, Pε = P̂ + ε(P − P̂). The quantity you want, ψ, traced along that walk is a curve. At the left end it is the plug-in (orange); at the right end, the truth (green).

The tangent at the left end is drawn by the influence function. Write it ϕ (ϕ̂ = ϕ(P̂) when it is computed at your fit). The tangent's rise over the walk is P ϕ̂ = ∫ ϕ̂ d(P − P̂), and the plug-in's first-order bias is its negative.

Extending the tangent to ε = 1 gives the one-step estimator (purple). What the tangent misses is the curvature: the second-order remainder R2.

ψ(P̂) − ψ(P) = −P ϕ̂ + R2(P̂, P) ⇒ ψone-step = ψ(P̂) + Pn ϕ̂ ≈ ψ(P) + R2 + (Pn − P) ϕ̂
plug-in ψ(P̂) tangent → one-step remainder R2 truth ψ(P)
Drag δ down and watch: the plug-in bias (first order) shrinks like δ, the remainder shrinks like δ². For efficient centered inference, the remainder must be oₚ(n−1/2). Two errors exactly of order n−1/4 only give an Oₚ(n−1/2) product; faster combined rates or a sharper remainder argument are needed.
Act 2

One step on data: g-computation, then the correction

Treatment is more likely for higher X, and the outcome is nonlinear in X. Two fitted models do the work, and the course names them the same way everywhere:

The outcome regression ma(x) = E[Y | A = a, X = x] is the average outcome in arm a at covariate x (m̂a when fitted). The propensity score g(x) = P(A = 1 | X = x) is the chance of treatment at x (ĝ when fitted). You may see these written π, μ or Q̄ elsewhere.

Here m̂ is an S-learner: one linear regression of Y on (X, A). The slider shrinks its treatment coefficient toward zero, the way a regularized learner would. The propensity model can be made anywhere from correct to useless.

ψ̂one-step = Pn[ m̂1(X) − m̂0(X) ] + Pn[ A/ĝ(X) · (Y − m̂1(X)) − (1−A)/(1−ĝ(X)) · (Y − m̂0(X)) ]
Step through the figure below. When the correction Pₙϕ̂ (the second bracket) is added to the S-learner plug-in, where does the estimate end up?

Watch where the correction comes from: it is the inverse-propensity-weighted average of the outcome model's own residuals.

treated (A=1) control (A=0) plug-in one-step truth = 2
This linear outcome model is misspecified even at λ=0: it omits X² and A×X. At λ=0 with fidelity 0, its correction is zero by the OLS normal equations, not evidence that the outcome model is correct. Symmetry can also cancel ATE bias here. The experiment below uses explicit correctness cases and an asymmetric population.

The same formula, one patient at a time

Now back to the course's 100 patients, where severity confounds treatment and the true average effect is 2. Same formula, same m̂ and ĝ: here ĝ(X) is simply the fraction treated in each severity stratum. The equation is the legend. Hover or focus a coloured term to light its pieces in the figure, or hover the figure to find its term. Each patient carries two sticks, and Play adds 1/n of every stick, tip to tail, into the two sums of the formula.

Exact, not simulated: every number is computed from the cohort. With ĝ equal to the stratum fractions, AIPW equals the severity-stratified estimate for any outcome model that depends only on arm and severity, because within a stratum the weighted residuals of an arm add up to exactly nx·(Ȳax − m̂a(x)), which cancels m̂. A smoother propensity model would repair the wrong outcome model only to first order, which is the general double-robustness story of Act 3.

Act 3

Repeat samples and separate the promises

One data set gives one estimate. To see bias, spread and interval coverage you need the whole sampling distribution, so repeat the study hundreds of times. This experiment crosses the two nuisance models: the outcome model is either right or wrong, and so is the propensity model.

In which panels will the AIPW (purple) histogram be centred away from the true ATE?

Centred is not the same as covered. In the default run (fitted, cross-fitted, n = 400, 300 samples) the outcome-only panel covers about 92%, a little short, because its influence-function SE (0.083) is smaller than the actual spread (0.089). The propensity-only panel is centred but over-covers: its IF-based SE averages 0.142 while the estimates actually spread with SD 0.096, because the formula treats the fitted propensity as if it were known, and with a wrong outcome model, estimating the propensity removes variance that the formula still counts. Getting a standard error you can report is its own problem: see Standard errors you can report.

Explain the correction

You have now derived the estimator you used in Your trial, adjusted. There, randomization fixed g, so the correction term had the right average whatever outcome model you fit. Its "Preview" boxes should now read as a summary of this lesson.

The tangent toward the truth has slope P₀ϕ̂; the plug-in's first-order error has the opposite sign. The one-step adds an empirical estimate of that slope. A small nuisance-error product gives double robustness of the point estimator. Efficient inference additionally requires a negligible remainder, empirical-process control (often through cross-fitting), and ϕ̂ converging to the efficient influence function.

R₂ = E₀[(ĝ−g₀){(m̂₁−m₁)/ĝ + (m̂₀−m₀)/(1−ĝ)}] (g, ma: the true functions)
ψ(P̂)−ψ₀ = −P₀ϕ̂ + R₂

TMLE changes the fitted distribution and plugs in; the one-step corrects the numerical estimate. Suitable versions share a first-order expansion under regularity conditions, without being exactly equal in every sample.