Foundations · start with something you can hold

Scores and Influence, From Scratch

How hard does one patient pull an estimate? This lesson starts from a seesaw you can push and ends with the formula every efficient estimator in the course is built on. Each idea gets a picture first and its proper name second.

Step 1

The mean is a balance point

Here are the course's 100 patients, each resting on a beam at their outcome. The beam balances at their mean, written ψ (psi).

One more patient lands 3.6 units to the right of the balance point and counts as 1 of 101 patients. How far does the balance point slide?

The slide is the patient's distance from the balance point, shrunk by their share of the mass. Per unit of mass, the pull is just that distance, z − ψ.

That pull has a name: the influence function, written ϕ (phi). For the mean, ϕ(z) = z − ψ.

Step 2

Trace the pull at every place

Swap the 100 dots for a smooth curve of possible values. Drop a droplet of mass anywhere and record how far the balance point moves per unit of mass.

Sweep the droplet across the line. The recorded points fall on a straight line of slope 1 through ψ: the influence function ϕ(z) = z − ψ, traced by hand.

In one sentenceThe influence function at z is how hard a little extra mass at z pulls the parameter, per unit of mass.
Step 3

A whole distribution is one point

Why does ϕ deserve a special name? To see it, picture every possible distribution as a single point in one big space. The curve is one point.

Move μ and σ and the point slides across a flat sheet: the Normal family, where two numbers pin the curve down. Add the bump and the point lifts off the sheet, because no Normal has that shape.

A model is the set of points you allow. The Normal model is the sheet. The nonparametric model allows every density, so it is the whole region around the sheet.

Go deeper

The model picture is schematic: the sheet's grid is exact in μ and σ, but the height of the lift is not a distance. The bump is a direction the Normal model has no way to move in. That difference is why models of different sizes have different sets of directions (tangent spaces), which the geometry lessons draw.

Step 4

A path: nudge the distribution a little at a time

From here on, p is the two-humped curve from the droplet step, not the Normal from the last step. At ε = 0 the path sits exactly on it.

Pick a direction h: a curve that says where to add mass (h above zero) and where to remove it (h below zero). Then slide ε to walk along the path.

pε(z) = p(z) · (1 + ε·h(z)) (h averages to zero under p, so the total mass stays 1)

The dashed curve is p, the solid curve is pε, and the shading shows mass moving from red to green. A path like this is also called a submodel: a one-dimensional slice of the model through p.

Go deeper

A finite-dimensional parametric model has finitely many independent directions, though infinitely many paths. The tangent space is the closed linear span of the scores of regular paths; in the full nonparametric model it is L²₀(P), every square-integrable function with mean zero. A linear path like this one also needs ε small enough that pε stays nonnegative; the displayed range guarantees it. All densities here live on one normalized grid on [−4, 4].

Step 5

The score is a slope, one number per z

Fix one value z₀. As ε moves, the height of the curve at z₀ goes up or down. The rate at which its logarithm changes, at ε = 0, is the score at z₀.

Drag the probe and ε. The log-height panel shows a line through ε = 0 whose slope is exactly h(z₀).

Collect that slope at every z and you have a function: the score of the path. For this path it is the direction h you chose, because log(1 + εh) ≈ εh.

Go deeper

In symbols, h(z) = d/dε log pε(z) at ε = 0. For a Normal(μ, σ²) moved along μ, log pε(z) = const − (z − μ − ε)²/2σ², whose ε-derivative at 0 is (z − μ)/σ²: the familiar parametric score, and the "tilt" direction here up to scale.

Step 6

Every score averages to zero

Take any direction h that keeps pε a density (total mass 1 for every ε). Weight h by the density and add it up: what is ∫ h(z) p(z) dz?

Watch the two shaded areas being measured. Green counts up, red counts down.

They cancel exactly. The total mass is 1 at every ε, so its rate of change, ∫ h p, is zero: every score has mean zero.

d/dε ∫ pε = ∫ p·h = EP[h] = 0
Go deeper

In Build a canonical gradient, an optional 3D view draws each density p as the point √p on a sphere. Mean zero is the same fact as ⟨h√p, √p⟩ = EP[h] = 0: every score direction is perpendicular to √p, so paths stay on the sphere.

Step 7

The parameter along the path: follow the mean

Now walk along the path and watch the mean ψ move. Its speed is the signed area of ϕ·h·p, where ϕ(z) = z − ψ is the same pull you traced on the seesaw.

Compare the two readout numbers: the actual move per unit ε and the area E[ϕh]. For the mean they agree exactly, because the mean is linear in p. This speed is called the pathwise derivative.

dψ/dε at 0 = ∫ ϕ(z) h(z) p(z) dz = EP[ϕ h], ϕ(z) = z − ψ
Go deeper

The derivation: ψ(Pε) = ∫ z p(z)(1 + εh(z)) dz = ψ(P) + ε ∫ z h p dz. Because ∫ h p = 0 you may subtract any constant from z inside the last integral, so the slope is ∫ (z − ψ) h p = EP[ϕh]. For nonlinear parameters the finite difference and E[ϕh] agree to first order in ε. The Mean Along a Path lesson takes this one line apart in slow motion.

In one sentenceAlong a path with score h, the mean moves at speed E[ϕ h], with ϕ(z) = z − ψ: a property of the parameter, found with no estimator in sight.
Step 8

The same ϕ works for every direction

Change h and the speed changes, but ϕ never has to be recomputed: the speed is always E[ϕh]. The table checks this for four directions against the actual move.

One function that gives the speed in every direction is what a gradient is. A parameter with such a ϕ is called pathwise differentiable, and that is the entry ticket to the whole theory.

Go deeper

In the nonparametric model the tangent space is every square-integrable mean-zero function, so the gradient is unique: the efficient influence function of the mean is z − μ with no alternatives. In a smaller model several gradients agree on the tangent space, and the efficient one is their projection onto it (Build a Canonical Gradient draws this).

Step 9

The seesaw was a path all along

Adding a droplet of mass ε at z₀ is itself a way to move the distribution: shrink p by (1 − ε) and put ε at z₀. Move the spike and watch the mean move by exactly ε·(z₀ − ψ).

So the seesaw pull and the pathwise speed are the same ϕ. Put mass 1/n at each of n observations and the estimate's error is, to first order, the average of ϕ over the sample.

Go deeper

For a discrete mass function with p(z₀) > 0, the contamination path Pε = (1 − ε)P + εδz₀ has score h(z) = 𝟙{z = z₀}/p(z₀) − 1, and E[ϕh] = ϕ(z₀). For a continuous density it is only a formal device (a Gateaux derivative), not a regular path with an L² score. For the mean, the error of the sample mean is exactly Pₙ(Z − μ). For a general asymptotically linear estimator the leading error is Pₙϕ under additional conditions; Pₙ − P is a signed empirical perturbation, not one small point mass.

In one sentenceThe influence function at z is the pull of a little extra mass at z. Averaged over a sample, it is the first-order error of any regular asymptotically linear estimator of this parameter.
Step 10

The same pull for the ATE, piece by piece

A treated patient's outcome is 1 unit above m₁(x₀). At an x₀ where the chance of treatment g(x₀) is 0.2, compared with a place where it is 0.8, how hard does this one observation pull the ATE through its own regression?

Now one observation is (x₀, a₀, y₀): covariate, arm, outcome. The target is the average treatment effect, ψ = average over x of m₁(x) − m₀(x), where ma(x) is the mean outcome in arm a at x.

Add a little mass at that observation and it pulls ψ two ways. First, x₀ gets a little more weight, which moves ψ by m₁(x₀) − m₀(x₀) − ψ. Second, the regression for its own arm is pulled toward y₀, by the residual divided by the chance of that arm at x₀.

ϕ(x₀, a₀, y₀) = [m₁(x₀) − m₀(x₀) − ψ] + (2a₀ − 1) · (y₀ − ma₀(x₀)) / P(A = a₀ | x₀)

Move x₀ to where the chosen arm is rare (a control at high x, say). The second piece grows, because one rare observation stands in for many. That is where inverse-propensity weights come from, and why positivity matters.

P(A = a₀ | x₀) is g(x₀) for a treated observation and 1 − g(x₀) for a control. The second piece is the correction term in AIPW.

In one sentenceOne observation pulls the ATE through where its x sits (m₁ − m₀ − ψ) and through its own arm's regression (the residual over the chance of its arm). Their sum is the efficient influence function of the ATE, found by asking what one observation does.

Where this leads: a parameter changes along any path at E[ϕ h], so ϕ is the one object that gives the first-order effect of every allowed change to the data. How ϕ also sets the best precision any estimator can reach comes later, in Build a canonical gradient and Efficiency Theory, Drawn.

Go deeper: the hardest direction and the efficiency bound

The hardest direction

Some directions tell you a lot about ψ and some tell you nothing. Along one path the problem has a single unknown, ε, so it has an ordinary variance floor.

Rotate h away from ϕ and watch that floor. It is highest when h points along ϕ, and there it equals Var(ϕ).

No estimator can beat the floor on the hardest path, so none can beat Var(ϕ) overall. That is the efficiency bound, and the hardest path is called the least favorable one.

Go deeper

Along a path with score h the problem is parametric in ε, so Cramér–Rao applies: no regular estimator can have asymptotic variance below (dψ/dε)²/I, with Fisher information I = E[h²]. Substituting dψ/dε = E[ϕh] gives E[ϕh]²/E[h²], which by Cauchy–Schwarz is at most E[ϕ²], with equality when h ∝ ϕ.

bound along h = (E[ϕh])² / E[h²] ≤ E[ϕ²] = Var(ϕ), equality iff h ∝ ϕ

TMLE fluctuates along a submodel whose score is the efficient influence function for this reason: it is the direction in which the data are least informative about ψ per unit of likelihood.

In one sentenceEvery path gives a variance floor of Var(ϕ)·cos²θ; the largest, Var(ϕ), is the efficiency bound, reached on the path whose score is ϕ itself.

What is exact, simulated and schematic: every density number is an exact sum on a fine grid, the seesaw uses the course's 100 cohort outcomes exactly, the ATE step uses a known law, and the model sheet in step 3 is schematic.

Further reading: the construction follows Tsiatis, Semiparametric Theory and Missing Data (ch. 3–4), and Schuler & van der Laan, Introduction to Modern Causal Inference (§3.1–3.3). The point-mass derivation of the ATE influence function is the Gateaux-derivative route in Hines et al. (2022), "Demystifying statistical learning based on efficient influence functions."