Least squares, and the road not to take

Two observations that hide each other

Two observations at the same place, wrong by the same amount, each look harmless when deleted alone, because a fit without one still has the other. Single deletion sees the shared error cut by (1 − 2h)/(1 − h) — measured at 261 times at the far end — and only the pair's two-by-two block of the hat matrix says what the two of them hold.

Worth reading first: Influence is decided before the data · The projection and the right angle.

Every influence diagnostic on a regression report is computed one observation at a time. The deleted residual asks what the fit would have predicted at observation i had observation i been left out. Cook’s distance asks how far the whole fitted curve moves when it is. The leverage asks how much of the fit’s budget of influence observation i holds, and the essay on that budget measured it as a property of the design, fixed before any response arrives and summing to the number of columns.

None of the three is asked about a set of observations, and the smallest set is two. Take two observations recorded at the same setting and wrong by the same amount — a measurement repeated within one faulty run, two units drawn from one bad batch, a row duplicated on entry. Delete one of them and the fit still contains the other, sitting at the same place and carrying the same error, and the fit leans towards it exactly as it leaned before. The deleted residual of the observation removed is then measured against a line its twin is still holding up. Each twin vouches for the other, and a diagnostic that questions them one at a time is asking each for a reference from its accomplice.

How much that hides is not a matter for judgement. It is a closed form in one number, and it holds to six significant figures at every position measured here.

Each twin covers for the other

The experiment is small on purpose. Thirty observations are scattered over [−1, 1] around the line 1 + x/2 with noise of standard deviation 0.1. Two more are placed together at x = 6, and both are raised by 3 above the line. The model is a straight line, so there are two columns, two units of influence, and thirty-two observations to share them.

The figure at the head of the page is that fit. Each twin has leverage 0.443, against an average of 2/32 = 0.0625, so the leverage column does mark them: each is seven times the average. The rest of the report does not. Fitted with all thirty-two observations, the line passes close to the pair, missing it by 0.34, because two observations holding most of a unit of influence between them get most of what they ask for. Fitted without one twin, the line still passes close to the other and predicts the removed one off by 0.62 on average — 0.54 for one and 0.69 for the other. Fitted without both, it follows the thirty ordinary observations and predicts the pair off by 3.01. The error the pair actually carries is 3.

So single deletion reports about a fifth of the shift that is there, on each observation, and it reports it on the observations whose influence is largest. The two refits agree with the closed forms below to better than 10⁻¹⁰, which is two routes to one number applied to a shortcut: the deleted residuals are not an artefact of the formula used to compute them.

The algebra says why, and it is short. The residuals of a set S of observations, predicted by a fit made without the whole set, are

(IHSS)1eS(I − H_{SS})^{-1}\, e_S

where eSe_S is the ordinary residuals of the set and HSSH_{SS} is the block of the hat matrix whose rows and columns belong to it. For a single observation the block is the number hih_i and the formula is the familiar ei/(1hi)e_i/(1 - h_i). For two identical rows the block is h times the two-by-two matrix of ones, and its eigenvalues are 2h, along the direction in which the two residuals move together, and 0, along the direction in which they move apart. A shift the twins share lies along the first direction, so joint deletion divides it by 1 − 2h and recovers it whole. Single deletion divides by 1 − h. What single deletion sees, as a fraction of what joint deletion sees, is (1 − 2h)/(1 − h), and at h = 0.443 that is one part in 4.86.

A pair can hold a whole unit, and a single row cannot say so

The hat matrix is the orthogonal projection onto the column space, and the eigenvalues of any principal block of a projection lie between zero and one, so the larger eigenvalue 2h can approach one and never pass it. Two identical rows can therefore hold at most one unit of influence between them, and each can hold at most half. That is the part of the arithmetic that makes the pair invisible rather than merely understated.

Single deletion divides by 1 − h, and for a twin 1 − h never falls below one half. The divisor that would expose an observation — the one that grows without bound as an observation takes over a direction of the fit — cannot become small for either twin, however far out the pair is placed, because neither twin ever holds more than half. The quantity that does approach its limit is the block’s larger eigenvalue, 0.885 at x = 6, and it is not on the diagonal of the hat matrix at all. It is a property of the pair, read from the off-diagonal entry that joins them.

The leverage column is caught the same way. A twin at h = 0.443 reads, on its own, as an observation holding somewhat less than half a unit — a large leverage, and a finite one. The same row read as a member of its pair holds 0.885 of a unit jointly with one other row, which is a direction of the fit that two observations nearly own. The diagonal cannot tell those two situations apart, since it holds one number per row and the distinction is about two rows at once. The budget of influence is still exactly two units; what it cannot report is that one of them has been spent on a single place by two rows together.

A pair of observations at x = 2, both raised by 3, and the straight lines fitted with and without themThirty observations on [−1, 1] and two more at x = 2, all near the line 1 + x/2 except that both of the pair are raised by 3. Fitted with every observation, the line misses the pair by 1.55. Fitted without one of them, it is pulled towards the other, and predicts the removed one off by 2.05. Fitted without both, it follows the thirty and predicts the pair off by 3.04. Each of the pair has leverage 0.245; together they hold 0.489 of the two units of influence.-101212345xyall 32without one twinwithout boththe pair at x = 2leverage of each0.24pair's joint share0.49deleted alone, mean2.1deleted together, mean3each twin covers for the otheronly removing both shows the shift
Fig. 1 The pair moved in to x = 2, where each twin has leverage 0.245 and the two hold 0.489 of a unit between them. The full fit misses the pair by 1.55, a fit without one twin predicts it off by 2.05, and a fit without both by 3.04.

Close in, the pair is still caught

Masking is a matter of degree before it is a matter of kind, and near the ordinary observations it is harmless.

At x = 2 each twin holds 0.245 and the pair’s block has a larger eigenvalue of 0.489 — half a unit, shared. The full fit misses the pair by 1.55, a fit without one twin predicts it off by 2.05, and a fit without both by 3.04. Two thirds of the shift is visible to single deletion, against a noise level of 0.1, and any procedure that deletes observations one at a time and looks at the largest deleted residuals finds both twins at the top of the list. At x = 1.5 the fraction visible is higher still: the masking factor is 1.29, each twin’s own deleted residual is about 2.36, and Cook’s distance reads 1.23 and 1.34 for the two of them against 0.052 for the most influential of the thirty.

Nothing in that regime is being hidden. The single diagnostics understate the pair, and they rank it correctly. The failure the algebra predicts needs the pair to hold most of a unit, and at x = 2 it holds under half.

The pair's deleted residuals, one at a time and together, as the pair moves out along xFor a pair of observations both raised by 3, placed at x from 1.5 to 50 beside thirty ordinary observations on [−1, 1], on logarithmic axes: the mean deleted residual of the pair with each removed alone, the mean with both removed together, and the largest deleted residual among the thirty. Together the pair's deleted residual stays between 2.7 and 3.04. Alone it falls from 2.36 to 0.0103, and from x = 10 it is smaller than the largest ordinary one.10¹10⁻²10⁻¹1position of the pair, xdeleted residualboth removedlargest ordinaryone removedthe dashed line is the shift the pair sharessingle deletion loses it as the pair's share nears one
Fig. 2 The pair’s deleted residuals as it moves out along x from 1.5 to 50: deleted together they stay between 2.70 and 3.04, while deleted alone they fall from 2.36 to 0.0103, and from x = 10 onwards each is below the largest deleted residual among the thirty ordinary observations.

Moved out, the pair ranks below the ordinary observations

The same pair, raised by the same 3, walked out along x, gives the figure above, and it has two regimes separated by one crossing.

Deleted together, the pair’s residual never strays far from the shift it carries: between 2.70 and 3.04 across a thirty-fold range of positions, never more than a tenth away. The small drift at the far end is the ordinary cost of predicting fifty units from a line fitted on [−1, 1], and it is the same size whether or not the pair is masked.

Deleted alone, the pair disappears. At x = 3 each twin’s own deleted residual is about 1.5; at x = 6 it is 0.62; at x = 10 it is 0.26 and — the crossing — smaller than the largest deleted residual among the thirty ordinary observations. From there on a list of observations ranked by deleted residual puts both twins below an ordinary point. At x = 20 the mean is 0.07, and at x = 50 it is 0.0103, a residual smaller than a tenth of the noise, on two observations each wrong by thirty times the noise.

Cook’s distance holds out longer, because it multiplies by h/(1 − h)² and so rewards exactly the leverage the twins have. At x = 6 it reads 0.86 and 1.41 for the twins against 0.060 for the largest ordinary observation, which is a clean flag. At x = 20 it reads 0.004 for one twin and 0.42 for the other, against 0.072 — one twin has fallen below every ordinary observation and the other is flagged only because the noise happened to separate their residuals. At x = 50 it reads 0.18 and 0.29 against 0.116, a factor of two and a half above an unremarkable point. The diagnostic built to combine leverage and residual is defeated by the same mechanism more slowly, because the leverage it multiplies by can reach one half and no further.

A pair of observations at x = 20, both raised by 3, and the straight lines fitted with and without themThirty observations on [−1, 1] and two more at x = 20, all near the line 1 + x/2 except that both of the pair are raised by 3. Fitted with every observation, the line passes close to the pair, missing it by 0.03. Fitted without one of them, it still passes close to the other, and predicts the removed one off by 0.07. Fitted without both, it follows the thirty and predicts the pair off by 2.91. Each of the pair has leverage 0.494; together they hold 0.988 of the two units of influence.-1491419123456789101112131415xyall 32without one twinwithout boththe pair at x = 20leverage of each0.49pair's joint share0.99deleted alone, mean0.068deleted together, mean2.9each twin covers for the otheronly removing both shows the shift
Fig. 3 At x = 20 the full fit misses the pair by 0.03 and a fit without one twin predicts it off by 0.07. Each twin holds 0.494 and the two hold 0.988 of a unit together; a fit without both predicts the pair off by 2.91.

That panel is a small residual that is not a small error in its purest form. The full fit misses each of two observations wrong by 3 by three hundredths, and it does so because they bought their own agreement: two rows holding 0.988 of a unit of influence between them get a line that goes where they are. A residual plot shows nothing, a deleted residual shows almost nothing, and the one quantity that shows the whole shift requires deleting the right two observations together.

The factor is (1 − h)/(1 − 2h), to six figures

The masking factor is the ratio of what joint deletion sees to what single deletion sees, and the closed form predicts it from the twins’ leverage alone.

How much a pair hides from single deletion, against how much of a unit of influence it holdsFor each position of the pair, the ratio of its joint deleted residual to its single deleted residual, against one minus the pair's joint share of influence — the larger eigenvalue of its two-by-two block of the hat matrix — on logarithmic axes. The curve is (1 − h)/(1 − 2h) for identical twins of leverage h. The measured ratios are 1.29, 1.48, 2.01, 2.75, 4.86, 11.6, 42.9, 261, each equal to the curve to six figures.10⁻²10⁻¹1110¹10²1 − the pair's joint share of influencejoint ÷ single deleted residual(1 − h)/(1 − 2h)x = 1.5, ratio1.3x = 6, ratio4.9x = 50, ratio261the pair's share approaches one unitand single deletion sees less and less of it
Fig. 4 The ratio of joint to single deleted residual at each of the eight positions, against one minus the pair’s joint share of influence, beside the curve (1 − h)/(1 − 2h). The measured ratios run 1.29, 1.48, 2.01, 2.75, 4.86, 11.6, 42.9 and 261, each on the curve to six figures.

The eight measured ratios are 1.29, 1.48, 2.01, 2.75, 4.86, 11.6, 42.9 and 261, at x = 1.5, 2, 3, 4, 6, 10, 20 and 50, and each agrees with (1 − h)/(1 − 2h) to six significant figures. That is closer than a measurement on noisy data has any right to be, and the reason is worth a sentence. The two twins’ ordinary residuals differ, because their noise differs. But the difference between them lies along the direction in which the pair’s block has eigenvalue zero, where joint and single deletion divide by the same thing, and the mean of the two lies along the direction with eigenvalue 2h. So the averaged residuals obey the formula exactly whatever noise was drawn, and the six figures are the arithmetic’s rather than the experiment’s.

The horizontal axis is the more useful reading. One minus the pair’s joint share runs from 0.633 at x = 1.5 to 0.0019 at x = 50, and the masking factor is, to leading order, one half divided by it. A pair that holds all but a tenth of a unit hides all but a fifth of its error from single deletion; a pair that holds all but a hundredth hides all but a fiftieth. The pole is at a joint share of one — at the point where two observations own a direction of the fit between them, which is precisely where neither of their diagonal leverages says anything alarming.

The pair’s influence is a matrix, and the matrix is already computed

A diagnostic for pairs sounds expensive and is not. The entry that joins two observations in the hat matrix is hij=aiT(ATA)1ajh_{ij} = a_i^{\mathsf T}(A^{\mathsf T}A)^{-1}a_j, computed by the same triangular solves that produce the diagonal, and the larger eigenvalue of the block [hihijhijhj]\begin{bmatrix} h_i & h_{ij} \\ h_{ij} & h_j \end{bmatrix} has a closed form. On thirty-two observations there are 496 pairs, and every one of their blocks is a lookup and a square root.

Deleting a pair from a fit is a rank-two change, and the Sherman–Morrison–Woodbury identity that makes a rank-one update cheaper than the problem makes it cheap in the same way: the refit without S needs only the inverse of IHSSI - H_{SS}, a two-by-two matrix. So the joint deleted residual is no more expensive than two single ones, and the refits in this essay are checks on it rather than the way to compute it.

The same algebra says what happens to a factorisation asked to lose the pair. Removing one twin leaves the other with a leverage, in the smaller fit, of h/(1 − h), which is 0.998 at x = 50. Removing the second twin is then a downdate at 1 − h of 0.0038 — the regime where removing an observation from a factorisation loses its accuracy in proportion to 1/(1 − h) — although neither observation, measured on the original fit, has a leverage above one half. A pair near a whole unit is a breakdown for the downdating algorithm as well as a blind spot for the diagnostic, and both are hidden from the diagonal for the same reason.

What does grow expensive is the size of the group. A group of k observations has a k-by-k block, and the number of groups of k among m observations grows as m to the k. Pairs are affordable exhaustively on any design a person would fit by hand; triples on thirty-two observations are already 4,960 blocks, and a masked group of five among a thousand rows is out of reach of enumeration. Screening pairs is a lookup; screening groups is a search.

The twins need not be identical

An argument about identical rows invites the objection that real observations are never identical, and the objection is measurable.

The two eigenvalues of the pair's leverage block as the pair is pulled apartFor the pair at x = 6 and x = 6 + ε, the two eigenvalues of its two-by-two block of the hat matrix against ε, on logarithmic axes. The larger stays between 0.885 and 0.914: the pair's joint share. The smaller rises from 4.48·10⁻⁸ at ε = 0.01 to 0.0013 at ε = 2, with a fitted slope of 1.94 — the square of the separation.10⁻²10⁻¹110⁻⁷10⁻⁵10⁻³10⁻¹separation of the pair, εeigenvalue of the pair's leverage blockjoint sharewhat tells them apartthe pair holds its share whether or not it is identicalthe difference between the twins is second order
Fig. 5 The pair at x = 6 and x = 6 + ε: the larger eigenvalue of its block stays between 0.885 and 0.914 as ε runs from 0.01 to 2, while the smaller rises from 4.48·10⁻⁸ to 0.0013 with a fitted slope of 1.94.

Pulling the twins apart by ε changes the block’s two eigenvalues in completely different ways. The larger, the pair’s joint share, stays between 0.885 and 0.914 while ε runs over more than two decades: a pair of observations near each other holds the same share of influence whether or not the two are at the same place. The smaller rises from 4.48·10⁻⁸ at ε = 0.01 to 0.0013 at ε = 2, with a fitted slope of 1.94 — the square of the separation. That is the exponent the geometry predicts: the two rows [1, 6] and [1, 6 + ε] differ by a vector of length ε, and the block’s quadratic form along the difference is a square of it.

So what distinguishes the twins from each other is a second-order quantity, and what they hold together is a first-order one. Masking needs the pair to be close compared with its distance from the rest of the design, and it does not need duplication. On a design with its bulk on [−1, 1], two observations two units apart, six units out, are close in that sense.

A pair of observations at x = 6 and x = 8, both raised by 3, and the straight lines fitted with and without themThirty observations on [−1, 1] and two more at x = 6 and x = 8, all near the line 1 + x/2 except that both of the pair are raised by 3. Fitted with every observation, the line passes close to the pair, missing it by 0.31. Fitted without one of them, it still passes close to the other, and predicts the removed one off by 0.44. Fitted without both, it follows the thirty and predicts the pair off by 3.00. The pair's leverages are 0.332 and 0.584; the two eigenvalues of their block are 0.914 and 0.0013.-10123456789123456789xyall 32without one twinwithout boththe pair at x = 6leverage of each0.46pair's joint share0.91deleted alone, mean0.44deleted together, mean3each twin covers for the otheronly removing both shows the shift
Fig. 6 The pair two units apart, at x = 6 and x = 8. Their leverages are 0.332 and 0.584 and their block’s eigenvalues 0.914 and 0.0013; a fit without one predicts it off by 0.44 on average, and a fit without both by 3.00.

The panel above is the pair at x = 6 and x = 8, and it is masked as thoroughly as the identical pair at x = 6. A fit without one twin predicts it off by 0.44 on average; a fit without both, by 3.00. What has changed is the diagonal. The twins’ leverages are now 0.332 and 0.584 — unequal, and the outer one above one half, which the identical pair could never reach. Read row by row, the report says that the observation at x = 8 is the influential one and the observation at x = 6 is a moderate neighbour. Read as a block, it says the two of them hold 0.914 of a unit, which is more than the identical pair held.

The diagonal moved by three quarters while the invariant moved by three per cent. The leverage of each row depends on how the pair’s joint share happens to be divided between its members, and that division is decided by a second-order separation; the share itself is not. It is the same distinction the valley with no bottom draws for coefficients — a direction the data barely determines is visible to the whole problem and invisible to any single parameter — drawn here along the observations instead.

What the diagonal cannot say

The practical consequences follow from one observation: influence belongs to directions of the fit, and a direction can be held by more than one row.

Screen pairs, not rows. For every pair of observations with large leverage, the larger eigenvalue of their two-by-two block is a lookup, and a value near one is two observations spending a unit of influence on one place. The rule of thumb that flags a leverage above twice the average still flags each twin here, and still cannot say that two flagged rows are one problem.

Delete in the groups the block suggests. A deleted residual computed for a pair whose block has an eigenvalue near one is the only diagnostic on this page that recovers the whole shift at every position, and it costs a two-by-two solve.

Distrust a small deleted residual on a high-leverage row. A twin at x = 50 has a deleted residual of 0.0103 and an error of 3. Where a row’s leverage is near one half, a small deleted residual is at least as consistent with a masked twin as with a good observation, and the two readings are distinguished by the off-diagonal entries of the hat matrix rather than by anything on its diagonal.

Treat a whole unit held by a group as a constraint. Two observations holding 0.988 of a unit force the fit through their common value almost as firmly as an equality would, which is the limit a constraint as a weight at infinity reaches from the other side: a row driven to h = 1 is a constraint, and a pair driven to a joint share of one is a constraint shared between two rows.

The masking factor is the one number the whole argument turns on, and it is worth restating in the form a report could print: a pair of observations with joint share s shows single deletion roughly (1 − s)/(1 − s/2) of the error it carries. At s = 0.5 that is two thirds, and nothing is lost. At s = 0.9 it is two elevenths. At s = 0.99 it is two hundredths, and the report prints two unremarkable rows.

Still open: the digits of 1 − h, and groups larger than two

Every diagnostic on this page divides by one minus something: 1 − h for a single deletion, 1 − 2h or 1 − s for a pair. Near the pole those are differences of nearly equal numbers, and a leverage computed in floating point and subtracted from one loses digits in proportion to how close to one it is. Whether the divisor can be computed without the subtraction, and what the alternative costs on a design whose far point is what made 1 − h small, is the question one minus a leverage is a subtraction measures.

The other open direction is size. Pairs were affordable here because a pair’s block is two-by-two and the number of pairs is quadratic. A masked group of three or more needs a larger block and a search over groups whose count grows as a power of the number of observations, and the measurement not yet made is how often screening pairs finds a masked triple anyway — whether the pairwise blocks of a masked group carry enough of its joint share to give it away, or whether a group can be arranged so that every pair within it looks ordinary.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Least-squaresLeverageNormal equationsOrthogonal projectionResidual