Regularisation, and the answer that is chosen

An estimate that shares the dip's luck

GCV's interior dips form where the residual per remaining degree of freedom is low, and the proposal was a guard that estimates the noise from the least-squares residual and refuses a minimum whose residual is implausibly small. Over 960 draws holding 45 dips it refuses none, and picks plain GCV's minimum on every draw. At the dip ρ is 0.84 with the true noise and 0.98 with the estimate, because the estimate is made from the residual at the dip's own end of the scale and inherits its luck. Told the true noise instead, the guard can refuse 14 dips only by refusing 78 real minima. Handed the same estimate, the discrepancy principle misses by 570,000 times.

Worth reading first: When the answer is a choice · Counting what cannot be looked at · A parameter that counts steps.

More samples take the floor and leave the dip gave generalised cross-validation more samples than unknowns and watched half of its failures go. The minimum at the floor of the λ scale, a zero divided by a zero on a square system, disappeared at every ratio. The interior dips did not: 20, 13, 8 and 4 draws in 240 at 1.25, 1.5, 2 and 4 samples per unknown had a global minimum to the left of the real one, and at four per unknown one of them still put plain GCV at 877 times the best error available.

The essay traced a dip to one number. Write ρ for the residual per remaining degree of freedom in units of the noise variance,

ρ=∥r∥2δ2(m−t),\rho = \frac{\lVert r\rVert^2}{\delta^2 (m - t)},

where δ is the noise per sample and m−tm - t is GCV’s trace term. Admitting one more direction lowers GCV when its noise coefficient’s square exceeds 2ρ2\rho, so a stretch of λ where ρ is low is a stretch where luck can build a dip. That needs δ, which GCV does not have — but “with m>nm > n, the least-squares residual … is an estimate of the noise with m−nm - n degrees of freedom. A rule that estimated δ from it, computed ρ at each candidate minimum and refused a dip when ρ was implausibly low would be GCV told the noise by its own data. Whether it removes the four survivors at four samples per unknown without becoming the discrepancy principle is the measurement.”

GCV's function for sixteen draws of 96 unknowns sampled at 384 points, at 0.01% noiseEach curve is one draw's GCV function of λ divided by its value at the draw's rightmost local minimum, on logarithmic axes, for 96 unknowns and 384 samples — 4 per unknown. 2 of sixteen draws have a deeper minimum to the left of the rightmost, at λ of ten to the -4.20 and -6.93, where plain GCV's error is 2.052 and 876.7 times the oracle's. The function's value at the smallest λ on the scale is at least 1.00 times its value at the rightmost minimum.4 per unknowndraws with a dip2left end, lowest1plain GCV ÷ oracleworst87710⁻⁸10⁻⁷10⁻⁶10⁻⁵10⁻⁴10⁻³10⁻²10⁻¹110¹10²10³λGCV ÷ at the rightmost minimumno deeper minimuma dip to the leftdashed: the rightmost minimum's own valuefilled dots: where plain GCV picks
Fig. 1 The previous essay’s picture of the survivor: GCV over its value at the rightmost minimum for sixteen draws on 96 unknowns sampled at 384 points, at 0.01% noise, the two draws with a deeper minimum to the left in the second colour.

The survivor is the draw whose curve falls away four decades to the left of its real minimum, at 10−6.9310^{-6.93}, and at four samples per unknown it is the only one of 240 that costs more than ten times the oracle. Every curve’s left end sits above its minimum, so nothing about the shape at the floor of the scale gives it away; the dip is an ordinary-looking minimum that happens to be deeper. That is what made a test on the residual attractive: if the dip’s shape is unremarkable, the size of what it leaves unexplained might not be.

It removes none of them, and it does not become the discrepancy principle either. It becomes GCV.

The guard

The estimate is the classical one. The least-squares solution leaves a residual orthogonal to the range of the m×nm \times n operator, and if the data are signal in the range plus white noise, that residual is pure noise in m−nm - n dimensions:

δ^2=∥rLS∥2m−n.\hat\delta^2 = \frac{\lVert r_{\text{LS}}\rVert^2}{m - n}.

It comes from the same singular value decomposition GCV uses, at no cost. With it, ρ^\hat\rho can be computed at every λ. An honest residual has ρ near one, scattered by the chi-squared spread of an average over m−tm - t terms divided by another over m−nm - n; the guard refuses a local minimum of GCV whose ρ^\hat\rho is below the lower 1% point of that spread in its normal approximation, 1−2.332/(m−t)+2/(m−n)1 - 2.33\sqrt{2/(m-t) + 2/(m-n)}, and takes the deepest local minimum it does not refuse. If it refuses all of them, it falls back to the rightmost.

The problems are the previous essay’s: a continuous Gaussian blur sampled at mm points of a signal represented on nn nodes, five grids from 24 to 96 unknowns, noise at 1%, 0.1% and 0.01% of the data’s root mean square, sixteen seeded draws of each, 240 draws at each of four ratios of samples to unknowns. Each draw is solved with Tikhonov at 121 values of λ, and every rule’s choice is scored by its error over the best error on the same draw.

None refused

The figure at the top of the page is every interior dip in the 960 draws, 45 of them, placed by ρ at the dip computed two ways: horizontally from the true noise variance, vertically from the estimate.

The proposal’s premise holds on the horizontal axis. With the true noise, ρ at a dip has a median of 0.84, and 18 of the 45 sit between 0.43 and 0.75 — the residual there really is smaller than its degrees of freedom should carry, which is what lets a lucky direction pull GCV down. On the vertical axis the same 45 dips have a median ρ^\hat\rho of 0.98, and only five are below 0.9. The estimate has moved every dip to one.

So the guard refuses none: zero of 20 dips at 1.25 samples per unknown, zero of 13 at 1.5, zero of 8 at 2 and zero of 4 at 4, at a threshold of 2.33 standard deviations and at 1.28 and 3.09 as well. It never refuses the rightmost minimum either, and at 2.33 its pick is plain GCV’s on all 960 draws. The 877-times draw is still 877 times.

The left end is the estimate’s own residual

ρ along λ on the worst interior dip at 4 samples per unknown, from the estimated and from the true noise, with the floor below which the guard refuses96 unknowns sampled at 384 points, noise 0.0001 of the data's root mean square, the draw on which plain GCV's error is 877 times the oracle's. At the dip, λ = 10^-6.93, ρ̂ is 0.98 against a floor of 0.74, and ρ from the true noise 0.81 against 0.82. At the rightmost minimum, 10^-2.93, ρ̂ is 1.04 and ρ 0.86. At the floor of the scale ρ̂ is 0.977. Floors below zero are drawn at zero.worst dip, 4 per unknownρ at the dip, estimated0.98ρ at the dip, true noise0.81true noise floor at the dip0.8210⁻⁸10⁻⁷10⁻⁶10⁻⁵10⁻⁴10⁻³10⁻²10⁻¹00.511.522.5λresidual per remaining degree of freedomdiprightmost minimumρ, estimated noiseρ, true noisedotted: the floor each would refuse belowthe left end is the estimate's own residual
Fig. 2 ρ along λ on the worst interior dip at each ratio, from the estimated noise and from the true one, each with the floor below which the guard would refuse; the dip and the rightmost minimum are marked. The dial sets the number of samples per unknown.

The reason is visible on any one draw. At four samples per unknown, the worst dip has 96 unknowns sampled at 384 points. Walk λ from the right: past the rightmost minimum at 10−2.9310^{-2.93} the residual is honest, and as λ falls towards the floor of the scale the filter admits every direction and the residual becomes the least-squares residual. That residual is, by definition, the one δ^\hat\delta was computed from:

ρ^(λ→0)=∥rLS∥2δ^2 (m−n)=1\hat\rho(\lambda \to 0) = \frac{\lVert r_{\text{LS}}\rVert^2}{\hat\delta^2\,(m - n)} = 1

exactly, whatever the noise did. On this draw ρ^\hat\rho at the floor of the scale is 0.98, the dip at 10−6.9310^{-6.93} reads 0.98, and the plateau between them is flat. The dip lives near the least-squares end of the scale — that is what makes it a dip, a minimum reached by admitting nearly everything — and near that end the estimate is measuring the residual against itself.

With the true noise the same walk reads 0.81 at the dip: low, as the proposal said. On this draw the noise in the 288 directions outside the range came out three per cent small, and every residual near the least-squares end inherits that. The estimate inherits it too. Turn the dial to 1.25 samples per unknown: the worst dip there, 1.77·10⁴ times the oracle, has ρ of 0.62 from the true noise and 0.98 from the estimate, because its 12 degrees of freedom outside the range came out 39 per cent small and the estimate is 0.61 of the truth.

That is the whole mechanism, and the next figure shows it is not one draw’s.

The noise estimated from the least-squares residual over the true noise, every draw at four ratios of samples to unknowns, the interior dips marked240 draws at each ratio, on a logarithmic axis. 1.25 per unknown: from 0.60 to 17.55, median 1.08, median on the dips 0.84; 1.5 per unknown: from 0.71 to 14.25, median 1.03, median on the dips 0.89; 2 per unknown: from 0.76 to 10.47, median 1.02, median on the dips 0.93; 4 per unknown: from 0.94 to 8.60, median 1.01, median on the dips 0.97. The largest overestimates are on the coarsest grid at the lowest noise, where the least-squares residual carries the discretisation's misfit as well as the noise.estimated ÷ true noise1.25 per unknown, dips' median0.841.5 per unknown, dips' median0.892 per unknown, dips' median0.934 per unknown, dips' median0.97110¹samples per unknownestimated ÷ true noise1.251.524large dots: the draws with an interior dipthe dips are the draws whose estimate came out low
Fig. 3 The noise estimated from the least-squares residual over the true noise, for every draw at each ratio; the draws with an interior dip are the large dots.

The dips are the draws whose estimate came out low. At 1.25 samples per unknown the median estimate over all 240 draws is 1.02 of the truth and over the 20 dips 0.84; at 1.5, 0.89 on the dips; at 2, 0.93; at 4, 0.97. A dip forms where the residual near the least-squares end is small for its degrees of freedom, and the least-squares residual is the residual near the least-squares end, so a draw that is lucky enough to dip is a draw whose estimate is low by the same luck. Dividing one by the other cancels exactly the thing the guard was meant to see. The estimate is unbiased over all draws and biased on precisely the draws that matter — the same selection, in a different costume, that a spread measured on probes it does not average removed from a trace estimator by keeping the spread’s probes apart from the mean’s.

The figure has a second population, high above the rest: estimates eight to seventeen times the truth. Those are the coarsest grid, 24 unknowns, at the lowest noise level. There the least-squares residual is not noise at all. The data are a continuous blur, and 24 nodes cannot represent it, so the part of the data outside the range is mostly the discretisation’s misfit, which at 0.01% noise is larger than the noise. An estimate made from the residual is an estimate of everything the model does not fit, which is the noise only when the model fits everything else. That error is harmless to the guard — an overestimate makes ρ smaller everywhere, and the guard refuses a minimum on none of those draws — but it is the reason the estimate cannot be read as “the noise” without knowing the model.

Told the true noise, still no separation

A guard on ρ as its threshold moves: GCV's interior dips refused against real minima refused, told the true noise and told the estimateOver 960 draws holding 45 interior dips. Told the true noise, at thresholds of 1.28, 2.33 and 3.09 standard deviations the guard refuses 14, 3, 0 dips and 78, 3, 0 real minima. Told the estimate, 0, 0, 0 dips and 1, 0, 0 real minima.true noise, of 45 dips1.28: dips refused141.28: real minima refused782.33: dips refused32.33: real minima refused33.09: dips refused03.09: real minima refused0020406080010203040real minima refusedinterior dips refused1.282.333.09every diptold the true noisetold the estimatelabels: the threshold in standard deviationsno threshold separates them
Fig. 4 The guard told the true noise variance, at thresholds of 1.28, 2.33 and 3.09 standard deviations: interior dips refused against real minima refused, over all 960 draws; the guard told the estimate is the point at the origin.

The obvious rescue is a better estimate. Give the guard the exact noise variance — the best any estimate could do — and its floor is 1−z2/(m−t)1 - z\sqrt{2/(m - t)} on the trace term alone. At the 1% threshold it refuses 3 of the 45 dips and 3 real minima. At 10%, 14 dips and 78 real minima. At 0.1%, nothing.

The dips sit near the least-squares end, where m−tm - t is close to m−nm - n: twelve degrees of freedom on the 48-unknown grid at 1.25 samples per unknown, twenty-four on 96 unknowns. The chi-squared spread on twelve degrees of freedom puts one honest residual in ten below 0.53, which is lower than most dips’ ρ. The dips are low, but they are low by less than the noise in ρ itself on the few directions left at that end of the scale. On the 1.25 worst draw the true-noise floor at the dip is 0.11 against a ρ of 0.62. And ρ at the real minimum, where many more directions remain, is not low at all on dip draws — 1.02 median at 1.25, 0.91 at 4 — so it is not a signature either. The dip’s ρ is the signature, and it is a signature whose noise is as large as the signal.

So the guard fails twice over: the estimate cancels the low ρ it was meant to see, and even with the true noise the low ρ is inside the spread of an honest one. The previous essay’s account of a dip is confirmed — a low ρ near the least-squares end — and its proposed detector is refuted by the same account: a quantity estimated on few degrees of freedom is a weak test of anything.

Handed the estimate, the discrepancy principle breaks

The worst draw of 240 at each ratio of samples to unknowns, error over the oracle's, for GCV, GCV with the ρ̂ guard, its rightmost minimum and the discrepancy principle told the estimated and the true noiseOn a logarithmic axis. plain GCV: 1.77·10⁴, 458, 60.1, 877; GCV with the guard: 1.77·10⁴, 458, 60.1, 877; discrepancy, told the estimate: 5.7·10⁵, 1.39, 5527, 1.18; rightmost minimum: 3.7, 2.98, 2.68, 2.57; discrepancy, told the noise: 1.16, 1.16, 1.15, 1.16 at 1.25, 1.5, 2 and 4 samples per unknown.worst of 240, over the oracleplain GCV, at 1.251.8·10⁴GCV with the guard, at 1.251.8·10⁴discrepancy, told the estimate, at 1.255.7·10⁵rightmost minimum, at 1.253.7discrepancy, told the noise, at 1.251.2110¹10²10³10⁴10⁵10⁶samples per unknownerror ÷ oracle's, worst draw1.251.524plain GCVGCV with the guarddiscrepancy, told the estimaterightmost minimumdiscrepancy, told the noisethe guard's line is plain GCV'sthe estimate helps the rule it was not meant for, sometimes
Fig. 5 The worst of 240 draws at each ratio, error over the oracle’s, for plain GCV, GCV with the guard, the rightmost local minimum, and the discrepancy principle told the estimated and the true noise.

The question’s other half was whether the guard would “become the discrepancy principle” — which is told a noise level and uses it for everything. The guard became nothing; its line is plain GCV’s at every ratio, 1.77·10⁴, 458, 60 and 877 times the oracle. But the estimate can be handed to the discrepancy principle directly, and the result is worth knowing because it is the obvious next idea and it is worse.

Told the true noise, the discrepancy principle’s worst draw is within 1.16 of the oracle at every ratio, as it has been on every grid in this series. Told the estimate, its worst draw is 5.7·10⁵ times the oracle at 1.25 samples per unknown, 1.39 at 1.5, 5,527 at 2 and 1.18 at 4. The catastrophes are on the fine grids, 48 unknowns and more, on the draws whose estimate came out 12 to 40 per cent low. Thirty-two coefficients instead of a noise level measured the same cliff on square systems: told too little noise, the discrepancy principle does not degrade, it falls past the best truncation into the directions that amplify noise, and on a fine grid those directions amplify by thousands. An estimate on m−nm - n degrees of freedom is too small by a fifth on a few draws in a hundred, and that is enough.

The rule that wins is the one that reads no noise level at all. The rightmost local minimum’s worst draw is 3.70, 2.98, 2.68 and 2.57 times the oracle at the four ratios, and its median is GCV’s. The minimum on the right proposed it to remove the degenerate minimum of square systems; here it also removes every interior dip, because a dip is by construction to the left of the rightmost minimum, and it needs no estimate to do it.

What the residual cannot be used for

The general lesson is about where an estimate comes from. A noise level read from the least-squares residual is a fine number to report and a poor one to decide with, at least when the decision is about the same end of the scale the residual comes from. It shares the residual’s luck with every candidate λ near that end. It is an estimate of everything outside the model’s range, not only the noise. And on a mildly oversampled system it has few degrees of freedom — six on the 24-unknown grid at 1.25 samples per unknown — so its own spread is as large as the effects it would be used to detect.

The data count their dimensions, not the steps found the discrepancy principle within 16 per cent of the oracle on every grid when told the noise, and GCV better on the median draw with thousand-fold failures. This essay adds the missing corner of that comparison: the discrepancy principle told an estimate of the noise made from the same data inherits GCV’s thousand-fold failures without its median, and the cheapest robust choice from the data alone remains GCV restricted to its rightmost minimum.

One draw in twenty measured GCV’s catastrophes as a tail — four to six draws in a hundred, at a thousand draws a level, while the median stayed among the best — and that is the shape every comparison here has had. The guard does not change GCV’s median, because it changes no pick; the discrepancy principle told the estimate moves its median by three to four per cent and buys a tail of its own. Choosing without knowing found the same split on the first problem the rules were set: GCV on the oracle’s λ in the median, the discrepancy principle six per cent off and steady. A rule’s median is set by how it reads the typical draw, and its tail by what it does on the draw where its one piece of information is wrong. An estimate made from the data is wrong exactly on the draws where the data are unusual, which are the draws the tail is made of.

What these draws do not show

One operator, a Gaussian blur, on one signal; noise that is white; five grids and three levels. An estimate of the noise that does not share the residual’s luck — repeated measurements, a held-out set of samples, or a noise level known from the instrument — is a different object and is not measured here. A blur with a longer tail would leave fewer directions at the least-squares end and move the dips; a coloured noise, which noise that spares the answer and fools the rules found fooling both rules on square systems, would bias the estimate in a direction that depends on its spectrum.

Still open: an estimate from samples held out, and the misses on the right

A held-out estimate. The estimate failed because it was made from the residual the dip lives in. Split the samples instead: fit on three quarters, and estimate the noise from how well that fit predicts the quarter it never saw. That estimate does not share the dip’s luck, because the held-out samples’ noise is independent of the fitted samples’ — the same separation that made Stein’s two-stage rule calibrated. The prediction with a sign is that a guard told a held-out estimate refuses the dips about as often as the guard told the true noise — 3 of 45 at the 1% threshold — and so still fails, because the second failure here, ρ’s own spread on few degrees of freedom, is not about the estimate at all; and that the discrepancy principle told a held-out estimate loses its catastrophes only where the held-out set has more degrees of freedom than m−nm - n.

The misses on the right. Two or three draws at every ratio are missed by the rightmost minimum by two to four times, and nothing here touches them. The proposal to read them through the mean square of the coefficients near the chosen λ stands, now with a warning attached: whatever noise level it is compared against must not be read from the same residual.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Discrepancy principleDiscretisationGeneralised cross-validationInfluence matrixLeast-squaresParameter choiceRegularisationTikhonov regularisation