Regularisation, and the answer that is chosen

An estimate that shares nothing with the fit

GCV's guard failed because the noise estimate it read came from the residual its dips live in, and the proposal was to estimate the noise from samples held out of the fit instead. Over the same 960 draws the held-out guard refuses none of the 45 dips at the 1% threshold — fewer than the guard told the true noise, which refuses three — because a quarter of the samples gives its threshold too few degrees of freedom to refuse anything. Loosened to 10%, it refuses sixteen dips and 170 real minima. But the discrepancy principle, which missed by 570,000 times when told the least-squares estimate, never misses by more than eighteen told the held-out one, at every ratio and not only where the held-out set is the larger. The held-out estimate is the worse estimate, biased upward by up to half. What makes it the better input is that it is independent of the residual it is compared with.

Worth reading first: When the answer is a choice · Counting what cannot be looked at · A parameter that counts steps.

An estimate that shares the dip’s luck gave generalised cross-validation a guard. A dip in GCV forms where the residual per remaining degree of freedom, ρ=∥r∥2/(δ2(m−t))\rho = \lVert r\rVert^2/(\delta^2(m-t)), is low, so the guard estimated the noise level δ\delta and refused a local minimum whose ρ^\hat\rho was implausibly small. With more samples than unknowns the noise has a free estimate, the least-squares residual over m−nm - n, and the guard read it. It refused none of 45 dips in 960 draws. At the dip ρ^\hat\rho sat near one because the estimate was made from the residual at the dip’s own end of the scale, and inherited its luck; told the same estimate, the discrepancy principle missed by 570,000 times.

The essay’s proposal was the separation that had worked once already in this subject. A spread measured on probes it does not average removed a trace estimator’s bias by estimating its spread from probes it did not then average; here, fit on three quarters of the samples and estimate the noise from how well that fit predicts the quarter it never saw. The trace estimator had needed it for the same reason: a rule that reads only its own probes stopped on a spread computed from the very samples it was averaging, and its unlucky runs were the ones that stopped early. An estimate and the thing it is used to judge must not share their noise. The prediction had two parts with a sign on each. The guard would refuse the dips about as often as the guard told the true noise — 3 of 45 — and so still fail, for a reason that is not about the estimate. And the discrepancy principle would lose its catastrophes only where the held-out set is larger than the least-squares residual’s m−nm - n degrees of freedom, which happens here only at 1.25 samples per unknown.

Both parts are wrong, in opposite directions. The figure at the top is the second: the discrepancy principle told the held-out estimate stays within eighteen times the oracle’s error at every ratio. The guard, told the same estimate, does nothing at all.

Holding out a quarter

The problems are the previous essays’: a Gaussian blur of a fixed signal sampled at mm points, reconstructed on nn unknowns at 1.25, 1.5, 2 and 4 samples per unknown, five grids from 24 to 96 unknowns, three noise levels from 10−210^{-2} to 10−410^{-4} of the data’s root mean square, sixteen draws each — the same 960 draws, with the same noise. Every fourth sample is held out. The other three quarters are fitted by Tikhonov regularisation on the usual grid of 121 values of λ\lambda, each fit predicts the held-out samples, and the noise estimate is the smallest held-out mean square over the grid:

δ^h2=min⁡λ1mh∥bh−Ahxλ∥2.\hat\delta_h^2 = \min_\lambda \frac{1}{m_h}\lVert b_h - A_h x_\lambda\rVert^2 .

The held-out samples’ noise never entered the fit, so the estimate shares nothing with the residual GCV and the guard read. It pays two prices for that. It has only mh=m/4m_h = m/4 samples behind it — from 7 at the smallest problem to 96 at the largest. And it is biased upward, since even the best fit predicts imperfectly, so the estimate includes the fit’s own error on the unseen samples as well as their noise.

Both reductions in this essay then take δ^h\hat\delta_h in place of the least-squares δ^\hat\delta. The guard’s floor — the lower 1% point of ρ^\hat\rho for an honest residual, 1−2.332/(m−t)+2/mh1 - 2.33\sqrt{2/(m-t) + 2/m_h} — now carries mhm_h where it carried m−nm - n. And the discrepancy principle takes the largest λ\lambda whose residual is at most 1.01m δ^h1.01\sqrt{m}\,\hat\delta_h.

The guard refuses nothing

At every interior dip, the residual per degree of freedom read with the held-out estimate, beside the floor below which the guard would refuse it, against the number of samples held out45 dips. ρ̂ at the dip runs from 0.27 to 1.22; the floor at the 1% threshold, at the dip's own degrees of freedom, from -0.34 to 0.62, and under every dip's ρ̂; the smallest margin between them is 0.14.-0.500.511.5samples held outρ̂ at the dip, and its floor12244896ρ̂ at the dipthe 1% floora dip is refused when its ρ̂ is under its floorthe floor sits below every dip
Fig. 1 At each of the 45 interior dips: ρ^\hat\rho read with the held-out estimate, and the 1% floor at the dip’s own degrees of freedom, against the number of samples held out.

At the 1% threshold the held-out guard refuses none of the 45 dips and none of the real minima, at any ratio, and its worst draw at each ratio is exactly the least-squares guard’s: 18,000 times the oracle at 1.25 samples per unknown. The guard told the true noise refuses three dips at that threshold. The prediction expected the held-out guard to match it and it does worse.

The figure says why, and it is not the estimate’s bias. ρ^\hat\rho at the dips now spreads from 0.27 to 1.22 — the held-out estimate does not pin them at one as the least-squares estimate did — so the dips are no longer disguised. But the floor has a 2/mh2/m_h term under its square root, and with twelve to forty samples held out on most dips that term is large: the floor runs from −0.34 to 0.62 across the dips and sits under every dip’s ρ^\hat\rho, by at least 0.14. A floor below zero refuses nothing, whatever the estimate says. The held-out estimate is independent of the residual and too thin to judge it.

That is the third failure this guard has met, and the other two did not go away. The least-squares estimate shared the dip’s luck. The true noise, which has neither problem, could not separate dips from real minima because ρ\rho itself spreads widely on few degrees of freedom. And an independent estimate built from a quarter of the samples has too few of its own.

Loosening it buys the wrong trade

The guard's trade at its 10% threshold over 960 draws holding 45 interior dips: dips and real minima refused, told the true noise and told the held-out estimatedips refused, true noise: 14; dips refused, held out: 16; real minima refused, true noise: 78; real minima refused, held out: 170, of 45 dips and 960 draws.050100150200dips refused, true noise14dips refused, held out16real minima refused, true noise78real minima refused, held out170draws45 dips in 960 drawsfew held-out samples, a wide floor
Fig. 2 Over the 960 draws: dips refused and real minima refused by the guard told the true noise and told the held-out estimate. The dial sets the threshold.

Raise the threshold to 10% and the floor rises far enough to bite. The held-out guard refuses sixteen of the 45 dips, two more than the true-noise guard’s fourteen, and pays with 170 refused real minima — draws where GCV’s rightmost local minimum, the one that was right, is thrown out — against the true noise’s 78. Turn the dial to 0.1% and neither guard refuses anything. At no threshold does the held-out guard refuse more than a third of the dips without refusing several times as many minima that should have stood.

The trade is worse than the true noise’s for the same reason the 1% threshold was empty. A floor that has to be raised far enough to bite on the dips, with a 2/mh2/m_h term that large under its square root, rises over many honest minima too, and the held-out estimate’s upward bias lowers every ρ^\hat\rho and pushes more of them under it. Both effects point the same way.

So the guard is finished as a remedy for these dips, by three independent routes. The minimum on the right had already found the remedy that works: ignore the deepest minimum and take the rightmost, which never misses by more than four times the oracle here — 3.7, 3.0, 2.7 and 2.6 at the four ratios — and needs no noise level at all.

The discrepancy principle comes back

The discrepancy principle had the opposite problem. It does need the noise level, and handed the least-squares estimate it missed by more than a thousand times at 1.25 and 2 samples per unknown. Handed the held-out estimate, its worst draws are 3.5, 18, 7.5 and 1.5 times the oracle at the four ratios. Across all 960 draws one misses by more than ten, against sixteen with the least-squares estimate.

The prediction placed the rescue where the held-out set is larger than m−nm - n. At 1.25 samples per unknown it is: seven against six samples at the smallest grid, thirty against twenty-four at the largest. At 1.5, 2 and 4 samples per unknown it is the smaller by a factor of 1.3 to 3, and the rescue happens anyway — 5,500 down to 7.5 at two per unknown. The degrees of freedom were not what made the least-squares estimate dangerous.

Every draw: the noise estimated from the least-squares residual and from the held-out samples, over the true noise, against the discrepancy principle's error over the oracle's when told each960 draws. The least-squares estimate is under 0.95 of the noise on 177 draws and the discrepancy principle told it misses by more than ten times on 16, all with estimates between 0.60 and 0.88. The held-out estimate is under 0.95 on 76 draws and misses by more than ten on 1. Held-out estimates run up to 63 times the noise, least-squares ones up to 18.960 drawsleast squares: misses over ten16held out: misses over ten1110¹110¹10²10³10⁴10⁵10⁶estimated noise ÷ true noisediscrepancy error ÷ oracle'sleast-squares estimateheld-out estimatedashed: the estimate equal to the noiselow is dangerous only when it is the residual's own
Fig. 3 Every draw: each estimate over the true noise, against the discrepancy principle’s error over the oracle’s when told it.

What made it dangerous is visible here. A noise estimate that is too low is the discrepancy principle’s only route to catastrophe, because it asks for a residual smaller than the noise and gets one only by under-regularising. The least-squares estimate is under 0.95 of the noise on 177 draws; the held-out estimate on 76. That is a difference, but not the difference: the least-squares estimate’s sixteen catastrophes all come from estimates between 0.60 and 0.88, and the held-out estimate reaches 0.61 too, with one miss above ten among its 76 low draws. Low is dangerous when it is low because the residual is: the least-squares estimate is computed from the least-squares residual, the floor the residual curve descends towards, so when that floor is low by luck the target falls with it, and the two move together. A held-out estimate that is low is low for an unrelated reason, and the residual curve it is compared against does not follow it down.

A target under the whole curve

The residual along the regularisation parameter on the draw where the discrepancy principle told the least-squares estimate misses worst, with the three noise levels it could be told1.25 samples per unknown, 96 unknowns, noise level 0.01. The residual's lowest value on the grid is 0.0545. Targets: least squares 0.0514, held out 0.0837, true 0.0732. Told each, the discrepancy principle lands at log λ -8.00, -1.13 and -1.27, errors 7.85e+4, 0.145 and 0.142; the oracle is at -1.60 with 0.138.error ÷ oracle'stold least squares5.7·10⁵told held out1.1told the truth1-8-6-4-2010⁻¹1log of the regularisation parameterresidual normleast-squares targettrue noiseheld-out targetsolid: the residual · dashed: three targetsa target under the whole curve
Fig. 4 The residual along λ\lambda on the draw where the least-squares estimate’s discrepancy principle misses worst, with the three targets it could be told.

The worst draw shows the mechanism at full size. At 1.25 samples per unknown, 96 unknowns and noise 10−210^{-2}, the true noise norm is 0.072 and the least-squares estimate puts it at 0.051 — seven tenths. The residual’s lowest value anywhere on the grid is 0.054. The target is under the whole curve: no λ\lambda satisfies it, and the rule falls back to the smallest λ\lambda on the grid, where the error is 570,000 times the oracle’s. The held-out target, 0.084, crosses the curve where it rises steeply out of its floor; the discrepancy principle lands at log⁡10λ=−1.13\log_{10}\lambda = -1.13 against the oracle’s −1.60-1.60, and its error is 1.05 times the oracle’s. Told the true noise it lands at −1.27-1.27 and 1.03.

The least-squares residual and the estimate are the same number up to a factor: the estimate is that residual divided by m−n\sqrt{m-n}, and the target multiplies it by 1.01m1.01\sqrt{m}, so the target is always exactly 1.01m/(m−n)1.01\sqrt{m/(m-n)} times the least-squares residual — 2.26 at 1.25 samples per unknown — however large the noise is. On this draw the least-squares residual is 0.023, so the target is 0.051. But the residual curve does not reach the least-squares residual on the grid: even at λ=10−8\lambda = 10^{-8} it still filters the components with the very smallest singular values, and their share of the data keeps it at 0.054. The target, tied to a floor the curve never reaches, lands under the curve. Nothing ties the held-out target to that floor.

Where the estimate is made

The held-out estimate is a minimum over λ\lambda, and the λ\lambda at which it is reached is itself a choice made from data the fit did not see. Measured against the oracle’s λ\lambda for the full problem, it sits at a median of 0.07 decades above it, with half of the 960 draws within half a decade either side. The estimate is made, in other words, where the fit on three quarters is about as good as a fit can be, and its upward bias is the error that remains at the best λ\lambda — the part of the unseen samples no smooth reconstruction predicts — rather than a symptom of fitting at the wrong one.

That says two things. The bias cannot be removed by searching harder; it is the floor of prediction error on samples the fit never saw, and it grows exactly where the figure above shows it growing, at low noise, where the reconstruction’s own error is no longer small beside the noise. And held-out validation is, here, a parameter rule in its own right, the one when the answer is a choice said something outside the fitted data would have to supply. It has not been scored as one in this essay, because it chooses λ\lambda for a fit on three quarters of the data and the question was the noise level. It is the obvious next measurement.

The worse estimate is the better input

The median of each noise estimate over the true noise, held out and least squares, at three noise levels, against the ratio of samples to unknownsheld out at noise 0.01: 1.227, 1.156, 1.122, 1.087; least squares at noise 0.01: 1.020, 1.021, 1.015, 1.010; held out at noise 0.001: 1.353, 1.285, 1.181, 1.131; least squares at noise 0.001: 1.080, 1.057, 1.018, 1.017; held out at noise 0.0001: 1.542, 1.392, 1.247, 1.163; least squares at noise 0.0001: 1.090, 1.057, 1.018, 1.017 at 1.25, 1.5, 2 and 4 samples per unknown.0.911.11.21.31.41.5samples per unknownmedian estimate ÷ true noise1.251.524held out, 0.01least squares, 0.01held out, 0.001least squares, 0.001held out, 0.0001least squares, 0.0001solid 0.01 · dashed 0.001 · dotted 0.0001biased upward, and further at low noise
Fig. 5 The median of each estimate over the true noise, by ratio of samples to unknowns and noise level.

Measured as an estimate, the held-out one loses. Its median is 1.11 to 1.35 times the noise across the ratios, where the least-squares one is 1.01 to 1.08, and it is worst at the lowest noise — up to 1.5 times at 10−410^{-4} and 1.25 samples per unknown, and 63 times on its worst draw — because there the best fit’s own error on unseen samples is no longer small beside the noise. A statistician choosing between them by bias and spread would take the least-squares estimate every time.

The discrepancy principle should not, and the reason is the subject’s recurring one. An input to a rule is judged by what the rule does with it, and the discrepancy principle is catastrophic for low estimates that track the residual and merely conservative for high ones: an estimate too large by a third regularises too much and loses a little, which is why the held-out rule’s median error is 1.04 to 1.06 times the oracle against the true noise’s 1.03. That asymmetry rewards an estimate whose errors are upward and independent, and the held-out estimate is exactly that. More samples take the floor and leave the dip found the same structure in GCV — the failure that mattered was a minimum the method could fall into, not its typical accuracy — and noise that spares the answer and fools the rules found that what breaks the parameter rules is what they assume about the noise rather than the noise’s effect on the answer, which is the same lesson from the other side. The data count their dimensions, not the step’s had the discrepancy principle as the robust rule on square grids, told the noise; the whole of its trouble since has been the noise level it is told.

What a solver would do

With more samples than unknowns, the measurements give an order. To choose λ\lambda without a noise level, take GCV’s rightmost local minimum: within four times the oracle on every draw. To use the discrepancy principle, never feed it the least-squares residual’s estimate; hold out a quarter of the samples, fit the rest, and use the best held-out mean square, which costs one more SVD of three quarters of the matrix and keeps the rule within eighteen times the oracle. And do not guard GCV with any noise estimate, held out or not: the guard has now failed with an estimate that shares the residual’s luck, with one too thin to judge it, and with the truth.

The held-out estimate gives up a quarter of the data to the fit, and on problems this small that costs something the measurements do not separate: the fit on three quarters is a slightly different problem from the fit on all of them, and the discrepancy principle then applies the held-out noise level to the full fit. That was deliberate — the estimate’s job is to say how large the noise is, which does not depend on which samples are fitted — and the median errors say it did no harm.

What these draws do not show

One blur, one signal, white noise of known distribution, a single held-out quarter chosen as every fourth sample. Several folds averaged would give the estimate more degrees of freedom and might rescue the guard’s floor; the held-out set chosen at random rather than interleaved might change the fit’s error on it; and a signal whose smoothness differs from the blur’s would change the upward bias. None is measured.

Still open: folds, and the misses on the right

Four folds instead of one. Averaging the held-out mean square over four folds, each holding out a different quarter, would give the estimate four times the samples — mm in total — at the price of four fits. The prediction with a sign is that the guard’s floor then rises above the dips’ ρ^\hat\rho on about as many dips as the true-noise guard refuses, three of 45 at 1%, which would show the thin estimate rather than the guard’s idea was the failure here; and that the discrepancy principle’s worst draw falls below four at every ratio, matching the rightmost minimum.

The misses on the right. Two or three draws at every ratio are missed by the rightmost minimum by two to four times, and the guard was never going to touch them. The proposal to read them through the mean square of the coefficients near the chosen λ\lambda stands, now with a noise level that can safely be compared against: the held-out one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Discrepancy principleDiscretisationGeneralised cross-validationLeast-squaresParameter choiceRegularisationTikhonov regularisation