An estimate that shares nothing with the fit
Worth reading first: When the answer is a choice · Counting what cannot be looked at · A parameter that counts steps.
An estimate that shares the dip’s luck gave generalised cross-validation a guard. A dip in GCV forms where the residual per remaining degree of freedom, , is low, so the guard estimated the noise level and refused a local minimum whose was implausibly small. With more samples than unknowns the noise has a free estimate, the least-squares residual over , and the guard read it. It refused none of 45 dips in 960 draws. At the dip sat near one because the estimate was made from the residual at the dip’s own end of the scale, and inherited its luck; told the same estimate, the discrepancy principle missed by 570,000 times.
The essay’s proposal was the separation that had worked once already in this subject. A spread measured on probes it does not average removed a trace estimator’s bias by estimating its spread from probes it did not then average; here, fit on three quarters of the samples and estimate the noise from how well that fit predicts the quarter it never saw. The trace estimator had needed it for the same reason: a rule that reads only its own probes stopped on a spread computed from the very samples it was averaging, and its unlucky runs were the ones that stopped early. An estimate and the thing it is used to judge must not share their noise. The prediction had two parts with a sign on each. The guard would refuse the dips about as often as the guard told the true noise — 3 of 45 — and so still fail, for a reason that is not about the estimate. And the discrepancy principle would lose its catastrophes only where the held-out set is larger than the least-squares residual’s degrees of freedom, which happens here only at 1.25 samples per unknown.
Both parts are wrong, in opposite directions. The figure at the top is the second: the discrepancy principle told the held-out estimate stays within eighteen times the oracle’s error at every ratio. The guard, told the same estimate, does nothing at all.
Holding out a quarter
The problems are the previous essays’: a Gaussian blur of a fixed signal sampled at points, reconstructed on unknowns at 1.25, 1.5, 2 and 4 samples per unknown, five grids from 24 to 96 unknowns, three noise levels from to of the data’s root mean square, sixteen draws each — the same 960 draws, with the same noise. Every fourth sample is held out. The other three quarters are fitted by Tikhonov regularisation on the usual grid of 121 values of , each fit predicts the held-out samples, and the noise estimate is the smallest held-out mean square over the grid:
The held-out samples’ noise never entered the fit, so the estimate shares nothing with the residual GCV and the guard read. It pays two prices for that. It has only samples behind it — from 7 at the smallest problem to 96 at the largest. And it is biased upward, since even the best fit predicts imperfectly, so the estimate includes the fit’s own error on the unseen samples as well as their noise.
Both reductions in this essay then take in place of the least-squares . The guard’s floor — the lower 1% point of for an honest residual, — now carries where it carried . And the discrepancy principle takes the largest whose residual is at most .
The guard refuses nothing
At the 1% threshold the held-out guard refuses none of the 45 dips and none of the real minima, at any ratio, and its worst draw at each ratio is exactly the least-squares guard’s: 18,000 times the oracle at 1.25 samples per unknown. The guard told the true noise refuses three dips at that threshold. The prediction expected the held-out guard to match it and it does worse.
The figure says why, and it is not the estimate’s bias. at the dips now spreads from 0.27 to 1.22 — the held-out estimate does not pin them at one as the least-squares estimate did — so the dips are no longer disguised. But the floor has a term under its square root, and with twelve to forty samples held out on most dips that term is large: the floor runs from −0.34 to 0.62 across the dips and sits under every dip’s , by at least 0.14. A floor below zero refuses nothing, whatever the estimate says. The held-out estimate is independent of the residual and too thin to judge it.
That is the third failure this guard has met, and the other two did not go away. The least-squares estimate shared the dip’s luck. The true noise, which has neither problem, could not separate dips from real minima because itself spreads widely on few degrees of freedom. And an independent estimate built from a quarter of the samples has too few of its own.
Loosening it buys the wrong trade
Raise the threshold to 10% and the floor rises far enough to bite. The held-out guard refuses sixteen of the 45 dips, two more than the true-noise guard’s fourteen, and pays with 170 refused real minima — draws where GCV’s rightmost local minimum, the one that was right, is thrown out — against the true noise’s 78. Turn the dial to 0.1% and neither guard refuses anything. At no threshold does the held-out guard refuse more than a third of the dips without refusing several times as many minima that should have stood.
The trade is worse than the true noise’s for the same reason the 1% threshold was empty. A floor that has to be raised far enough to bite on the dips, with a term that large under its square root, rises over many honest minima too, and the held-out estimate’s upward bias lowers every and pushes more of them under it. Both effects point the same way.
So the guard is finished as a remedy for these dips, by three independent routes. The minimum on the right had already found the remedy that works: ignore the deepest minimum and take the rightmost, which never misses by more than four times the oracle here — 3.7, 3.0, 2.7 and 2.6 at the four ratios — and needs no noise level at all.
The discrepancy principle comes back
The discrepancy principle had the opposite problem. It does need the noise level, and handed the least-squares estimate it missed by more than a thousand times at 1.25 and 2 samples per unknown. Handed the held-out estimate, its worst draws are 3.5, 18, 7.5 and 1.5 times the oracle at the four ratios. Across all 960 draws one misses by more than ten, against sixteen with the least-squares estimate.
The prediction placed the rescue where the held-out set is larger than . At 1.25 samples per unknown it is: seven against six samples at the smallest grid, thirty against twenty-four at the largest. At 1.5, 2 and 4 samples per unknown it is the smaller by a factor of 1.3 to 3, and the rescue happens anyway — 5,500 down to 7.5 at two per unknown. The degrees of freedom were not what made the least-squares estimate dangerous.
What made it dangerous is visible here. A noise estimate that is too low is the discrepancy principle’s only route to catastrophe, because it asks for a residual smaller than the noise and gets one only by under-regularising. The least-squares estimate is under 0.95 of the noise on 177 draws; the held-out estimate on 76. That is a difference, but not the difference: the least-squares estimate’s sixteen catastrophes all come from estimates between 0.60 and 0.88, and the held-out estimate reaches 0.61 too, with one miss above ten among its 76 low draws. Low is dangerous when it is low because the residual is: the least-squares estimate is computed from the least-squares residual, the floor the residual curve descends towards, so when that floor is low by luck the target falls with it, and the two move together. A held-out estimate that is low is low for an unrelated reason, and the residual curve it is compared against does not follow it down.
A target under the whole curve
The worst draw shows the mechanism at full size. At 1.25 samples per unknown, 96 unknowns and noise , the true noise norm is 0.072 and the least-squares estimate puts it at 0.051 — seven tenths. The residual’s lowest value anywhere on the grid is 0.054. The target is under the whole curve: no satisfies it, and the rule falls back to the smallest on the grid, where the error is 570,000 times the oracle’s. The held-out target, 0.084, crosses the curve where it rises steeply out of its floor; the discrepancy principle lands at against the oracle’s , and its error is 1.05 times the oracle’s. Told the true noise it lands at and 1.03.
The least-squares residual and the estimate are the same number up to a factor: the estimate is that residual divided by , and the target multiplies it by , so the target is always exactly times the least-squares residual — 2.26 at 1.25 samples per unknown — however large the noise is. On this draw the least-squares residual is 0.023, so the target is 0.051. But the residual curve does not reach the least-squares residual on the grid: even at it still filters the components with the very smallest singular values, and their share of the data keeps it at 0.054. The target, tied to a floor the curve never reaches, lands under the curve. Nothing ties the held-out target to that floor.
Where the estimate is made
The held-out estimate is a minimum over , and the at which it is reached is itself a choice made from data the fit did not see. Measured against the oracle’s for the full problem, it sits at a median of 0.07 decades above it, with half of the 960 draws within half a decade either side. The estimate is made, in other words, where the fit on three quarters is about as good as a fit can be, and its upward bias is the error that remains at the best — the part of the unseen samples no smooth reconstruction predicts — rather than a symptom of fitting at the wrong one.
That says two things. The bias cannot be removed by searching harder; it is the floor of prediction error on samples the fit never saw, and it grows exactly where the figure above shows it growing, at low noise, where the reconstruction’s own error is no longer small beside the noise. And held-out validation is, here, a parameter rule in its own right, the one when the answer is a choice said something outside the fitted data would have to supply. It has not been scored as one in this essay, because it chooses for a fit on three quarters of the data and the question was the noise level. It is the obvious next measurement.
The worse estimate is the better input
Measured as an estimate, the held-out one loses. Its median is 1.11 to 1.35 times the noise across the ratios, where the least-squares one is 1.01 to 1.08, and it is worst at the lowest noise — up to 1.5 times at and 1.25 samples per unknown, and 63 times on its worst draw — because there the best fit’s own error on unseen samples is no longer small beside the noise. A statistician choosing between them by bias and spread would take the least-squares estimate every time.
The discrepancy principle should not, and the reason is the subject’s recurring one. An input to a rule is judged by what the rule does with it, and the discrepancy principle is catastrophic for low estimates that track the residual and merely conservative for high ones: an estimate too large by a third regularises too much and loses a little, which is why the held-out rule’s median error is 1.04 to 1.06 times the oracle against the true noise’s 1.03. That asymmetry rewards an estimate whose errors are upward and independent, and the held-out estimate is exactly that. More samples take the floor and leave the dip found the same structure in GCV — the failure that mattered was a minimum the method could fall into, not its typical accuracy — and noise that spares the answer and fools the rules found that what breaks the parameter rules is what they assume about the noise rather than the noise’s effect on the answer, which is the same lesson from the other side. The data count their dimensions, not the step’s had the discrepancy principle as the robust rule on square grids, told the noise; the whole of its trouble since has been the noise level it is told.
What a solver would do
With more samples than unknowns, the measurements give an order. To choose without a noise level, take GCV’s rightmost local minimum: within four times the oracle on every draw. To use the discrepancy principle, never feed it the least-squares residual’s estimate; hold out a quarter of the samples, fit the rest, and use the best held-out mean square, which costs one more SVD of three quarters of the matrix and keeps the rule within eighteen times the oracle. And do not guard GCV with any noise estimate, held out or not: the guard has now failed with an estimate that shares the residual’s luck, with one too thin to judge it, and with the truth.
The held-out estimate gives up a quarter of the data to the fit, and on problems this small that costs something the measurements do not separate: the fit on three quarters is a slightly different problem from the fit on all of them, and the discrepancy principle then applies the held-out noise level to the full fit. That was deliberate — the estimate’s job is to say how large the noise is, which does not depend on which samples are fitted — and the median errors say it did no harm.
What these draws do not show
One blur, one signal, white noise of known distribution, a single held-out quarter chosen as every fourth sample. Several folds averaged would give the estimate more degrees of freedom and might rescue the guard’s floor; the held-out set chosen at random rather than interleaved might change the fit’s error on it; and a signal whose smoothness differs from the blur’s would change the upward bias. None is measured.
Still open: folds, and the misses on the right
Four folds instead of one. Averaging the held-out mean square over four folds, each holding out a different quarter, would give the estimate four times the samples — in total — at the price of four fits. The prediction with a sign is that the guard’s floor then rises above the dips’ on about as many dips as the true-noise guard refuses, three of 45 at 1%, which would show the thin estimate rather than the guard’s idea was the failure here; and that the discrepancy principle’s worst draw falls below four at every ratio, matching the rightmost minimum.
The misses on the right. Two or three draws at every ratio are missed by the rightmost minimum by two to four times, and the guard was never going to touch them. The proposal to read them through the mean square of the coefficients near the chosen stands, now with a noise level that can safely be compared against: the held-out one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A rule that has to be told how good its answer will be — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- The corner reads the norm it is drawn in — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- A corner chosen without the iteration — both name discrepancy principle, generalised cross-validation, parameter choice, tikhonov regularisation
- A corner the penalty can afford — both name discrepancy principle, discretisation, regularisation, tikhonov regularisation
- Choosing without knowing — both name discrepancy principle, parameter choice, regularisation, tikhonov regularisation
- One draw in twenty — both name discrepancy principle, generalised cross-validation, parameter choice, tikhonov regularisation
Named objects
A flat tag is an object no other essay names yet.
Discrepancy principleDiscretisationGeneralised cross-validationLeast-squaresParameter choiceRegularisationTikhonov regularisation