An estimate that shares the dip's luck
Worth reading first: When the answer is a choice · Counting what cannot be looked at · A parameter that counts steps.
More samples take the floor and leave the dip gave generalised cross-validation more samples than unknowns and watched half of its failures go. The minimum at the floor of the λ scale, a zero divided by a zero on a square system, disappeared at every ratio. The interior dips did not: 20, 13, 8 and 4 draws in 240 at 1.25, 1.5, 2 and 4 samples per unknown had a global minimum to the left of the real one, and at four per unknown one of them still put plain GCV at 877 times the best error available.
The essay traced a dip to one number. Write ρ for the residual per remaining degree of freedom in units of the noise variance,
where δ is the noise per sample and is GCV’s trace term. Admitting one more direction lowers GCV when its noise coefficient’s square exceeds , so a stretch of λ where ρ is low is a stretch where luck can build a dip. That needs δ, which GCV does not have — but “with , the least-squares residual … is an estimate of the noise with degrees of freedom. A rule that estimated δ from it, computed ρ at each candidate minimum and refused a dip when ρ was implausibly low would be GCV told the noise by its own data. Whether it removes the four survivors at four samples per unknown without becoming the discrepancy principle is the measurement.”
The survivor is the draw whose curve falls away four decades to the left of its real minimum, at , and at four samples per unknown it is the only one of 240 that costs more than ten times the oracle. Every curve’s left end sits above its minimum, so nothing about the shape at the floor of the scale gives it away; the dip is an ordinary-looking minimum that happens to be deeper. That is what made a test on the residual attractive: if the dip’s shape is unremarkable, the size of what it leaves unexplained might not be.
It removes none of them, and it does not become the discrepancy principle either. It becomes GCV.
The guard
The estimate is the classical one. The least-squares solution leaves a residual orthogonal to the range of the operator, and if the data are signal in the range plus white noise, that residual is pure noise in dimensions:
It comes from the same singular value decomposition GCV uses, at no cost. With it, can be computed at every λ. An honest residual has ρ near one, scattered by the chi-squared spread of an average over terms divided by another over ; the guard refuses a local minimum of GCV whose is below the lower 1% point of that spread in its normal approximation, , and takes the deepest local minimum it does not refuse. If it refuses all of them, it falls back to the rightmost.
The problems are the previous essay’s: a continuous Gaussian blur sampled at points of a signal represented on nodes, five grids from 24 to 96 unknowns, noise at 1%, 0.1% and 0.01% of the data’s root mean square, sixteen seeded draws of each, 240 draws at each of four ratios of samples to unknowns. Each draw is solved with Tikhonov at 121 values of λ, and every rule’s choice is scored by its error over the best error on the same draw.
None refused
The figure at the top of the page is every interior dip in the 960 draws, 45 of them, placed by ρ at the dip computed two ways: horizontally from the true noise variance, vertically from the estimate.
The proposal’s premise holds on the horizontal axis. With the true noise, ρ at a dip has a median of 0.84, and 18 of the 45 sit between 0.43 and 0.75 — the residual there really is smaller than its degrees of freedom should carry, which is what lets a lucky direction pull GCV down. On the vertical axis the same 45 dips have a median of 0.98, and only five are below 0.9. The estimate has moved every dip to one.
So the guard refuses none: zero of 20 dips at 1.25 samples per unknown, zero of 13 at 1.5, zero of 8 at 2 and zero of 4 at 4, at a threshold of 2.33 standard deviations and at 1.28 and 3.09 as well. It never refuses the rightmost minimum either, and at 2.33 its pick is plain GCV’s on all 960 draws. The 877-times draw is still 877 times.
The left end is the estimate’s own residual
The reason is visible on any one draw. At four samples per unknown, the worst dip has 96 unknowns sampled at 384 points. Walk λ from the right: past the rightmost minimum at the residual is honest, and as λ falls towards the floor of the scale the filter admits every direction and the residual becomes the least-squares residual. That residual is, by definition, the one was computed from:
exactly, whatever the noise did. On this draw at the floor of the scale is 0.98, the dip at reads 0.98, and the plateau between them is flat. The dip lives near the least-squares end of the scale — that is what makes it a dip, a minimum reached by admitting nearly everything — and near that end the estimate is measuring the residual against itself.
With the true noise the same walk reads 0.81 at the dip: low, as the proposal said. On this draw the noise in the 288 directions outside the range came out three per cent small, and every residual near the least-squares end inherits that. The estimate inherits it too. Turn the dial to 1.25 samples per unknown: the worst dip there, 1.77·10⁴ times the oracle, has ρ of 0.62 from the true noise and 0.98 from the estimate, because its 12 degrees of freedom outside the range came out 39 per cent small and the estimate is 0.61 of the truth.
That is the whole mechanism, and the next figure shows it is not one draw’s.
The dips are the draws whose estimate came out low. At 1.25 samples per unknown the median estimate over all 240 draws is 1.02 of the truth and over the 20 dips 0.84; at 1.5, 0.89 on the dips; at 2, 0.93; at 4, 0.97. A dip forms where the residual near the least-squares end is small for its degrees of freedom, and the least-squares residual is the residual near the least-squares end, so a draw that is lucky enough to dip is a draw whose estimate is low by the same luck. Dividing one by the other cancels exactly the thing the guard was meant to see. The estimate is unbiased over all draws and biased on precisely the draws that matter — the same selection, in a different costume, that a spread measured on probes it does not average removed from a trace estimator by keeping the spread’s probes apart from the mean’s.
The figure has a second population, high above the rest: estimates eight to seventeen times the truth. Those are the coarsest grid, 24 unknowns, at the lowest noise level. There the least-squares residual is not noise at all. The data are a continuous blur, and 24 nodes cannot represent it, so the part of the data outside the range is mostly the discretisation’s misfit, which at 0.01% noise is larger than the noise. An estimate made from the residual is an estimate of everything the model does not fit, which is the noise only when the model fits everything else. That error is harmless to the guard — an overestimate makes ρ smaller everywhere, and the guard refuses a minimum on none of those draws — but it is the reason the estimate cannot be read as “the noise” without knowing the model.
Told the true noise, still no separation
The obvious rescue is a better estimate. Give the guard the exact noise variance — the best any estimate could do — and its floor is on the trace term alone. At the 1% threshold it refuses 3 of the 45 dips and 3 real minima. At 10%, 14 dips and 78 real minima. At 0.1%, nothing.
The dips sit near the least-squares end, where is close to : twelve degrees of freedom on the 48-unknown grid at 1.25 samples per unknown, twenty-four on 96 unknowns. The chi-squared spread on twelve degrees of freedom puts one honest residual in ten below 0.53, which is lower than most dips’ ρ. The dips are low, but they are low by less than the noise in ρ itself on the few directions left at that end of the scale. On the 1.25 worst draw the true-noise floor at the dip is 0.11 against a ρ of 0.62. And ρ at the real minimum, where many more directions remain, is not low at all on dip draws — 1.02 median at 1.25, 0.91 at 4 — so it is not a signature either. The dip’s ρ is the signature, and it is a signature whose noise is as large as the signal.
So the guard fails twice over: the estimate cancels the low ρ it was meant to see, and even with the true noise the low ρ is inside the spread of an honest one. The previous essay’s account of a dip is confirmed — a low ρ near the least-squares end — and its proposed detector is refuted by the same account: a quantity estimated on few degrees of freedom is a weak test of anything.
Handed the estimate, the discrepancy principle breaks
The question’s other half was whether the guard would “become the discrepancy principle” — which is told a noise level and uses it for everything. The guard became nothing; its line is plain GCV’s at every ratio, 1.77·10⁴, 458, 60 and 877 times the oracle. But the estimate can be handed to the discrepancy principle directly, and the result is worth knowing because it is the obvious next idea and it is worse.
Told the true noise, the discrepancy principle’s worst draw is within 1.16 of the oracle at every ratio, as it has been on every grid in this series. Told the estimate, its worst draw is 5.7·10⁵ times the oracle at 1.25 samples per unknown, 1.39 at 1.5, 5,527 at 2 and 1.18 at 4. The catastrophes are on the fine grids, 48 unknowns and more, on the draws whose estimate came out 12 to 40 per cent low. Thirty-two coefficients instead of a noise level measured the same cliff on square systems: told too little noise, the discrepancy principle does not degrade, it falls past the best truncation into the directions that amplify noise, and on a fine grid those directions amplify by thousands. An estimate on degrees of freedom is too small by a fifth on a few draws in a hundred, and that is enough.
The rule that wins is the one that reads no noise level at all. The rightmost local minimum’s worst draw is 3.70, 2.98, 2.68 and 2.57 times the oracle at the four ratios, and its median is GCV’s. The minimum on the right proposed it to remove the degenerate minimum of square systems; here it also removes every interior dip, because a dip is by construction to the left of the rightmost minimum, and it needs no estimate to do it.
What the residual cannot be used for
The general lesson is about where an estimate comes from. A noise level read from the least-squares residual is a fine number to report and a poor one to decide with, at least when the decision is about the same end of the scale the residual comes from. It shares the residual’s luck with every candidate λ near that end. It is an estimate of everything outside the model’s range, not only the noise. And on a mildly oversampled system it has few degrees of freedom — six on the 24-unknown grid at 1.25 samples per unknown — so its own spread is as large as the effects it would be used to detect.
The data count their dimensions, not the steps found the discrepancy principle within 16 per cent of the oracle on every grid when told the noise, and GCV better on the median draw with thousand-fold failures. This essay adds the missing corner of that comparison: the discrepancy principle told an estimate of the noise made from the same data inherits GCV’s thousand-fold failures without its median, and the cheapest robust choice from the data alone remains GCV restricted to its rightmost minimum.
One draw in twenty measured GCV’s catastrophes as a tail — four to six draws in a hundred, at a thousand draws a level, while the median stayed among the best — and that is the shape every comparison here has had. The guard does not change GCV’s median, because it changes no pick; the discrepancy principle told the estimate moves its median by three to four per cent and buys a tail of its own. Choosing without knowing found the same split on the first problem the rules were set: GCV on the oracle’s λ in the median, the discrepancy principle six per cent off and steady. A rule’s median is set by how it reads the typical draw, and its tail by what it does on the draw where its one piece of information is wrong. An estimate made from the data is wrong exactly on the draws where the data are unusual, which are the draws the tail is made of.
What these draws do not show
One operator, a Gaussian blur, on one signal; noise that is white; five grids and three levels. An estimate of the noise that does not share the residual’s luck — repeated measurements, a held-out set of samples, or a noise level known from the instrument — is a different object and is not measured here. A blur with a longer tail would leave fewer directions at the least-squares end and move the dips; a coloured noise, which noise that spares the answer and fools the rules found fooling both rules on square systems, would bias the estimate in a direction that depends on its spectrum.
Still open: an estimate from samples held out, and the misses on the right
A held-out estimate. The estimate failed because it was made from the residual the dip lives in. Split the samples instead: fit on three quarters, and estimate the noise from how well that fit predicts the quarter it never saw. That estimate does not share the dip’s luck, because the held-out samples’ noise is independent of the fitted samples’ — the same separation that made Stein’s two-stage rule calibrated. The prediction with a sign is that a guard told a held-out estimate refuses the dips about as often as the guard told the true noise — 3 of 45 at the 1% threshold — and so still fails, because the second failure here, ρ’s own spread on few degrees of freedom, is not about the estimate at all; and that the discrepancy principle told a held-out estimate loses its catastrophes only where the held-out set has more degrees of freedom than .
The misses on the right. Two or three draws at every ratio are missed by the rightmost minimum by two to four times, and nothing here touches them. The proposal to read them through the mean square of the coefficients near the chosen λ stands, now with a warning attached: whatever noise level it is compared against must not be read from the same residual.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A rule that has to be told how good its answer will be — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- The corner reads the norm it is drawn in — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- A corner the penalty can afford — both name discrepancy principle, discretisation, regularisation, tikhonov regularisation
- A parameter chosen on a smaller problem — both name discrepancy principle, generalised cross-validation, influence matrix, parameter choice
- The rule that is wrong in the right direction — both name discrepancy principle, generalised cross-validation, parameter choice, tikhonov regularisation
- Where the grid hands over to λ — both name discretisation, parameter choice, regularisation, tikhonov regularisation
Named objects
A flat tag is an object no other essay names yet.
Discrepancy principleDiscretisationGeneralised cross-validationInfluence matrixLeast-squaresParameter choiceRegularisationTikhonov regularisation