Four folds feed one rule
Worth reading first: When the answer is a choice · Counting what cannot be looked at · A parameter that counts steps.
An estimate that shares nothing with the fit told two parameter-choice rules a noise level estimated from samples held out of the fit. Generalised cross-validation’s guard, which refuses a local minimum whose residual per remaining degree of freedom is implausibly low, refused none of 45 dips; the discrepancy principle, which had missed by 570,000 times when told the least-squares estimate, never missed by more than eighteen. The essay blamed the guard’s failure on the estimate’s size: a quarter of the samples, from 7 to 96 of them, gave its threshold too few degrees of freedom to refuse anything.
The proposal was the textbook remedy. “Averaging the held-out mean square over four folds, each holding out a different quarter, would give the estimate four times the samples — in total — at the price of four fits. The prediction with a sign is that the guard’s floor then rises above the dips’ on about as many dips as the true-noise guard refuses, three of 45 at 1%, which would show the thin estimate rather than the guard’s idea was the failure here; and that the discrepancy principle’s worst draw falls below four at every ratio.”
The second half holds by a wide margin. The first fails, and the way it fails shows that the guard’s idea, not the estimate’s size, was the problem — and that the earlier essay’s title was wrong about the quarter too.
Four folds, the same 960 draws
The problems are the familiar ones: a Gaussian blur of a fixed signal sampled at points and reconstructed on unknowns at 1.25, 1.5, 2 and 4 samples per unknown, five grids from 24 to 96 unknowns, three noise levels from to of the data’s root mean square, sixteen draws each, with the same noise as before. The samples are split into four quarters by their index. Each quarter is held out in turn, the other three are fitted by Tikhonov regularisation on the usual grid of 121 values of , and the fit predicts the held-out quarter. The four held-out sums of squares are added at each , and the estimate is the smallest total per sample,
which is four-fold cross-validation’s error at its minimum. Every sample is now predicted once by a fit that never saw it. Both rules take in place of the earlier estimates: the guard’s floor, the lower point of for an honest residual, carries where it carried the quarter’s , and the discrepancy principle takes the largest whose residual is at most .
The figure at the top of the page is the result in one place: the worst draw of 240 at each ratio for five rules. The discrepancy principle told four folds and told the true noise lie on top of each other at the bottom.
Never low
The four-fold estimate is not more accurate than the quarter’s. Its median over the true noise is 1.41, 1.30, 1.21 and 1.10 at the four ratios, slightly higher than the quarter’s 1.35, 1.26, 1.18 and 1.11 at the first three; both are biased upward, because even the best fit predicts unseen samples imperfectly and the estimate includes that error along with the noise. What four folds change is the bottom of the distribution. The quarter’s estimate falls below the true noise on 131 of the 960 draws, as low as 0.609 at 1.25 samples per unknown. The four-fold estimate falls below it on none. Its lowest draw at each ratio is 1.071, 1.062, 1.051 and 1.022 times the noise.
The discrepancy principle cares about nothing else. Plotted against the estimate it was told, its error is flat across the whole range of overestimates — an estimate 7.6 times the noise, on a small grid at the lowest noise level, still lands within 1.01 of the oracle, because there the residual curve is flat for a long way on the side of large — and rises only on the left. All five of the held-out quarter’s misses by more than twice the oracle were told an estimate below the noise, between 0.61 and 0.90 of it. Four folds never go there, and so the principle’s worst draw is 1.194, 1.193, 1.184 and 1.193 at the four ratios, against 1.163, 1.162, 1.155 and 1.164 told the noise itself. For this rule, an estimate that is reliably a little high is as good as the truth.
The five misses make the point draw by draw. At 1.25 samples per unknown, a 40-unknown grid at the highest noise level was told 0.609 of the noise by its quarter and missed by 3.51; four folds told it 1.369 and it missed by 1.028. At 1.5 per unknown on the 96-unknown grid the quarter said 0.832 and the miss was 17.8; four folds said 1.064 and the miss was 1.122. The remaining three — quarters of 0.896, 0.832 and 0.832, misses of 2.52, 2.21 and 7.50 — became 1.075, 1.089 and 1.057 when told four folds. On every one the four-fold estimate was above the noise and the quarter’s below it, and on every one the difference between the two was the whole of the miss.
That is a sharper version of what an estimate that shares the dip’s luck found from the other side. There the least-squares estimate was low on exactly the draws where the residual was low, and the principle, aiming at a target below the whole residual curve, ran to the bottom of the scale and missed by 570,000 times. An estimate’s accuracy on average was never what decided this rule’s worst case; its lowest value was.
All or nothing
The guard is the opposite case. At the 1% threshold, told four folds, it refuses no dip and one real minimum; told the true noise it refuses three dips and three real minima. At 0.1% neither refuses anything. At 10% the true-noise guard refuses 14 dips and 78 real minima, a trade the earlier essays measured and found poor. The four-fold guard refuses 44 of the 45 dips and 520 real minima — more than half of all 960 draws’ real minima. More samples take the floor and leave the dip counted where the dips come from, 28, 20, 13, 8 and 4 draws in 240 as the samples per unknown rose, and every real minimum refused to catch one of them sends the rule to whichever minimum is left. Its worst draw is over the oracle by 17,659 times at 1.25 samples per unknown, the same dip the unguarded rule falls into. On the dial it moves from refusing nothing to refusing nearly everything with no setting in between that separates the two.
The reason is in the quantity the guard reads. Told the true noise, at the rightmost minima — the real ones — spreads from 0.81 at the tenth percentile to 2.04 at the ninetieth, with a median of 1.007, as an honest residual should. Told four folds, it sits between 0.38 and 0.87 over the same percentiles, with a median of 0.70 and a maximum of 1.011. Part of that is the bias: an estimate 10 to 40 per cent high pushes every down by the square of it. The rest is worse than bias. The logarithms of the two values of are negatively correlated, over 960 draws: where the residual happens to be large, the four-fold estimate is large too, and by more, and the ratio falls. The estimate rises with the residual’s luck and cancels it, so the one thing the guard needs to see — whether this residual is low for its noise — is exactly what it cannot see.
Even told the truth the two populations overlap, which is why the true-noise guard’s trade was poor before any estimate entered: at the 45 dips its has a median of 0.84 and a tenth of them are above 1.05, while a tenth of the real minima are below 0.81.
At the dips the four-fold has a median of 0.47 against the real minima’s 0.70, so the dips are lower on average, but the ranges overlap: a tenth of the dips are above 0.75 and a tenth of the real minima below 0.38. A threshold inside the cluster refuses both, which is the 10% result; a threshold below it refuses neither, which is the 1% result.
The quarter was never independent either
The earlier essay called the held-out quarter “an estimate that shares nothing with the fit”, and its argument was that the held-out samples’ noise never entered the three-quarter fit that predicted them. That is true of the fit that made the estimate. It is not true of the fit the guard judges, which is GCV’s fit to all samples, the held-out quarter included. So the quarter’s samples are in both the estimate and the residual it is compared with.
Measured, the quarter’s at the real minima is anti-correlated with the true one by , as much as the four folds’. What the quarter had that four folds lack is its own scatter: with as few as seven samples behind it, its ranged up to 2.99, and that scatter is what occasionally lifted a real minimum above the floor while a dip stayed below it. Four folds take the scatter away and leave the anti-correlation, and the guard, which was failing for one reason with a quarter, fails more cleanly for the same reason with the whole sample.
So the guard’s premise needs an estimate of the noise that is independent of the residual it divides, and no estimate made from the same samples is. The previous essay’s analogy was the right one, applied one step further: a spread measured on probes it does not average worked because the probes behind the spread were not the probes behind the average. Here every estimate available from the data is made from the samples behind the residual, which is a rule that reads only its own probes again with samples in place of probes. An estimate from outside them — replicate readings of the instrument, a calibration run, a specification — is the only kind the guard can use, the same conclusion a unit is a statement about the noise reached for total least squares, where two replicate readings a column were enough to recover weights the data alone could not supply; and with it the guard is the true-noise guard, whose trade the earlier essays already measured and found poor.
What the bias costs, and where
The four-fold estimate’s upward bias is not one number. It is largest where the noise is smallest and the samples per unknown fewest: at a noise level of and 1.25 samples per unknown the median estimate is 1.60 times the noise, at and four per unknown 1.085. That is what a cross-validation error should do. At low noise the error of predicting unseen samples is dominated by the fit’s own error rather than the noise, and with few samples per unknown each fit, trained on three quarters, has fewer samples to fit with. The discrepancy principle shrugs off all of it, because every one of these estimates is high, and an overestimate on these problems costs at most 19 per cent.
Why the two rules part company
The two rules ask the noise estimate different questions. The discrepancy principle asks where the residual curve crosses a level, and on these problems the curve is steep on the left and flat on the right, so a level too high costs almost nothing and a level too low costs everything. What it needs is a level that is never low, and four folds supply one. Choosing without knowing found the principle catastrophic when told a noise level ten times too small; it never asked about one ten times too large, and here one 7.6 times too large cost one per cent.
The guard asks whether this residual is unusually small for this noise, which is a question about the residual’s luck, and it needs a noise level that does not share that luck. No estimate from the samples qualifies, and four folds, by using every sample, make the sharing complete. The rightmost minimum, which the minimum on the right proposed as a plainer rule, needs no noise level at all and stays within 2.57 to 3.70 of the oracle at every ratio; told four folds, the discrepancy principle beats it on the worst draw at every ratio. It does not beat it on the typical draw. The rightmost minimum’s median is 1.004 or 1.005 at every ratio, the four-fold discrepancy principle’s 1.051 to 1.078 and the true-noise principle’s 1.028 to 1.034: aiming at a level a little high costs a few per cent on every draw, and the rightmost minimum, when it is right, is right to the grid. The choice between them is between a few per cent everywhere and a factor of three somewhere.
What four folds cost
The price of the estimate is four more factorisations. Each fold fits three quarters of the samples, so it needs a singular value decomposition of a matrix with rows and columns, and the four together cost about three times the one decomposition of the full matrix that GCV and the discrepancy principle already share; on these problems, with at most 384, that is a fraction of a second. After the decompositions every value of on the grid costs a few vector operations per fold. So the discrepancy principle told four folds is a complete parameter choice for about four times the price of GCV, which needs no noise level, and on these 960 draws it has a better worst case than GCV’s rightmost minimum by a factor of two to three and a worse typical case by about five per cent.
What 960 draws do not show
One family of problems: a Gaussian blur of one smooth signal, Tikhonov regularisation, white Gaussian noise. On a problem whose residual curve is steep on the right as well, an overestimate would cost the discrepancy principle more than it does here, and the four-fold estimate’s bias, up to 60 per cent, would start to matter. Folds by index, which on these uniform grids makes each fold a regular quarter of the domain; random folds would give each fit a different coverage and might change the bias. Four folds and no other number: more folds would reduce the bias, since each fit would train on more samples, and the correlation with the residual would stay. And the oracle is the best on the grid, so a ratio of 1.19 is to the best this grid of 121 values offers.
Still open: more folds, and noise from outside the sample
More folds. With folds each fit trains on of the samples, so the bias should shrink toward the noise as grows while the estimate stays above it. The prediction with a sign is that leave-one-out, the limit, gives an estimate whose median over the true noise is under 1.1 at every ratio and noise level, still never below the noise on any of the 960 draws, and that the discrepancy principle told it keeps its worst draw under 1.2.
Noise from outside the sample. The guard needs a noise level independent of the residual. The prediction is that a guard told an estimate from four replicate readings of each sample — independent of the fitted data by construction — refuses more dips than real minima at the 1% threshold, the first configuration in these essays to do so, and that the discrepancy principle told the same estimate does no better than told four folds, because four folds were already never low.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A rule that has to be told how good its answer will be — both name discrepancy principle, generalised cross-validation, noise level, parameter choice, regularisation, tikhonov regularisation
- Noise that spares the answer and fools the rules — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- The corner reads the norm it is drawn in — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- The data count their dimensions, not the step's — both name discrepancy principle, generalised cross-validation, parameter choice, regularisation, tikhonov regularisation
- A corner chosen without the iteration — both name discrepancy principle, generalised cross-validation, parameter choice, tikhonov regularisation
- One draw in twenty — both name discrepancy principle, generalised cross-validation, parameter choice, tikhonov regularisation
Named objects
A flat tag is an object no other essay names yet.
Cross-validationDiscrepancy principleGeneralised cross-validationNoise levelParameter choiceRegularisationTikhonov regularisation