Noise that spares the answer and fools the rules
Worth reading first: When the answer is a choice · Choosing without knowing.
Where the answer stops being in the data ends with a warning it does not test. Everything in it assumes the noise is spread equally over the singular directions, which is what makes the Picard floor flat; noise with structure, it says, “concentrates in particular directions and does not produce a flat floor at all, so the crossing is not a crossing and the whole detector has nothing to detect. That is the realistic case and it is not measured here.”
This essay measures it. The warning turns out to be right about the floor and wrong about the detector, and the damage it predicted lands somewhere else entirely — not on the Picard picture and not on the answer, but on the rules for choosing λ that choosing without knowing scored against an oracle on white noise.
A floor with a slope
The noise model is the simplest one that is not white: each sample is ρ times the previous sample plus a fresh independent part, scaled so the variance is the same at every point. Its covariance is ρ raised to the distance between two samples, so ρ = 0 is white noise and ρ near one is a slow wander — an instrument whose error at one reading remembers the error at the last, which is what drift, a temperature-sensitive baseline and a detector with a time constant all look like.
Every draw is scaled to the same relative size, 0.1% of the exact data, so the comparisons below are at equal ‖e‖ and only the noise’s spectrum moves. At ρ = 0 the draw is identical to the white draw the earlier essays used, number for number, which is checked rather than assumed; so anything that changes as ρ rises is a statement about correlation and not about a changed setup.
What the figure shows first is the floor. For white noise the coefficients uₖᵀe of the noise alone are flat across the spectrum, at about ‖e‖/√n each, and the tilt between the first eight and the last sixteen is 0.07 decades — noise in the measurement of noise. At ρ = 0.9 the tilt is 1.14 decades: the noise is more than ten times larger in the leading directions than in the trailing ones. The Picard essay was right that the floor stops being flat.
Where correlated noise goes, and why it is cheap there
The direction of the tilt is the whole story, and it is not an accident of this operator.
Correlated noise is slow. Neighbouring samples move together, so its energy sits at low frequencies. The blur’s leading singular vectors are its smooth ones — the blur passes slow variation and destroys fast variation, which is what a blur is — so the leading directions and the slow directions are the same directions. At ρ = 0 the first 28 singular directions carry 43% of the noise’s energy, roughly their share of 64. At ρ = 0.9 they carry 98%.
Those are the directions every sensible filter keeps whole, and they are also the directions where σₖ is close to one. A term of the regularised sum contributes fₖ(uₖᵀe)/σₖ of noise to the answer, and dividing by a σ near one does not amplify anything. So the noise has moved, at fixed size, from the places where it was expensive to the places where it is nearly free.
The measured consequence is that correlated noise makes the best available answer better, not worse. The median over 48 draws of the least error any λ achieves is 0.1056 for white noise, 0.1049 at ρ = 0.5, 0.1010 at ρ = 0.9 and 0.0984 at ρ = 0.99. The improvement is small, 7% at the far end, and it is in the opposite direction from the intuition that structured noise is harder noise. It is harder to detect. It is not harder to survive, on an operator that smooths, because the operator’s inverse amplifies exactly the fast variation that correlated noise lacks.
The resolution side of the answer does not move at all. A second blur, narrower than the first splits the error into what the averaging kernel does to the truth and what the filter lets through of the noise, and the first half contains no e. At any given λ the kernel is the same matrix whatever the noise’s spectrum is. Everything correlation does, it does to the second half and to the choice of λ.
The crossing survives
The Picard essay’s second prediction was that the detector would have nothing to detect. On this problem it does not come close to that.
The detector finds the crossing as the minimum of the smoothed coefficient curve, on the reasoning that below the noise floor the curve cannot fall further. A tilted floor still has a minimum. It sits further right than a flat floor’s would, because the floor itself is falling there — on the draw drawn above, the crossing moves from 32 at ρ = 0 to 43 at ρ = 0.9 while the best truncation moves from 28 to 31. Across 48 draws the median distance from the crossing to the best truncation is 22 for white noise and 24.5 at ρ = 0.9, and the crossing is later than the best truncation on all 48 of the correlated draws.
The reason the tilt does so little damage is scale. A tilt of 1.1 decades across 64 indices is a gentle slope next to a spectrum that falls thirteen decades over the same range. The signal’s coefficients drop through the tilted floor at nearly the same place they dropped through the flat one, because the thing falling fast is the signal and not the floor. A noise model steep enough to hide the crossing would need a tilt comparable to the decay of the singular values themselves — noise whose spectrum is shaped like the blur’s — and at that point it is no longer usefully called noise, since no measurement could separate it from a smoother signal.
So the honest correction to the Picard essay is narrower than its warning. The floor tilts; the crossing is still a crossing; and the ρ = 0.99 draw, which is nearly a random walk and tilts by 1.57 decades, puts the crossing at 43 and the best truncation at 31, which is the same picture again.
The rule that is told nothing is told the wrong thing
Generalised cross-validation chooses the λ that minimises ‖Ax_λ − b‖² divided by the square of the number of components the filter discards. It needs no noise level and no knowledge of the answer, which is why it did best of the three rules on white noise, landing on the oracle’s λ on most draws.
The quantity it reads is the residual, and the residual is made of the noise in the directions the filter throws away. For white noise that part is a fair sample of the whole. For slow noise it is not: almost none of the noise is in the discarded directions, so the residual is small, and GCV reads a small residual as evidence of a small noise level. Evidence of a small noise level is a reason to regularise lightly. So GCV under-smooths correlated noise, systematically, and the size of the miss is measurable: the median ratio of its λ to the best λ is 1.00 for white noise, 0.74 at ρ = 0.3 and 0.54 at ρ = 0.9.
An under-smoothing miss by a factor of two in λ would cost little on its own — the error curve is shallow near its minimum. What it costs in practice is the tail. GCV’s objective on this problem can have a second valley at very small λ, where the solution is fitting noise and the residual is tiny: on the worst white-noise draw it has a local minimum at a sensible 8.9·10⁻⁴ and its global minimum at 3.2·10⁻⁸, lower by a third. A rule already biased towards small λ falls into that valley more often. The count of draws on which GCV more than doubles the best error rises from 3 of 48 for white noise to 6 at ρ = 0.5, 10 at ρ = 0.7 and 14 at ρ = 0.9. The worst of them are not doublings: GCV’s λ lands at the bottom of the grid, 2.5·10⁻⁸, and the error is thousands of times the best.
The median hides this completely. GCV’s median cost goes from 1.005 to 1.027 — two per cent, which a comparison on one draw or on a median would report as correlation doing no harm. A reader who took what a single draw cannot report seriously would count the doubled draws instead, and the count goes up by a factor of nearly five.
The white-noise tail is not correlation’s doing, and it deserves its own sentence. Three draws in 48 already defeat GCV on white noise, the worst of them at 25,000 times the best error, and the sixteen draws choosing without knowing ran across did not happen to include one. That tail is a property of the rule on this problem; what correlation does is multiply it.
The rule that is told the size misses the other way
The discrepancy principle is told ‖e‖ and picks the largest λ whose residual is no bigger than it. It does not double the error on a single draw at any correlation measured, from white noise to ρ = 0.99. Its typical cost does rise: the median multiple of the best error goes from 1.063 for white noise to 1.122 at ρ = 0.7, 1.187 at ρ = 0.9 and 1.220 at ρ = 0.99 — the six-per-cent cost roughly tripled.
The mechanism is GCV’s in reverse. The rule is told the true size of the noise and looks for a residual of that size. But with slow noise most of e is absorbed into the solution — the filter keeps the directions it lives in — so the residual at any sensible λ is much smaller than ‖e‖. The only way to make the residual as large as ‖e‖ is to discard directions that contain signal, which is to say to over-smooth. The discrepancy principle’s λ is a median 2.5 times the best for white noise and 8.6 times the best at ρ = 0.9. It misses high by a larger factor than GCV misses low, and costs less for it, because the error curve rises gently on the over-smoothing side and steeply on the other.
So the two rules fail in opposite directions for one reason. Both read the residual as a sample of the noise, one to estimate its size and one to match a known size, and correlated noise makes the residual an unrepresentative sample in a known direction: too small. A rule that under-reads the noise regularises too lightly; a rule that is told the right noise and cannot find it in the residual regularises too heavily. The L-curve corner, which reads the residual against the solution norm, gets worse with correlation as well — doubling the error on 8 draws of 48 for white noise and on 22 at ρ = 0.9.
Counting draws rather than reading one
The figure below is the measurement the essay rests on, and it is a count.
Two features of the counts are worth reading rather than smoothing over. GCV’s count is not monotone: it peaks at ρ = 0.9 and falls to 11 at ρ = 0.99. Forty-eight draws is enough to see a trend of a factor of four and not enough to say whether a difference of three is real, and the counts are reported as counts for that reason. And the L-curve’s count keeps rising to the end, reaching more than half the draws, which makes it the rule most damaged by correlation in absolute terms — though it was already the worst of the three on white noise, as choosing without knowing found.
The discrepancy principle’s zero is the most robust number on the page, and the reason is the one that made it the weakest rule in the earlier measurement: it is the only rule that is told something from outside the data, and being told the noise’s size protects it from the catastrophe while its reading of the residual still biases it. The two numbers a caller has makes the general version of this point for a different decision — that the quantities computable from a solve are blind to where the noise is — and here is the same blindness moving a parameter rather than choosing a method.
Whitening, and what it has to know
The standard repair is to change the problem until the noise is white. If the noise covariance C is known, a matrix W for which W C Wᵀ is the identity turns b = Ax + e into Wb = (WA)x + W·e, whose noise term is white, and every rule can be run on the whitened problem as though nothing had happened. For this noise model W is bidiagonal: the first row passes the first sample, and every later row takes the difference between a sample and ρ times its predecessor, rescaled.
The Picard picture comes back flat, and the rules come back with it. After whitening, GCV doubles the error on 3 draws of 48 at ρ = 0.9 and on 2 or 3 at every other correlation, which is the white-noise count; its λ ratio recovers to 0.74. The discrepancy principle’s median cost falls from 1.187 to 1.013. The best available error does not move — 0.1010 before whitening and 0.1014 after — which confirms that whitening repairs the choice rather than the problem.
The repair is not the same size for every rule, and the strip of draws shows where it is partial.
Read row by row, the whitened strip says three different things. GCV’s tail has collapsed from 14 draws past the doubling line to 3 — the white-noise count — and its median cost is 1.010, so the rule is back to being the one that usually lands on the oracle and occasionally falls into the valley at small λ. The discrepancy principle’s cluster has moved left as a block, to a median of 1.013. And the L-curve corner has improved only part of the way: 22 doubled draws become 14, against 8 for white noise, and its median cost is 1.771 against 1.506.
That last row is the useful caution. Whitening makes the noise white; it does not make the whitened problem the white-noise problem. The operator is now the blur followed by a differencing, its singular values fall from 0.55 rather than from 1 and are 2.3 times flatter over the first sixteen, and a rule that reads the shape of a curve — the L-curve’s corner is a feature of how the solution norm trades against the residual norm — reads a different shape on a different operator. GCV and the discrepancy principle read the residual’s size against the noise’s, and for them white noise is the whole of what was missing. The L-curve reads geometry, and part of what moved its corner was the geometry.
It is not free, and the cost is information. The discrepancy principle had to be told one number, ‖e‖. Whitening has to be told the whole covariance — here a single ρ, because the model is known to have one parameter, but in general a matrix. That is strictly more than any of the three rules was given, and a wrong ρ produces a W that leaves the noise coloured in some other way. When the matrix is wrong too makes the parallel point for total least squares: choosing a method is choosing a noise model, and here choosing to whiten is choosing a covariance.
What whitening did that was not asked of it
One number from the whitened runs is better than it has any right to be, and it is reported rather than explained.
The discrepancy principle’s median cost on the whitened problem is 1.013 at ρ = 0.9 and 1.002 at ρ = 0.99 — lower than its 1.063 on white noise, where whitening does nothing. So whitening has not merely restored the rule to its white-noise behaviour; at strong correlation it has made it better than it ever was. The whitened problem has white noise, but it also has a different operator, a differencing applied after the blur, and that operator’s spectrum is flatter at the top than the blur’s. Where the discrepancy principle lands depends on the operator as well as the noise, and on this operator it lands closer to the oracle.
That is a measurement on one blur and one signal, and it would be a mistake to read it as a recommendation to whiten noise that is already white. What it does show is that a rule’s cost is a joint property of the rule, the noise and the operator — which is the reason a bound that holds with probability insists on stating what a number was measured over before it is quoted.
Where this goes from here
Three continuations follow from this one, and each of them is a measurement this essay needed and did not make.
Estimating the covariance from the data. Whitening here was handed ρ. A real instrument’s correlation has to be estimated, and the natural estimate — from the residual of a fit — is biased by exactly the absorption this essay measured, since slow noise is the part a fit swallows. Whether an estimated ρ whitens well enough to recover GCV’s count, and how wrong it can be before the repair stops working, is open.
Noise that is not stationary. The model here has the same variance and the same correlation at every sample. Noise whose size grows along the signal, or a drift that switches on part-way, is not whitened by any bidiagonal W, and the tilt of the floor would then vary along the spectrum rather than being one slope.
The grid’s own part in this. Correlation between neighbouring samples is correlation per grid step, so the same physical instrument has a different ρ on a finer grid. The grid was the first filter measures what the grid itself does to an unregularised solve with white noise per sample; the correlated version, where refining the grid also makes the noise slower, is not drawn. And the averaging kernel of a second blur, narrower than the first would let a rule choose λ for resolution, which reads no residual and would be immune to everything above — at the price of needing a bound on the noise that correlation, by construction, makes harder to state.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The corner reads the norm it is drawn in — both name discrepancy principle, generalised cross-validation, l-curve, noise floor, parameter choice, regularisation, tikhonov regularisation
- One draw in twenty — both name discrepancy principle, generalised cross-validation, l-curve, noise floor, parameter choice, tikhonov regularisation
- Thirty-two coefficients instead of a noise level — both name discrepancy principle, noise floor, parameter choice, picard condition, regularisation, tikhonov regularisation
- A parameter chosen on a smaller problem — both name discrepancy principle, generalised cross-validation, parameter choice
- A parameter that counts steps — both name discrepancy principle, tikhonov regularisation
- A ranking that is an eigenvector — both name parameter choice, regularisation
Named objects
A flat tag is an object no other essay names yet.
Correlated noiseDiscrepancy principleGeneralised cross-validationL-curveNoise floorParameter choicePicard conditionPrewhiteningRegularisationTikhonov regularisation