Where the grid hands over to λ
Worth reading first: When the answer is a choice · The stencil that is not symmetric.
The grid was the first filter took a continuous deconvolution, solved it with no regularisation on every grid from 12 to 48 points, and found an optimum: at 0.1% noise per sample the error is least at 26 points, beside the best truncation of a 64-point operator, and past 34 points it is already twice as bad. Choosing a grid, it concluded, is choosing a truncation. It ended with the question its setup kept out: every point on its curve was either a grid with no λ or a fine grid with λ, and nobody had put the two together.
The question has an obvious candidate answer, and it is a trade. A grid and a λ both decide how much of the answer’s detail survives, so a moderately coarse grid with a little regularisation ought to be interchangeable with a fine grid and a lot, and the joint choice ought to be a ridge on a surface: many pairs, all about as good. The aliasing the grid essay found — components past a grid’s size folded back into its answer, where no λ on that grid can reach them — pointed the other way, but only weakly.
The measurement does not find a trade. It finds a handover. Below a certain grid size the grid does all the filtering and no λ improves it by as much as 2 per cent. Above 40 points λ does all of it, and the λ that does it is the same number on every grid to within a quarter of a decade, so the grid stops mattering. In between, each grid with its own best λ reaches something between the two, and none reaches the fine grids.
The sweep, and what is held fixed
Everything is the grid essay’s setup. A Gaussian blur of fixed width in the continuous variable acts on a signal of two bumps and a step on [0, 1]. On a grid of n points the operator is the sampled kernel with each row normalised to sum to one. The data at each grid point are the continuous blur of the continuous signal, computed on two thousand quadrature points, so a coarse grid is handed measurements it cannot quite explain. The noise per sample is a fixed fraction of the data’s root mean square, and the same sixteen seeds are drawn on every grid. Every answer is joined into a piecewise-linear function and scored against the continuous signal.
To that, one parameter is added. On each grid the Tikhonov solution is computed at 31 values of λ from 10⁻⁸ to 10⁻⁰·⁵, four to a decade, and the λ reported as best for that grid is the one whose median error over the sixteen draws is least — a single λ for the grid, as an instrument would have, rather than the best λ for each draw separately. The grids run from 16 to 96 points, which takes the matrix’s condition number from 2.8 to 9·10¹⁷ and the unregularised error from 0.18 to something astronomically large.
Below the best unregularised grid, λ has nothing to do
The left-hand side of the hero figure is the first finding and the least surprising one. On the grids of 16, 20, 24 and 26 points the best λ is the smallest one tried, and the error with it is the error without it to four figures. The same holds on 28 points. At those sizes the matrices’ condition numbers are 2.8, 8.0, 29, 61 and 134, and the noise at 0.1% per sample is amplified to nothing a filter could recover: the grid essay’s noise-free and noisy curves coincide there, because the error is resolution rather than noise.
A λ on such a grid can only smooth an answer that is already too smooth. The curve of error against λ for 24 points shows it directly.
On 24 points the error is 0.1370 from λ = 10⁻⁸ up to about 10⁻³ and rises from there. On 34, 48 and 96 points the curve has the familiar shape of every regularisation curve on this problem — enormous at small λ, a minimum, a slow rise — and the three minima sit at 5.6·10⁻³, 5.6·10⁻³ and 3.2·10⁻³. The 34-point grid’s minimum is the highest of the three, at 0.1297; the 48-point grid’s is 0.1206 and the 96-point grid’s 0.1178.
Up to its best unregularised size, then, a grid is a filter that leaves λ nothing to filter. That was implicit in the grid essay’s picture of a grid as a truncation: a truncation below the noise’s own cut has discarded nothing the noise could have spoiled, and a second filter on top of it can only discard signal.
Above 40 points the grid stops mattering
The right-hand side is the finding that decides the question. On grids of 40, 48, 64 and 96 points the best λ at 0.1% noise is 3.2·10⁻³, 5.6·10⁻³, 5.6·10⁻³ and 3.2·10⁻³, and the errors it reaches are 0.1163, 0.1206, 0.1200 and 0.1178 — all within 4 per cent of each other. The matrices over that range differ in condition number by thirteen orders of magnitude, and the unregularised solves from 5 to 4·10¹⁴ in error. With λ, none of that reaches the answer.
Two single draws make the contrast concrete.
On 34 points with its best λ the answer’s error is 0.129, and with noise-free data on the same grid and the same λ it is also 0.129. The regularisation has done its job completely — nothing of the noise remains — and what is left is a 34-point piecewise-linear function’s inability to put the step’s corners where they are and a λ’s rounding of the bumps. The grid essay’s unregularised 34-point solve at this noise level was 0.267.
On 96 points the same comparison reads 0.114 and 0.113. The noise is gone here too, and the resolution left is better: the step’s sides are samples a ninety-fifth of the interval apart rather than a thirty-third, and λ, not the grid, sets how sharp they may be. The matrix has a condition number of 9.3·10¹⁷, which is to say it is singular to working precision, and the regularised answer on it is the best on the page. A fine grid’s ill-conditioning is not a problem once λ is present, because the components that make it ill-conditioned are exactly the ones λ removes — the grid essay’s picture of the unregularised 36-point solve oscillating between −1.1 and 2.1 was a picture of those components arriving with nothing to stop them.
At one per cent, the same shape with less room
Ten times the noise moves every number and keeps the handover.
At 1% noise per sample the best unregularised grid is 24 points, at 0.1472, and on every grid up to 24 points λ changes nothing. On 40, 48, 64 and 96 points the best λ is 3.2·10⁻² on all four, and the errors are 0.1402, 0.1435, 0.1448 and 0.1375 — within 5 per cent. The 96-point grid with its λ is 7 per cent better than the best grid without one.
The four curves say the same thing as at 0.1%, and the three fine-grid minima now sit at the same λ exactly on the grid of λ values used. The gain from the fine grid is smaller than at lower noise — 7 per cent against 9 — because at 1% noise the right number of components to keep is about 24, a 24-point grid can nearly hold them, and there is less resolution for a finer grid to add.
How much the best λ moves between fine grids
“The same number” deserves a closer look than a grid of four λ values a decade gives it, and it gets one. Read at twenty values a decade on grids of 40, 48, 64, 80 and 96 points, the best λ at 1% noise is 10⁻¹·⁵⁰, 10⁻¹·⁵⁵, 10⁻¹·⁴⁵, 10⁻¹·⁷⁰ and 10⁻¹·⁶⁰; at 0.1% it is 10⁻²·⁵⁰, 10⁻²·³⁰, 10⁻²·³⁰, 10⁻²·⁴⁵ and 10⁻²·⁴⁵; at 0.01% it is 10⁻³·¹⁰ three times, 10⁻³·⁰⁵ and 10⁻³·⁰⁰. The spread is a quarter of a decade at most, a tenth at the lowest noise, and it has no consistent direction: fitted against the number of grid points, the slopes are −0.35, −0.06 and +0.25.
There is a short argument that λ ought to drift. A finer grid with the same noise per sample carries more measurements, each coefficient of the data grows like the square root of the number of samples while each coefficient of the noise does not, and a Tikhonov parameter that balances the two should then fall like n^(−1/2) — a fifth of a decade between 40 and 96 points, a slope of −0.5 at every noise level. The measured slopes are not that. What the argument leaves out is how flat the error is near its minimum: on every fine grid at every noise level, moving λ a quarter of a decade either side of its best costs at most 7.5 per cent, and at 0.01% noise at most 1 per cent. A drift of a fifth of a decade lives inside that flat stretch, and sixteen draws place the minimum within it no better than to about its width.
So the honest statement is narrower than the heading might suggest and more useful than a law would be. The λ worth using on a fine grid can be chosen once, on any fine grid, to within a factor of two; the cost of not knowing which fine grid it was chosen on is a few per cent; and whether a real drift of the size the argument predicts sits underneath is not something this measurement can say.
The grids in between
Between the best unregularised grid and 40 points each grid, given its own best λ, reaches an answer better than the coarse grid’s and worse than the fine grids’. At 1% noise the best of the grids from 26 to 36 points is 34 points with λ = 3.2·10⁻², at 0.1451 — 1.4 per cent better than 24 points without λ and 5.5 per cent worse than 96 points with it. At 0.1% the best in between is 36 points, at 0.1249: 3.9 per cent better than 26 without λ and 6 per cent worse than 96 with it.
That is the shape the trade hypothesis did not predict. If a grid and a λ were two settings of one amount of smoothing, a grid of 30 points with a light λ would reach what 96 points with a heavier one reaches, and it does not: at 0.1% noise it reaches 0.1384 against 0.1178, 17 per cent worse. A grid in that range is too fine to be its own filter — its unregularised solve has already crossed the grid essay’s cliff — and too coarse to represent what a well-regularised answer can show. λ repairs the first problem and cannot touch the second.
The grid essay’s filter factors say why the second is the one that binds. A grid of n points keeps the fine operator’s leading components softly — about 0.7 of each just below its size — and folds part of the components just above its size back in. λ on that grid acts on the grid’s own singular vectors, which already carry that softness and that folding; it can remove noise from them and cannot sharpen them. On a fine grid the soft shoulder sits far past any component λ keeps, and λ’s filter is the only one in play.
Where λ starts to help, at three noise levels
Read across the three noise levels at once, the handover moves with the noise, and the λ it hands over to moves in a regular way.
The grid at which λ first helps is 24 points at 1% noise, 26 at 0.1% and 34 at 0.01% — the grid essay’s best unregularised grids, to the point. That is not a coincidence to note but a restatement: the best unregularised grid is by definition the finest grid on which the noise has not yet reached the answer, and λ has something to do precisely when it has. What the map adds is the λ on the other side. On fine grids it is 3.2·10⁻², 5.6·10⁻³ and 10⁻³, a factor of about 5.6 for each factor of ten in the noise, so the best λ on this problem falls like the noise to the power of about three quarters. That is a fact about this signal and its step, measured at three points, and not a rate anything here predicts; the method that cannot use a smooth answer is the reminder that such rates belong to the smoothness of the answer and change when it does.
At a hundredth of a per cent the fine grid reaches the fine grid’s Tikhonov answer
The lowest noise level completes the picture and adds a comparison worth having.
At 0.01% noise the best unregularised grid is 34 points, at 0.1288; λ helps the 32- and 34-point grids by less than 2 per cent and the 36-point grid by 8. From 36 points on the best λ is 10⁻³, and the fine grids reach 0.1122, 0.1145, 0.1116 and 0.1115 at 40, 48, 64 and 96 points — 13 per cent better than the best coarse grid, and equal, to three figures, to the best Tikhonov answer on 64 points with its λ chosen separately for every draw.
That last equality is the handover stated in one number. The grid essay compared its best coarse grid with a truncation and a Tikhonov solution on a fine grid and found gaps of 4, 11 and 16 per cent. Every one of those gaps is closed by the fine grid with the right λ, and at the lowest noise the 96-point grid with a single λ for all sixteen draws reaches what the 64-point grid reaches with sixteen λ chosen against the truth. Four knobs and one floor found four regularisers meeting at one floor on a fixed grid; on this evidence the grid itself is a knob that stops mattering once it is past that floor’s resolution.
What the handover means for choosing a grid
Put in the order a person solving such a problem makes the decisions, the three findings reduce to one rule and one warning.
Choose the grid for resolution and let λ handle the noise. Past about 40 points on this problem, any grid with its best λ is within 5 per cent of any other, and the finest is usually the best. The grid’s ill-conditioning, which on 96 points is total, costs nothing, because the components that carry it are the components λ removes — the regularised answer is a second blur, narrower than the first, and its width is set by λ rather than by the grid.
A coarse grid is a complete choice, not a partial one. On a grid at or below the noise’s own cut no λ adds anything, so a coarse grid chosen for memory or speed is also a choice of regularisation that nothing later can adjust. The grid essay’s advice — ask how many components the grid admits before choosing λ — is sharpened into a test with two outcomes: if the grid admits fewer components than the noise can afford, λ is irrelevant and the answer is as good as the grid; if more, λ is necessary and the grid is irrelevant.
The number that decides which outcome applies is available before any λ is chosen, and from two places. One is the grid’s condition number against the noise: the grid essay found the unregularised error doubling once κ reached a few times the reciprocal of the noise, which is the condition number as an amplifier read as a threshold rather than a price, and on this problem it puts the handover at 24, 26 and 34 points to within a grid step. The other is the data’s own coefficients: the index at which the Picard plot finds the coefficients stop falling is an estimate, with its known overshoot, of how many components the noise leaves usable, and a grid with fewer points than that has nothing for λ to take away. Neither needs the truth, and neither needs a λ. A convergence study done by refining a grid with a fixed λ, or with none, sees neither: a different equation on every grid is the field-level warning that such a study measures the discretised equations rather than the problem, and here the discretised equations change their character exactly at the handover.
The range between is the one to avoid. A grid fine enough to need λ and coarse enough to limit the answer pays for both, and on this problem it is 5 to 17 per cent worse than a finer grid with the same kind of λ.
Choosing the λ itself is the problem every essay in choosing without knowing’s line has measured, and nothing here makes it easier. What the measurement does say is that the λ chosen by any of those rules on one fine grid transfers to another fine grid to within a factor of two, which is the resolution at which the rules disagree with the oracle anyway.
What was not measured
The λ reported for each grid is the best single λ for sixteen draws, chosen against the truth. No parameter-choice rule was run on the grids, and whether the discrepancy principle or cross-validation lands on the same λ on every fine grid — whether the rules share the grid-independence the oracle shows — is not measured. The drift of λ with the number of samples that a simple argument predicts is neither confirmed nor excluded, for the reason given above.
Only Tikhonov regularisation was put on the grids, and only the sampled kernel read as hat functions. A better discretisation is a weaker filter measures what an integrated kernel and a cubic-spline reading do to the unregularised solve; the same sweep with λ on those discretisations is the direct combination, and it is not drawn. The noise is independent per sample on every grid, and a real instrument’s noise correlated over a fixed physical length would become more correlated per sample on finer grids, which is exactly the regime the fine grids occupy.
Still open: do the rules find the λ the oracle does, and does the handover survive correlated noise?
The parameter rules across grids. The discrepancy principle, generalised cross-validation and the L-curve were scored on a 64-point grid. The same three on grids of 40, 64 and 96 points at each noise level would say whether their choices are as grid-independent as the oracle’s, and whether a rule’s cost against the oracle grows or shrinks as the grid is refined past the handover.
Noise with a physical correlation length. Holding the noise’s correlation length fixed in the continuous variable, rather than holding it independent per sample, makes each refinement add samples whose noise is more nearly shared. Noise that spares the answer and fools the rules found correlated noise helping the best answer at a fixed grid; whether it moves the handover finer, and whether the fine-grid λ stays one number, is the continuation.
The handover in two dimensions. An image grid of n × n points has a condition number and a number of unknowns that both grow faster with n, and a much larger share of its components near any cut. Whether the range of grids to avoid widens is the measurement that would say how much of this carries to the problems where the grid is most often chosen for memory.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Thirty-two coefficients instead of a noise level — both name filter factors, ill-posed problem, parameter choice, regularisation, tikhonov regularisation
- The corner reads the norm it is drawn in — both name filter factors, parameter choice, regularisation, tikhonov regularisation
- An answer that changes with the seed — both name filter factors, ill-posed problem, truncated svd
- One draw in twenty — both name ill-posed problem, parameter choice, tikhonov regularisation
- What a cheap preconditioner has to leave alone — both name deconvolution, filter factors, tikhonov regularisation
- A nearest point that is not there — both name ill-posed problem, truncated svd
Named objects
A flat tag is an object no other essay names yet.
Condition numberDeconvolutionDiscretisationFilter factorsIll-posed problemParameter choiceRegularisationTikhonov regularisationTruncated SVD