Regularisation, and the answer that is chosen

A rule that has to be told how good its answer will be

The L-curve's corner reads a noise share of 0.10 to 0.21 across fifteen pairings of signal and penalty, and the share the best λ sits at runs from 0.0037 to 0.43 — a factor of a hundred and fifteen. A rule aimed at the right share is within a few per cent of the oracle on every one of them. The right share is about a third to four-fifths of the relative error that λ will achieve, which is the number the answer was wanted for.

Worth reading first: Choosing without knowing · When the answer is a choice.

The corner reads the norm it is drawn in left the L-curve with a diagnosis rather than a verdict. The corner is not a detector of the best λ and never was; across five signals and three penalties it lands wherever the amplified noise is between a tenth and a fifth of the norm being plotted, and it finds the oracle only when the oracle happens to sit there. That reading explained a cost — 1.53 times the oracle under ‖x‖ and 1.003 under ‖L₁x‖ on the same draws — without saying whether the cost was reparable.

A diagnosis of that shape suggests a rule immediately. If the corner is a share detector, then aim at the share directly: estimate how much of the plotted norm is amplified noise, and stop where that reaches a chosen level. Such a rule would not need a corner at all, and so would not need the smoothing that finding a corner on a discrete curve requires — which is a parameter of the rule for picking a parameter, and was named as such when the corner was first scored.

The rule works. It is nearly exact, on every pairing, including the one where the corner costs twenty-five times the oracle. What it needs is a number, and the number turns out to be the one the whole exercise exists to obtain.

The share the rule can estimate, two bumps under ‖x‖, 0.1% noiseAgainst λ on logarithmic axes: the true noise share of the plotted norm on one draw, and the share estimated from the noise norm alone, which is what a rule would have. The two agree to within a factor of 1.09 at the best λ. The L-curve's corner sits at a share of 0.102 and costs 25.10 times the oracle; the best λ sits at 0.0037. A rule that stops at the first λ whose estimated share has come down to 0.0035 costs 1.000 times the oracle over 20 draws.where each λ sitsoracle's share0.0037corner's share0.1corner's cost25and what a target costsbest fixed target0.0035its cost1estimate ÷ truth at the oracle0.9210⁻⁸10⁻⁶10⁻⁴10⁻²110⁻³10⁻²10⁻¹110¹10²10³10⁴λnoise part ÷ signal parttarget 0.0035oracle λ: share 0.0037corner: share 0.102solid: the share on this draw · dashed: the share a rule can estimatethe rule is a level set of the dashed curve
Fig. 1 The noise share of the plotted norm against λ, and the same share as a rule could estimate it, with the oracle’s λ and the corner’s marked on the curve. The distance between those two marks is the corner’s cost.

The share a rule can have

The share the earlier measurement measured is ‖L R(λ)e‖ over ‖L R(λ)x‖ — the noise’s part of the plotted norm against the signal’s — and both halves of it require knowing which part of the data is noise. Nobody has that.

What a code can have is δ = ‖e‖, the single number the discrepancy principle already needs. For noise that is white and scaled to that norm, the expected size of its contribution is (δ/n)LR(λ)(\delta/\sqrt{n})\,\|LR(\lambda)\| in the Frobenius norm, because R(λ) acting on the data’s coefficients is a fixed matrix at each λ and the noise coefficients are exchangeable. The denominator is harder to have and easier to approximate: Lxλ\|Lx_\lambda\| is on the vertical axis of the L-curve already, and wherever the share is small it is the signal’s part to within that share.

So the estimate is δLR(λ)F/(nLxλ)\delta\|LR(\lambda)\|_F / (\sqrt{n}\,\|Lx_\lambda\|), and it costs one Frobenius norm per λ — which depends on the operator and the penalty and not on the draw, so it is computed once for a problem rather than once for a data set.

The exchangeability is the only probabilistic step and it is worth naming. The noise is drawn white and then scaled so that ‖e‖ is exactly δ, which makes its coefficients in the operator’s left singular basis a fixed-norm vector with no preferred direction. Under that, the expected squared size of R(λ)eR(\lambda)e is δ2/n\delta^2/n times the squared Frobenius norm of R(λ)R(\lambda) — an exact statement about the ensemble and an approximate one about any single draw, with a relative spread that falls like 1/n1/\sqrt{n} and is therefore a few per cent at n=64n = 64. Everything else in the estimate is deterministic.

It tracks the truth closely. At the best λ, on the two-bump signal under ‖x‖ at 0.1% noise, the true share is 0.0037 and the estimate is 0.0034. On four spikes it is 0.1642 against 0.1886. Across the fifteen pairings the estimate is within about twenty per cent of the share everywhere, which is more than enough for a rule that is going to be aimed at a level rather than at a crossing.

The estimated share falls as λ rises, monotonically over the whole grid, and that direction is the whole reason damping is worth doing: more damping suppresses amplified noise faster than it suppresses signal. So the rule is a level set — walk up the grid and stop at the first λ whose estimated share has come down to the target.

What the rule is worth when it is aimed correctly

Sweep the target across three decades and score the rule on each pairing against that pairing’s own oracle. Every curve has a minimum, and every minimum is at essentially one.

What a fixed share target costs, four signals under ‖x‖, 0.1% noiseThe median cost over 20 draws of choosing λ where the estimated noise share reaches a target, against that target on logarithmic axes, for four signals. Every curve has a minimum at or near one — the rule is nearly exact given the right target — and the minima are at 0.0035 for two bumps, 0.0141 for seven oscillations, 0.0398 for bumps and a step, 0.2239 for four spikes. The L-curve's corner reads a share near 0.14 on all four, and costs 25.10, 8.03, 1.38, 1.00 times the oracle.the best target, per signaltwo bumps0.0035seven oscillations0.014and what it costs theretwo bumps1seven oscillations110⁻³10⁻²10⁻¹1110¹10²target shareerror ÷ the oracle'sthe corner reads heretwo bumpsseven oscillationsbumps and a stepfour spikesevery curve bottoms out at about oneand they bottom out in different places
Fig. 2 The rule’s median cost against its target, for four signals under ‖x‖. Each curve bottoms out at about one, and they bottom out in four different places.

On the two-bump signal under ‖x‖ — where the corner costs 25.1 times the oracle, the most expensive cell in the whole table — the best target is 0.0035 and the rule costs 1.000. On the oscillating signal the corner costs 8.03 and the rule at a target of 0.0141 costs 1.006. On four spikes the corner costs 1.00 and the rule costs 1.003. Across all fifteen pairings the worst the rule does, aimed at its own best target, is 1.025.

That settles the first half of the question and it settles it against the obvious reading of the earlier measurement. The corner’s cost is not a detection error. A share is the right thing to aim at; the amplified noise’s fraction of the plotted norm is a sufficient statistic for the best λ, to a few tenths of a per cent, on every signal and penalty tried. What the corner gets wrong is where to aim.

And the target is not one number

The best targets are 0.0035, 0.0141, 0.0398 and 0.2239 for the four signals under ‖x‖. Across all fifteen pairings the share the oracle sits at runs from 0.0037 to 0.4282 — a factor of one hundred and fifteen. The corner’s own share runs from 0.101 to 0.210, a factor of 2.09.

The noise share of the plotted norm at the L-curve's corner and at the oracle, five signals, three penaltiesFor each of five signals under each of three penalties, the median over 60 draws of how large the noise's part of the plotted norm is against the signal's, at the L-curve's corner and at the best λ. The corners sit between 0.11 and 0.21 in all fifteen. The oracles range from 0.004 to 0.42, and the corner's cost grows with the distance between the two.10⁻³10⁻²10⁻¹1noise part ÷ signal part, in the plotted normsignal · penaltycorner costevery cornerbumps and a step · ‖x‖×1.53bumps and a step · ‖L₁x‖×1.00bumps and a step · ‖L₂x‖×1.02bumps, step, offset · ‖x‖×2.81bumps, step, offset · ‖L₁x‖×1.01bumps, step, offset · ‖L₂x‖×1.12two bumps · ‖x‖×29.0two bumps · ‖L₁x‖×4.45two bumps · ‖L₂x‖×2.38seven oscillations · ‖x‖×8.97seven oscillations · ‖L₁x‖×2.29seven oscillations · ‖L₂x‖×2.51four spikes · ‖x‖×1.00four spikes · ‖L₁x‖×1.04four spikes · ‖L₂x‖×1.05blue: at the corner · red: at the oracle · medians over drawsthe corners line up
Fig. 3 The two shares in all fifteen pairings: where the corner lands, and where the best λ is. The corners line up and the oracles do not, and the corner’s cost is the distance between them.

So the corner is an accurate instrument reading a quantity that is not constant, which is the least useful combination available. An inaccurate instrument reading a constant would at least be improvable.

The practical consequence is measurable directly: take the target that is best for one signal and use it on another.

A share target chosen on one signal, applied to the other three, 0.1% noiseA four-by-four table. Each row is the target share that is best for one signal under ‖x‖; each column is the signal that target is then used on; each cell is the median cost over 20 draws as a multiple of that column's own oracle. The diagonal runs from 1.000 to 1.006 — every target is nearly exact where it was chosen. Off the diagonal the largest is 37.8, the target chosen on four spikes applied to two bumps. The targets themselves run from 0.0035 to 0.224.aimed, and carriedworst on the diagonal1worst off it38draws20target chosen onapplied totwosevenbumpsfourtwo bumps — 0.00351.002.211.281.25seven oscillations — 0.01412.121.011.061.15bumps and a step — 0.03986.481.601.001.09four spikes — 0.223937.811.31.851.00median error ÷ that signal's own oraclethe diagonal is the rule aimed correctlythe smallest target is the safe one to carry
Fig. 4 Four targets, each chosen to be best for one signal, applied to all four. The diagonal is the rule aimed correctly; everything off it is what a constant costs.

The diagonal runs from 1.000 to 1.006. Off it, the target chosen on four spikes — 0.2239, which sits just above the band the corner reads — costs 37.8 times the oracle on the two-bump signal and 11.3 on the oscillating one. That is worse than the corner itself, which costs 25.1 and 8.03 on those two, and it is worse for the same reason and a little more of it.

The table also says which way to err, and the answer is unambiguous: downward. The smallest of the four targets, 0.0035, costs at most 2.12 anywhere in the table. The largest costs 37.8. A target below the right one stops at a larger λ than the oracle’s — more damping, a smoother answer, some signal thrown away — and the L-curve’s own geometry says why that is cheap: to the right of the corner the curve is nearly flat, so a long walk along λ buys a small change in the answer. To the left it is steep. This is the field’s oldest piece of practical advice, that over-regularising is safer than under-regularising, arriving as a number rather than as a maxim.

The rule has no tail

A median of 1.000 is the statistic that hid two earlier findings in this field, so it is worth asking what the rule’s worst draw looks like before anything is concluded from its middle.

Over forty draws at 0.1% noise, aimed at its own best target, the rule’s ninetieth percentile is 1.066, 1.016, 1.014 and 1.097 on the four signals under ‖x‖, and its maximum is 1.164, 1.024, 1.044 and 1.164. There is no tail. The worst of a hundred and sixty draws is sixteen per cent above the oracle.

That is a sharper contrast than the medians give. One draw in twenty found generalised cross-validation returning a second answer on four to six draws in a hundred, ten to seven million times worse than the oracle, while its median stayed among the best of five rules — a rule whose failure mode is invisible in every summary except a count. The share rule cannot fail that way, because it is not minimising anything. It walks a monotone curve and stops at a level, so there is no second minimum for it to find and no search floor to be sensitive to. The quasi-optimality criterion, which never cost more than 1.41 in five thousand draws, is the other rule in this field with that property and it has it for a different reason.

So the rule’s whole difficulty is concentrated in one number, which is an unusual and rather clean place for a difficulty to be. Everything about the draw, the smoothing, the search and the tail is settled; what is left is the target.

What the target is a property of

Fifteen numbers spanning two orders of magnitude are not fifteen accidents. Ordered by their share, the pairings are ordered by something, and the something is visible as soon as the shares are plotted against the accuracy each best λ actually achieves — measured in the same norm the share is measured in, which for a derivative penalty is not the ordinary relative error.

The share the best λ sits at, against the accuracy it achieves, fifteen pairings at 0.1% noiseFifteen points on logarithmic axes, one for each of five signals under each of three penalties: the noise share of the plotted norm at the best λ, against the relative error that λ achieves measured in the same norm. The shares span 0.0037 to 0.428, a factor of 115. Divided by the error they span 0.268 to 0.813, a factor of 3.04. The two lines are those extremes as proportionalities.the share at the best λsmallest0.0037largest0.43spread115divided by the error theresmallest0.27largest0.81spread310⁻²10⁻¹110⁻²10⁻¹relative error at the best λnoise part ÷ signal part therebumps and a stepbumps, step, offsettwo bumpsseven oscillationsfour spikesblue: under ‖x‖ · red: under a derivative penaltythe band between the two lines is a factor of three
Fig. 5 The share at the best λ against the relative error that λ achieves, fifteen pairings, with the extremes of their ratio drawn as two proportionalities. The band between the lines is a factor of three.

The fifteen points lie in a band. Divided by that error the shares run from 0.268 to 0.813 — a factor of 3.04, against their own factor of 115. The share the best λ sits at is, to within a factor of three, between a third and four-fifths of the relative error the answer at that λ will have.

The relation has a reason and the reason is why it cannot be exploited. At the best λ the two halves of the error are balanced: the part from throwing signal away and the part from letting noise through are of the same order, which is what makes the λ best. The second of those parts is the amplified noise, and the share measures it against the signal’s own size. So the share at the oracle is the noise half of a balanced error, expressed as a fraction — which is the error, up to the balance ratio and up to the difference between measuring a norm and measuring a difference of norms. Neither of those is constant, and the band is three wide rather than one.

The relation is not an artefact of one noise level either. At 1% noise the shares span 33 and their ratio to the error spans 4.55; at 0.01% they span 646 and 7.30. The band widens as the problem gets easier, which is what a proportionality with a drifting constant looks like, and the shares move by nearly three orders of magnitude while the ratio moves by a factor of two.

The share the best λ sits at, against the accuracy it achieves, fifteen pairings at 1% noiseFifteen points on logarithmic axes, one for each of five signals under each of three penalties: the noise share of the plotted norm at the best λ, against the relative error that λ achieves measured in the same norm. The shares span 0.0192 to 0.639, a factor of 33. Divided by the error they span 0.193 to 0.880, a factor of 4.55. The two lines are those extremes as proportionalities.the share at the best λsmallest0.019largest0.64spread33divided by the error theresmallest0.19largest0.88spread4.510⁻¹110⁻¹1relative error at the best λnoise part ÷ signal part therebumps and a stepbumps, step, offsettwo bumpsseven oscillationsfour spikesblue: under ‖x‖ · red: under a derivative penaltythe band between the two lines is a factor of three
Fig. 6 The same fifteen pairings at ten times the noise. The points have moved up and to the right together, which is what a proportionality does and is not what fifteen coincidences would do.

What that means the rule needs

A rule aimed at a target share is a rule that must be told, within a factor of three, the relative error its own answer is going to have.

That is not a rule. It is a restatement of the problem with an extra step in it. The accuracy of the regularised answer is the quantity the regularisation is being done to obtain, and a method that needs it as an input has gained nothing — the same closed circle the oracle λ itself sits in, and the reason the oracle is described in this field as a reference rather than a method.

It is worth being precise about what has and has not been shown. It has not been shown that no data-driven quantity predicts the target; three were tried here and a fourth in the earlier measurement, and what has been shown is that the obvious candidate — the corner’s own share — is the worst constant in the table rather than the best. The roughness of the truth in the penalty’s norm does not predict it: that quantity falls as the penalty order rises and the target rises, and the basis a penalty works in is where that mismatch would have to be resolved if it can be. Nor does the noise level alone: the fifteen targets at one noise level already span a hundredfold.

What has been shown is stronger than a failed search, because it identifies what a successful one would have to find. Any quantity that predicts the target predicts the achievable accuracy to within a factor of three. A rule that estimated the achievable accuracy that well would be a remarkable thing to have independently of any of this, and its existence is not obviously ruled out — the noise level itself is estimable from the data to within a factor that does not break the discrepancy principle, which is the same shape of problem one level down.

Six estimates of the noise level, against the discrepancy principle's cliff, 0.1% noiseFor each of six estimators, the estimate of the noise norm divided by the truth over 400 draws, drawn as a bar for the middle eighty per cent and a whisker for the middle ninety-eight. The shaded column is where the principle's error doubles. The root mean square of the last thirty-two coefficients keeps its lowest whisker at 0.77 and sends 0 draws over; the estimates read past the Picard crossing reach down to 0.30.00.40.81.21.622.4estimate ÷ the true noise normthe cliff, 0.69the truthdoubled the errorrms, last 835 of 400rms, last 163 of 400rms, last 320 of 400median, last 324 of 400rms, past the crossing44 of 400median, past the crossing49 of 400bar: middle 80% · whisker: middle 98% · dot: medianshaded: where the error doubles
Fig. 7 Six estimates of the noise level from the data alone, scored against the cliff the discrepancy principle falls off when it is told a wrong one. The rule above needs δ, so it inherits every one of these.

What the corner is for, after this

Nothing here recovers the L-curve as a rule and nothing here needs to. Its corner reads one quantity accurately and that is a useful thing to know about a curve everybody draws.

What it does supply is a reading of the corner’s failures that predicts them rather than describing them. The corner costs little when the answer is inaccurate and much when the answer is accurate, because the share it reads is fixed near 0.14 and the share it should read is proportional to the accuracy. On four spikes, where the best achievable relative error is 0.571, the target is 0.20 and the corner costs 1.00. On two bumps, where the best achievable error is 0.0045, the target is 0.0040 and the corner costs 25.1. That is a prediction about any new pairing: the corner will be expensive exactly where the problem is easy.

Four knobs and one floor is the reason the prediction is worth having across methods rather than inside one: a truncation, a Tikhonov parameter, a step count and a randomised rank reach within three per cent of each other on this problem, so a statement about where a rule for one of them fails is a statement about a floor all four are aiming at.

It is a slightly uncomfortable conclusion for a rule that was designed to be read by eye. The corner is most trusted where the L-curve is most sharply cornered, and the L-curve is most sharply cornered where the answer is well determined — which is where the corner is furthest from the best λ.

What this does not settle

One operator throughout: the 64-point Gaussian blur, at three noise levels, with five signals and three penalties. The share-to-error relation is measured on 45 pairings in total and every one of them shares a singular value spectrum.

The estimate of the share uses δ exactly. Every criticism made of the discrepancy principle for needing δ applies here unchanged, and the rule has not been swept over misstated δ — the cliff measured for the discrepancy principle may be in a different place for a rule that reads a Frobenius norm rather than a residual.

The targets reported are the minima of a sweep over twenty draws each, so each is the best of a discrete grid at a resolution of 0.15 of a decade, widened from a tenth after the first version of the grid stopped at 0.004 and put the easiest pairing’s own best target outside it, and “1.000 times the oracle” means the median over those draws. The transport table reads costs at the nearest grid target rather than by re-sweeping, which is why the diagonal is 1.000 to 1.006 rather than exactly one.

The noise is white in every measurement. Correlated noise changes the Frobenius-norm estimate of the share in a way that is not measured, and it already breaks two of the standard rules while improving the best answer available.

Still open: whether the target can be estimated, and what a surface would have

Estimating the achievable accuracy. The rule needs the answer’s own error to within a factor of three. An estimate of it exists in principle — the unbiased predictive risk estimator estimates the predictive error from δ and the trace, and that is a different error in a different norm, but the two are related by the operator. Whether the risk estimate, converted through the operator, lands inside the band measured here is a direct test and would turn the whole construction into a method.

The share on the projected problem. A parameter chosen on a smaller problem found the residual-reading rule transferring exactly from a projected problem to the full one, and the trace-reading rule biased by a fixed number of grid steps. The share is a Frobenius norm of a matrix, so it is a trace-like quantity and ought to inherit the second behaviour rather than the first. Whether it does, and by how many steps, is what would say whether this rule can live inside a hybrid method.

Two penalties and no corner. The corner reads the share of the plotted norm, and a problem penalised in both ‖x‖ and ‖L₁x‖ has two norms and no single curve. Whether such a problem has a corner at all, what it would be a corner of, and whether the two share conditions can be satisfied at once is the question the L-curve’s whole geometry runs out at.

Which way the constant of proportionality moves. The ratio of share to error runs 0.19 to 0.88 across three noise levels and fifteen pairings, and the band widens as the noise falls. Whether that widening is a property of the balance at the optimum — which would make it predictable — or of the grid resolution at very small λ, is separable by one finer sweep.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Discrepancy principleFilter factorsGeneral form regularisationGeneralised cross-validationL-curveNoise levelOracleParameter choiceRegularisationTikhonov regularisation