A rule that has to be told how good its answer will be
Worth reading first: Choosing without knowing · When the answer is a choice.
The corner reads the norm it is drawn in left the L-curve with a diagnosis rather than a verdict. The corner is not a detector of the best λ and never was; across five signals and three penalties it lands wherever the amplified noise is between a tenth and a fifth of the norm being plotted, and it finds the oracle only when the oracle happens to sit there. That reading explained a cost — 1.53 times the oracle under ‖x‖ and 1.003 under ‖L₁x‖ on the same draws — without saying whether the cost was reparable.
A diagnosis of that shape suggests a rule immediately. If the corner is a share detector, then aim at the share directly: estimate how much of the plotted norm is amplified noise, and stop where that reaches a chosen level. Such a rule would not need a corner at all, and so would not need the smoothing that finding a corner on a discrete curve requires — which is a parameter of the rule for picking a parameter, and was named as such when the corner was first scored.
The rule works. It is nearly exact, on every pairing, including the one where the corner costs twenty-five times the oracle. What it needs is a number, and the number turns out to be the one the whole exercise exists to obtain.
The share a rule can have
The share the earlier measurement measured is ‖L R(λ)e‖ over ‖L R(λ)x‖ — the noise’s part of the plotted norm against the signal’s — and both halves of it require knowing which part of the data is noise. Nobody has that.
What a code can have is δ = ‖e‖, the single number the discrepancy principle already needs. For noise that is white and scaled to that norm, the expected size of its contribution is in the Frobenius norm, because R(λ) acting on the data’s coefficients is a fixed matrix at each λ and the noise coefficients are exchangeable. The denominator is harder to have and easier to approximate: is on the vertical axis of the L-curve already, and wherever the share is small it is the signal’s part to within that share.
So the estimate is , and it costs one Frobenius norm per λ — which depends on the operator and the penalty and not on the draw, so it is computed once for a problem rather than once for a data set.
The exchangeability is the only probabilistic step and it is worth naming. The noise is drawn white and then scaled so that ‖e‖ is exactly δ, which makes its coefficients in the operator’s left singular basis a fixed-norm vector with no preferred direction. Under that, the expected squared size of is times the squared Frobenius norm of — an exact statement about the ensemble and an approximate one about any single draw, with a relative spread that falls like and is therefore a few per cent at . Everything else in the estimate is deterministic.
It tracks the truth closely. At the best λ, on the two-bump signal under ‖x‖ at 0.1% noise, the true share is 0.0037 and the estimate is 0.0034. On four spikes it is 0.1642 against 0.1886. Across the fifteen pairings the estimate is within about twenty per cent of the share everywhere, which is more than enough for a rule that is going to be aimed at a level rather than at a crossing.
The estimated share falls as λ rises, monotonically over the whole grid, and that direction is the whole reason damping is worth doing: more damping suppresses amplified noise faster than it suppresses signal. So the rule is a level set — walk up the grid and stop at the first λ whose estimated share has come down to the target.
What the rule is worth when it is aimed correctly
Sweep the target across three decades and score the rule on each pairing against that pairing’s own oracle. Every curve has a minimum, and every minimum is at essentially one.
On the two-bump signal under ‖x‖ — where the corner costs 25.1 times the oracle, the most expensive cell in the whole table — the best target is 0.0035 and the rule costs 1.000. On the oscillating signal the corner costs 8.03 and the rule at a target of 0.0141 costs 1.006. On four spikes the corner costs 1.00 and the rule costs 1.003. Across all fifteen pairings the worst the rule does, aimed at its own best target, is 1.025.
That settles the first half of the question and it settles it against the obvious reading of the earlier measurement. The corner’s cost is not a detection error. A share is the right thing to aim at; the amplified noise’s fraction of the plotted norm is a sufficient statistic for the best λ, to a few tenths of a per cent, on every signal and penalty tried. What the corner gets wrong is where to aim.
And the target is not one number
The best targets are 0.0035, 0.0141, 0.0398 and 0.2239 for the four signals under ‖x‖. Across all fifteen pairings the share the oracle sits at runs from 0.0037 to 0.4282 — a factor of one hundred and fifteen. The corner’s own share runs from 0.101 to 0.210, a factor of 2.09.
So the corner is an accurate instrument reading a quantity that is not constant, which is the least useful combination available. An inaccurate instrument reading a constant would at least be improvable.
The practical consequence is measurable directly: take the target that is best for one signal and use it on another.
The diagonal runs from 1.000 to 1.006. Off it, the target chosen on four spikes — 0.2239, which sits just above the band the corner reads — costs 37.8 times the oracle on the two-bump signal and 11.3 on the oscillating one. That is worse than the corner itself, which costs 25.1 and 8.03 on those two, and it is worse for the same reason and a little more of it.
The table also says which way to err, and the answer is unambiguous: downward. The smallest of the four targets, 0.0035, costs at most 2.12 anywhere in the table. The largest costs 37.8. A target below the right one stops at a larger λ than the oracle’s — more damping, a smoother answer, some signal thrown away — and the L-curve’s own geometry says why that is cheap: to the right of the corner the curve is nearly flat, so a long walk along λ buys a small change in the answer. To the left it is steep. This is the field’s oldest piece of practical advice, that over-regularising is safer than under-regularising, arriving as a number rather than as a maxim.
The rule has no tail
A median of 1.000 is the statistic that hid two earlier findings in this field, so it is worth asking what the rule’s worst draw looks like before anything is concluded from its middle.
Over forty draws at 0.1% noise, aimed at its own best target, the rule’s ninetieth percentile is 1.066, 1.016, 1.014 and 1.097 on the four signals under ‖x‖, and its maximum is 1.164, 1.024, 1.044 and 1.164. There is no tail. The worst of a hundred and sixty draws is sixteen per cent above the oracle.
That is a sharper contrast than the medians give. One draw in twenty found generalised cross-validation returning a second answer on four to six draws in a hundred, ten to seven million times worse than the oracle, while its median stayed among the best of five rules — a rule whose failure mode is invisible in every summary except a count. The share rule cannot fail that way, because it is not minimising anything. It walks a monotone curve and stops at a level, so there is no second minimum for it to find and no search floor to be sensitive to. The quasi-optimality criterion, which never cost more than 1.41 in five thousand draws, is the other rule in this field with that property and it has it for a different reason.
So the rule’s whole difficulty is concentrated in one number, which is an unusual and rather clean place for a difficulty to be. Everything about the draw, the smoothing, the search and the tail is settled; what is left is the target.
What the target is a property of
Fifteen numbers spanning two orders of magnitude are not fifteen accidents. Ordered by their share, the pairings are ordered by something, and the something is visible as soon as the shares are plotted against the accuracy each best λ actually achieves — measured in the same norm the share is measured in, which for a derivative penalty is not the ordinary relative error.
The fifteen points lie in a band. Divided by that error the shares run from 0.268 to 0.813 — a factor of 3.04, against their own factor of 115. The share the best λ sits at is, to within a factor of three, between a third and four-fifths of the relative error the answer at that λ will have.
The relation has a reason and the reason is why it cannot be exploited. At the best λ the two halves of the error are balanced: the part from throwing signal away and the part from letting noise through are of the same order, which is what makes the λ best. The second of those parts is the amplified noise, and the share measures it against the signal’s own size. So the share at the oracle is the noise half of a balanced error, expressed as a fraction — which is the error, up to the balance ratio and up to the difference between measuring a norm and measuring a difference of norms. Neither of those is constant, and the band is three wide rather than one.
The relation is not an artefact of one noise level either. At 1% noise the shares span 33 and their ratio to the error spans 4.55; at 0.01% they span 646 and 7.30. The band widens as the problem gets easier, which is what a proportionality with a drifting constant looks like, and the shares move by nearly three orders of magnitude while the ratio moves by a factor of two.
What that means the rule needs
A rule aimed at a target share is a rule that must be told, within a factor of three, the relative error its own answer is going to have.
That is not a rule. It is a restatement of the problem with an extra step in it. The accuracy of the regularised answer is the quantity the regularisation is being done to obtain, and a method that needs it as an input has gained nothing — the same closed circle the oracle λ itself sits in, and the reason the oracle is described in this field as a reference rather than a method.
It is worth being precise about what has and has not been shown. It has not been shown that no data-driven quantity predicts the target; three were tried here and a fourth in the earlier measurement, and what has been shown is that the obvious candidate — the corner’s own share — is the worst constant in the table rather than the best. The roughness of the truth in the penalty’s norm does not predict it: that quantity falls as the penalty order rises and the target rises, and the basis a penalty works in is where that mismatch would have to be resolved if it can be. Nor does the noise level alone: the fifteen targets at one noise level already span a hundredfold.
What has been shown is stronger than a failed search, because it identifies what a successful one would have to find. Any quantity that predicts the target predicts the achievable accuracy to within a factor of three. A rule that estimated the achievable accuracy that well would be a remarkable thing to have independently of any of this, and its existence is not obviously ruled out — the noise level itself is estimable from the data to within a factor that does not break the discrepancy principle, which is the same shape of problem one level down.
What the corner is for, after this
Nothing here recovers the L-curve as a rule and nothing here needs to. Its corner reads one quantity accurately and that is a useful thing to know about a curve everybody draws.
What it does supply is a reading of the corner’s failures that predicts them rather than describing them. The corner costs little when the answer is inaccurate and much when the answer is accurate, because the share it reads is fixed near 0.14 and the share it should read is proportional to the accuracy. On four spikes, where the best achievable relative error is 0.571, the target is 0.20 and the corner costs 1.00. On two bumps, where the best achievable error is 0.0045, the target is 0.0040 and the corner costs 25.1. That is a prediction about any new pairing: the corner will be expensive exactly where the problem is easy.
Four knobs and one floor is the reason the prediction is worth having across methods rather than inside one: a truncation, a Tikhonov parameter, a step count and a randomised rank reach within three per cent of each other on this problem, so a statement about where a rule for one of them fails is a statement about a floor all four are aiming at.
It is a slightly uncomfortable conclusion for a rule that was designed to be read by eye. The corner is most trusted where the L-curve is most sharply cornered, and the L-curve is most sharply cornered where the answer is well determined — which is where the corner is furthest from the best λ.
What this does not settle
One operator throughout: the 64-point Gaussian blur, at three noise levels, with five signals and three penalties. The share-to-error relation is measured on 45 pairings in total and every one of them shares a singular value spectrum.
The estimate of the share uses δ exactly. Every criticism made of the discrepancy principle for needing δ applies here unchanged, and the rule has not been swept over misstated δ — the cliff measured for the discrepancy principle may be in a different place for a rule that reads a Frobenius norm rather than a residual.
The targets reported are the minima of a sweep over twenty draws each, so each is the best of a discrete grid at a resolution of 0.15 of a decade, widened from a tenth after the first version of the grid stopped at 0.004 and put the easiest pairing’s own best target outside it, and “1.000 times the oracle” means the median over those draws. The transport table reads costs at the nearest grid target rather than by re-sweeping, which is why the diagonal is 1.000 to 1.006 rather than exactly one.
The noise is white in every measurement. Correlated noise changes the Frobenius-norm estimate of the share in a way that is not measured, and it already breaks two of the standard rules while improving the best answer available.
Still open: whether the target can be estimated, and what a surface would have
Estimating the achievable accuracy. The rule needs the answer’s own error to within a factor of three. An estimate of it exists in principle — the unbiased predictive risk estimator estimates the predictive error from δ and the trace, and that is a different error in a different norm, but the two are related by the operator. Whether the risk estimate, converted through the operator, lands inside the band measured here is a direct test and would turn the whole construction into a method.
The share on the projected problem. A parameter chosen on a smaller problem found the residual-reading rule transferring exactly from a projected problem to the full one, and the trace-reading rule biased by a fixed number of grid steps. The share is a Frobenius norm of a matrix, so it is a trace-like quantity and ought to inherit the second behaviour rather than the first. Whether it does, and by how many steps, is what would say whether this rule can live inside a hybrid method.
Two penalties and no corner. The corner reads the share of the plotted norm, and a problem penalised in both ‖x‖ and ‖L₁x‖ has two norms and no single curve. Whether such a problem has a corner at all, what it would be a corner of, and whether the two share conditions can be satisfied at once is the question the L-curve’s whole geometry runs out at.
Which way the constant of proportionality moves. The ratio of share to error runs 0.19 to 0.88 across three noise levels and fifteen pairings, and the band widens as the noise falls. Whether that widening is a property of the balance at the optimum — which would make it predictable — or of the grid resolution at very small λ, is separable by one finer sweep.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rule that is wrong in the right direction — both name discrepancy principle, generalised cross-validation, l-curve, parameter choice, tikhonov regularisation
- A second blur, narrower than the first — both name filter factors, generalised cross-validation, regularisation, tikhonov regularisation
- Where the grid hands over to λ — both name filter factors, parameter choice, regularisation, tikhonov regularisation
- A parameter that counts steps — both name discrepancy principle, filter factors, tikhonov regularisation
- A step that is not a unit of work — both name discrepancy principle, filter factors, tikhonov regularisation
- A stopping rule that follows the run it is given — both name discrepancy principle, parameter choice, tikhonov regularisation
Named objects
A flat tag is an object no other essay names yet.
Discrepancy principleFilter factorsGeneral form regularisationGeneralised cross-validationL-curveNoise levelOracleParameter choiceRegularisationTikhonov regularisation