The L-curve, and where four rules put λ
At its defaults it draws the l-curve, and where four rules put λ. The norm of the solution against the norm of its residual, on logarithmic axes, as λ sweeps eight decades. The curve has a corner: to the left of it the noise is being amplified and to the right the signal is being thrown away. Four points are marked — the three rules that use only the data, and the oracle, which requires the exact answer and is not a method.
parameter-rules is one function in lib/figures/regular.js —
regularisation — the filter, the corner, and the answer that is a choice. Everything below came out of it during this build, at
arguments taken from the essays rather than invented for this page. A figure here is the
figure a reader meets in an essay, and if the generator changes, this page changes with it.
At its defaults
Drawn even though every essay passes arguments — which on this site is every essay, at 100% of placements since the standard pass. A default nothing exercises is a trap for the next essay to call this with none, and this is the page where a default that has drifted from the figures around it becomes visible.
The norm of the solution against the norm of its residual, on logarithmic axes, as λ sweeps eight decades. The curve has a corner: to the left of it the noise is being amplified and to the right the signal is being thrown away. Four points are marked — the three rules that use only the data, and the oracle, which requires the exact answer and is not a method.
noise: 0.001
The arguments are the ones A parameter chosen on a smaller problem passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
The norm of the solution against the norm of its residual, on logarithmic axes, as λ sweeps eight decades. The curve has a corner: to the left of it the noise is being amplified and to the right the signal is being thrown away. Four points are marked — the three rules that use only the data, and the oracle, which requires the exact answer and is not a method.
show: "target", truth: "smooth", order: 0
The arguments are the ones A rule that has to be told how good its answer will be passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
Against λ on logarithmic axes: the true noise share of the plotted norm on one draw, and the share estimated from the noise norm alone, which is what a rule would have. The two agree to within a factor of 1.09 at the best λ. The L-curve's corner sits at a share of 0.102 and costs 25.10 times the oracle; the best λ sits at 0.0037. A rule that stops at the first λ whose estimated share has come down to 0.0035 costs 1.000 times the oracle over 20 draws.
show: "target-sweep"
The arguments are the ones A rule that has to be told how good its answer will be passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
The median cost over 20 draws of choosing λ where the estimated noise share reaches a target, against that target on logarithmic axes, for four signals. Every curve has a minimum at or near one — the rule is nearly exact given the right target — and the minima are at 0.0035 for two bumps, 0.0141 for seven oscillations, 0.0398 for bumps and a step, 0.2239 for four spikes. The L-curve's corner reads a share near 0.14 on all four, and costs 25.10, 8.03, 1.38, 1.00 times the oracle.
show: "shares"
The arguments are the ones A rule that has to be told how good its answer will be passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
For each of five signals under each of three penalties, the median over 60 draws of how large the noise's part of the plotted norm is against the signal's, at the L-curve's corner and at the best λ. The corners sit between 0.11 and 0.21 in all fifteen. The oracles range from 0.004 to 0.42, and the corner's cost grows with the distance between the two.
show: "target-transport"
The arguments are the ones A rule that has to be told how good its answer will be passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A four-by-four table. Each row is the target share that is best for one signal under ‖x‖; each column is the signal that target is then used on; each cell is the median cost over 20 draws as a multiple of that column's own oracle. The diagonal runs from 1.000 to 1.006 — every target is nearly exact where it was chosen. Off the diagonal the largest is 37.8, the target chosen on four spikes applied to two bumps. The targets themselves run from 0.0035 to 0.224.
What it checked while drawing
Every figure above checked its own claims on the way to being drawn, and a claim that failed
would have stopped the picture rather than shipped a wrong one. Those checks used to leave
no trace at all: a passing one returned true and the only evidence the figure had
checked anything was that nothing crashed. The list below is what they actually said, collected
by running this generator with an observer installed — not a description of
what it is believed to check.
52 distinct claims across 6 sets of arguments, grouped below by shape — because most of them are one sentence with a different number in it, and how many separate times that sentence was put to the test is the informative part.
the percentiles of rms, last 8 are ordered — checked 3 times
a corner view that exists: share or curve
a derivative penalty to pair with ‖x‖
a draw inside the thousands the tails are measured over
a grid fine enough to place the optimum and coarse enough to draw
a lower floor never makes the rule fail less often
a noise level small enough to be noise
a parameter-rules mode that exists
a penalty of order zero, one or two
a penalty this study draws
a positive definite normal matrix
a signal the dial draws
a signal this study draws
a target exists that costs almost nothing on every pairing
a truth this field constructs
and at least one is ruinous on another
and at least one pays a real price for not knowing the answer
and choosing the wrong one of them is a real cost
and one of the minima is near the oracle
and the optimum is localised rather than a plateau
and the share divided by that error does not
and the targets are not one target
and told a third of the noise, it is not
by several times less spread
cross-validation has a tail the quasi-optimality criterion does not
discrepancy principle does not beat the oracle
enough draws for a band, and few enough to redraw on a slider
enough draws for a median and few enough for fifteen studies
enough draws for a percentile and few enough to redraw
enough draws for a share to be a share
enough draws to reach a one-in-a-hundred tail
every corner sits near one share
every target is nearly exact on its own signal
generalised cross-validation does not beat the oracle
L-curve corner does not beat the oracle
mixing the two penalties barely beats the better one alone
the best pair is no worse than either penalty alone
the cliff is an understatement and its spread is ordered
the corner does not beat the oracle
the corner reads the share this ladder found it reads
the corner sits at a noise share of the order the study measures
the estimated share falls as λ rises
the oracle's share spans orders of magnitude
the percentiles of median, last 32 are ordered
the percentiles of median, past the crossing are ordered
the percentiles of rms, past the crossing are ordered
the quasi-optimality criterion reads the solutions, so the sweep must keep them
the rule's choice is a local minimum of its objective
thirty-two coefficients keep the principle off its cliff
told the truth, the principle is close to its oracle
Against the rule
The rule does not apply to it. It factorises nothing, so there is no residual it could be withholding. That is worth stating rather than leaving blank: a site that reported the rule as satisfied by every generator would be counting mostly generators the rule never reached.
Across the library: the rule bites on 217
of 397 generators —
199 print a residual and
18 are exempt with a published reason;
180 factorise nothing.
Read from lib/residual-rule.js, which is the same body the gate enforces from,
and the gate's last check fails the build if this page and it disagree about any generator.
Where it is called
Changing this generator changes every figure on this list. That is what makes the list worth publishing rather than keeping in a check script.
A parameter chosen on a smaller problem
Inside a hybrid method the regularisation parameter is chosen on a 25×24 problem rather than a 64×64 one. The rule that reads a residual transfers exactly; the rule that reads a trace is biased by exactly two grid steps at twenty-four steps and one at forty, at every noise level from 10% to 0.1%.
Regularisation, and the answer that is chosenA rule that has to be told how good its answer will be
The L-curve's corner reads a noise share of 0.10 to 0.21 across fifteen pairings of signal and penalty, and the share the best λ sits at runs from 0.0037 to 0.43 — a factor of a hundred and fifteen. A rule aimed at the right share is within a few per cent of the oracle on every one of them. The right share is about a third to four-fifths of the relative error that λ will achieve, which is the number the answer was wanted for.
Regularisation, and the answer that is chosenA second penalty is not a second parameter
Penalise ‖x‖ and ‖L₁x‖ at once and there are two λ to choose. Over ten draws on five signals the best pair beats the better single penalty by between 0.00% and 3.1%, and one of its two parameters is exactly zero on 30 to 70 per cent of draws. Choosing the wrong one of the two costs up to 54%. The surface is a choice between two curves with a knob nobody needs.
Regularisation, and the answer that is chosenA third penalty on a flat floor
Two penalties at once were worth three per cent at most, and the explanation offered was that one penalty already does the work — which predicts that a third buys less still and that the best choice sits on a face of the parameter cube. The third buys a median of exactly nothing on all five signals and 0.66% on its best draw. But on twelve draws of forty the best triple does use all three, and on every one of them the nearest point with a penalty switched off is within that same 0.66%. The minimum is not on a face; it is on a floor so flat that where it lands is noise. Choosing the right single penalty is worth a factor of two.
Regularisation, and the answer that is chosenChoosing without knowing
Three published rules for choosing a regularisation parameter, scored against an oracle that requires the exact answer and is therefore not a method. Generalised cross-validation lands on the oracle's λ exactly; the discrepancy principle costs 6%; the L-curve costs 129%. And told a noise level ten times too small, the discrepancy principle's error goes from 0.112 to 10,449.
Regularisation, and the answer that is chosenOne draw in twenty
Sixteen draws gave generalised cross-validation a worst case of 12%. A thousand draws at each of five noise levels give it a second answer on four to six in every hundred, ten to seven million times worse than the oracle, while its median stays among the best of five rules. The quasi-optimality criterion, told nothing either, never costs more than 1.41 in five thousand draws. The share settles by a thousand draws, and letting the search look further down more than triples it.
Regularisation, and the answer that is chosenThe corner reads the norm it is drawn in
The L-curve was the costliest rule this field scored, and the cost was not the rule's. On the same sixty draws, with the same best achievable error, the corner of ‖x‖ against the residual costs 1.53 times the oracle and the corner of ‖L₁x‖ costs 1.003. Across five signals and three penalties the corner lands wherever amplified noise is between a tenth and a fifth of the norm being plotted, and it finds the oracle only when the oracle happens to sit there — twenty-nine times too costly on a smooth signal under ‖x‖, within half a per cent on four spikes.
Regularisation, and the answer that is chosenThirty-two coefficients instead of a noise level
The discrepancy principle has to be told the noise, and told too little it does not degrade — it falls off a cliff, at 0.80 of the truth when the noise is 10% and at 0.58 when it is 0.001%, exactly where the understatement forces the filter past its best truncation. The missing number is in the data. The root mean square of the last thirty-two coefficients never sends the rule over the cliff at or below 1% noise in four hundred draws, where eight coefficients with the same median do so thirty-five times.