Concept

Noise level — where it appears

The size of the error in the data a computation is given, measured in the same norm as the residual. Regularisation rules such as the discrepancy principle stop or set a parameter where the residual reaches it, so a level that is guessed wrong moves the answer with it.

Named by 3 essays across 2 fields — each of them below, with the objects they name alongside it.

where each λ sitsoracle's share0.0037corner's share0.1corner's cost25and what a target costsbest fixed target0.0035its cost1estimate ÷ truth at the oracle0.9210⁻⁸10⁻⁶10⁻⁴10⁻²110⁻³10⁻²10⁻¹110¹10²10³10⁴λnoise part ÷ signal parttarget 0.0035oracle λ: share 0.0037corner: share 0.102solid: the share on this draw · dashed: the share a rule can estimatethe rule is a level set of the dashed curve

A rule that has to be told how good its answer will be

The L-curve's corner reads a noise share of 0.10 to 0.21 across fifteen pairings of signal and penalty, and the share the best λ sits at runs from 0.0037 to 0.43 — a factor of a hundred and fifteen. A rule aimed at the right share is within a few per cent of the oracle on every one of them. The right share is about a third to four-fifths of the relative error that λ will achieve, which is the number the answer was wanted for.

regularisation · Parameter choice
worst of 960, over the oraclediscrepancy, four folds, worst1.2discrepancy, told the noise, worst1.2discrepancy, one held-out quarter, worst18110¹10²10³10⁴samples per unknownerror ÷ oracle's, worst draw1.251.524GCV with the four-fold guarddiscrepancy, one held-out quarterrightmost minimumdiscrepancy, four foldsdiscrepancy, told the noise960 draws, four ratiosfolds fix one rule and not the other

Four folds feed one rule

GCV's guard failed when told a noise level estimated from one held-out quarter of the samples, and the proposal was four folds: hold out each quarter in turn, so the estimate rests on every sample. The prediction was that the guard would then refuse about as many dips as it does told the true noise, three of 45, and that the discrepancy principle's worst draw would fall below four. The second half holds and then some: the four-fold estimate is never below the true noise on any of 960 draws, and the discrepancy principle told it never misses by more than 1.19 times the oracle — the same as told the truth. The first half fails. The guard refuses no dips at the 1% threshold and 44 of 45 at 10%, along with 520 real minima, because the ratio it reads moves against the true one: at the real minima the guard's ρ̂ is confined between 0.38 and 1.01 and anti-correlated with the ρ̂ the true noise gives. The one held-out quarter was anti-correlated the same way.

regularisation · Regularisation
the same in every columnleast squares0.32total, weighted0.11total, c → 00.16total, c → ∞0.1710⁻⁴10⁻²110²10⁴00.050.10.150.20.250.30.35factor c the first column's units are changed byrelative error, medianreverse regressionrecorded unitsleast squarescolumn 1 exactweighted by the noisedots: plain total least squaresthe units choose the answer

A unit is a statement about the noise

Total least squares minimises one Frobenius norm over every column of [A b], so the unit a column is written in is a claim about how noisy it is. On a 60-row fit with two measured regressors, rewriting one column in units from 10⁻⁴ to 10⁴ times the recorded ones leaves ordinary least squares exactly where it was and moves total least squares by half again, continuously, between two estimators with their own names: the reverse regression of that column on the others, and the fit that treats it as exact. Its correction is split among the columns as the squares of the coefficients, which the units set and the noise never enters. Dividing each column by its noise level makes every unit give one answer, the best of five estimators on all three placements of the noise tried, and two replicate readings a column are enough to get most of the way there when the recorded units were badly wrong.

leastsquares · Total least-squares

Named alongside it

The objects these essays reach for when they reach for this one.

Discrepancy principleGeneralised cross-validationParameter choiceRegularisationTikhonov regularisationCross-validationErrors-in-variablesExact ground truthFilter factorsGeneral form regularisationL-curveLeast-squares

All concepts