Noise level — where it appears
Named by 3 essays across 2 fields — each of them below, with the objects they name alongside it.
A rule that has to be told how good its answer will be
The L-curve's corner reads a noise share of 0.10 to 0.21 across fifteen pairings of signal and penalty, and the share the best λ sits at runs from 0.0037 to 0.43 — a factor of a hundred and fifteen. A rule aimed at the right share is within a few per cent of the oracle on every one of them. The right share is about a third to four-fifths of the relative error that λ will achieve, which is the number the answer was wanted for.
Four folds feed one rule
GCV's guard failed when told a noise level estimated from one held-out quarter of the samples, and the proposal was four folds: hold out each quarter in turn, so the estimate rests on every sample. The prediction was that the guard would then refuse about as many dips as it does told the true noise, three of 45, and that the discrepancy principle's worst draw would fall below four. The second half holds and then some: the four-fold estimate is never below the true noise on any of 960 draws, and the discrepancy principle told it never misses by more than 1.19 times the oracle — the same as told the truth. The first half fails. The guard refuses no dips at the 1% threshold and 44 of 45 at 10%, along with 520 real minima, because the ratio it reads moves against the true one: at the real minima the guard's ρ̂ is confined between 0.38 and 1.01 and anti-correlated with the ρ̂ the true noise gives. The one held-out quarter was anti-correlated the same way.
A unit is a statement about the noise
Total least squares minimises one Frobenius norm over every column of [A b], so the unit a column is written in is a claim about how noisy it is. On a 60-row fit with two measured regressors, rewriting one column in units from 10⁻⁴ to 10⁴ times the recorded ones leaves ordinary least squares exactly where it was and moves total least squares by half again, continuously, between two estimators with their own names: the reverse regression of that column on the others, and the fit that treats it as exact. Its correction is split among the columns as the squares of the coefficients, which the units set and the noise never enters. Dividing each column by its noise level makes every unit give one answer, the best of five estimators on all three placements of the noise tried, and two replicate readings a column are enough to get most of the way there when the recorded units were badly wrong.
Named alongside it
The objects these essays reach for when they reach for this one.
Discrepancy principleGeneralised cross-validationParameter choiceRegularisationTikhonov regularisationCross-validationErrors-in-variablesExact ground truthFilter factorsGeneral form regularisationL-curveLeast-squares