Least squares by QR — where it appears
Named by 2 essays across one field — each of them below, with the objects they name alongside it.
The road that squares the problem
The normal equations are the first method every course teaches and the method no library uses. Forming AᵀA squares the condition number, and below ε = √u it does not degrade — it produces a matrix that is exactly singular, from data that was perfectly usable.
The degree that is safe to overshoot
The rules that choose a Tikhonov parameter miss by factors of millions on one draw in twenty. Transplanted to the degree of a polynomial fit, in a basis orthonormal on the data, the same rules never cost more than 2.7 times the best degree's error in three hundred draws. The reason is the shape of the valley they search: six degrees too few costs from 44 to 16,000 times the best error, forty degrees too many costs about twice it. The one rule with a tail, the discrepancy principle, has its threshold half a standard deviation above the residual it is waiting for.
Named alongside it
The objects these essays reach for when they reach for this one.
Bias varianceCholeskyCondition squaringCross-validationDiscrepancy principleError accumulationGeneralised cross-validationGram matrixHilbert matrixLäuchli's matrixLeast-squaresLeverage