A unit is a statement about the noise
Worth reading first: When the matrix is wrong too · The projection and the right angle · The units the matrix is measured in.
When the matrix is wrong too introduced total least squares as the answer to a fit whose matrix is measured as well as its right-hand side, and noted in passing that it has a price ordinary least squares does not. Its objective is , one Frobenius norm over the whole augmented matrix, “which treats a unit of error in the first column as interchangeable with a unit in the last. If the first column is a temperature in kelvin and the last a pressure in pascals, that is a claim nobody made.” The repair it named is a scaling: divide each column by its own noise level. The two numbers a caller has then found that the quantity governing the method’s error is a gap between two singular values, and that the residual and a caller can see are blind to it.
Both essays drew every column’s noise at the same size, so the claim the units make was always true. Neither measured what happens when it is false, and the two questions that follow are concrete. How far does a wrong unit move the answer, and is the damage bounded? And does the gap that predicted the error survive a change of units? This essay measures both, on a fit where the true coefficients are known.
Least squares does not see the units and total least squares does
The fit has 60 rows and two measured regressors. The exact design has a first column scattered about 1 and a second about 2, the true coefficients are 1.5 and −0.7, and Gaussian noise is added to both columns and to the right-hand side, with a standard deviation set separately for each in the units the data were recorded in. The experiment rewrites the first column in units times smaller: every entry is multiplied by , so its true coefficient becomes . Each estimate is then converted back by multiplying its first coefficient by , and its error is read in the recorded units — the root mean square of the two coefficients’ relative errors, median over forty fits.
The figure at the top of the page is the case both earlier essays assumed: noise with standard deviation 0.6 in every column. Ordinary least squares returns 0.3171 at every unit factor, and the nine-digit agreement across eight decades is not a coincidence. Multiplying a column of by multiplies the corresponding coefficient of the least-squares solution by exactly and changes nothing else, because the residual it minimises is a vector in the units of and the column space of is the same set at every . The projection and the right angle is the geometry of that: the answer is the projection of onto a space, and rescaling a basis vector of the space does not move the space.
Total least squares returns 0.1139 in the recorded units, and that number is not special to the method; it is special to the units, because in them the three noise levels happen to be equal. Rewrite the first column a hundred times larger and the error is 0.1716; a hundred times smaller, 0.1607. The curve is flat at both ends and dips only where the units are right, and the higher plateau stands half as high again as the dip.
Between two estimators somebody has named
The flat ends are limits, and both are estimators in their own right.
As the first column shrinks toward nothing in Frobenius terms, and the cheapest way to make the system consistent is to change that column. The correction goes there entirely, so the method behaves as if the first column were the only noisy quantity and everything else exact. That is the reverse regression: regress the first column on the second column and on by ordinary least squares, then solve the fitted relation for . As the first column is too large to correct at any affordable price, so it is treated as exact. That is the mixed fit: project the first column out, solve total least squares on what is left, and recover the first coefficient by back-substitution, which is the limit a constraint is a weight at infinity takes for a constraint row, taken here for a column.
The measurement computes both directly and compares them with total least squares at and , fit by fit, on all 120 fits of the three noise placements. Every coefficient agrees with its limit to better than one part in a thousand. With equal noise the reverse regression has median error 0.1607 and the mixed fit 0.1716, and the curve in the hero figure runs between them with its minimum at . So the damage a unit can do here is bounded, by whichever of the two named estimators is worse, and with equal noise both are still better than ordinary least squares by almost a factor of two. A unit chosen badly picks a point on a one-parameter family whose ends are known. It does not pick an arbitrary answer.
The correction is split by the coefficients
What the units actually move is where total least squares puts its correction, and that has a closed form worth having. The correction is rank one: with residual it is and . The column of belonging to coefficient is a multiple of by , and the correction to is itself, so the squared corrections stand in the ratio — in whatever units the problem is written in.
In the recorded units, with coefficients near 1.5 and −0.7, that ratio is 2.25 : 0.49 : 1, and the measured split is 59.8, 13.2 and 27.0 per cent. The noise is a third in each column throughout. So even in the units where the method is at its best the correction is not an estimate of where the noise was; it is the smallest perturbation that makes the system consistent, and smallest is measured with coefficients as weights. Turn the dial to 0.1 and the first coefficient in the new units is fifteen, its share of the correction rounds to 100 per cent, and the fit is the reverse regression. Turn it to 100 and the coefficient is 0.015, the share rounds to zero, and the other two settle at 28.2 and 71.8 per cent, the second coefficient’s square against one, with the first column treated as exact.
That is the precise sense in which a unit is a statement about the noise. Writing the first column in units a hundred times larger says its readings are a hundred times more precise relative to their size than they were, and the method believes it. The noise did not change; the claim did.
Where the matrix is nearly exact no unit rescues it
This sweep is the case when the matrix is wrong too used to show total least squares losing: noise of 0.05 in each column of and 0.6 in . Ordinary least squares has median error 0.0429. Plain total least squares in the recorded units has 0.1108, 2.6 times worse, and no unit factor brings it close: its best is 0.0757 as , where the first column is treated as exact and the second is still corrected as though it were as noisy as , and its worst 0.1475 as .
The earlier essay drew the conclusion that the model matters more than the method, and this sharpens it. Weight each column by its own noise level — divide the columns of by 0.05 and by 0.6 before solving, and undo the scaling afterwards — and total least squares has median error 0.0424, within 1.2 per cent of ordinary least squares. The loss on this problem was never total least squares failing on an exact matrix. It was the plain method’s implicit claim that the matrix is as noisy as the right-hand side, and with the claim corrected the two estimators agree to the precision of a median over forty fits, as they should: when is nearly exact, weighted total least squares is nearly ordinary least squares.
When one column carries the noise the reverse regression is right
The third placement puts the noise mostly in the first column: 0.6 there, 0.15 in the second, 0.3 in .
Here the curve is low on the left and high on the right. The reverse regression, the limit, has median error 0.0914 and the weighted fit 0.0912. That is the limit’s own logic: the reverse regression treats the first column as the only noisy quantity, and on this problem it nearly is, carrying 76 per cent of the noise variance. The mixed fit, which treats that column as exact, has 0.1779, nearly twice as much. In the recorded units plain total least squares is at 0.0921, close to the best by luck, and a factor of three in the first column’s units, , takes it to 0.1494. Ordinary least squares is at 0.2640.
So a named estimator that is bad on one problem is near-optimal on another, and which one is good depends on where the noise is. That is the same statement the earlier essays made about ordinary and total least squares, made now about the whole family the units sweep through.
Five estimators, three placements, one that is never beaten
Put side by side, the five estimators each win somewhere except one that always wins. Ordinary least squares is the best of the unweighted ones when is nearly exact and the worst by far otherwise. Plain total least squares in the recorded units is best when the recorded units happen to equalise the noise and middling elsewhere. The reverse regression is right when the first column carries the noise and worst when is exact. The mixed fit is the best unweighted total least-squares answer when is nearly exact and the worst of them when the first column is noisy. The weighted fit is at or within half a per cent of the best on every placement: 0.1139, 0.0912 and 0.0424. The measurement checks that no unit factor anywhere in the sweep beats it by more than half a per cent, and none beats it at all; the closest any comes is the reverse regression on the second placement, at 0.0914 against 0.0912.
That is the result in its plainest form. Among all the answers total least squares can give by changing units, the one that divides each column by its noise is the best on every problem measured, and it is the only one with that property. The weighting is not a refinement of the method. It is what makes the norm the method minimises a sum of like quantities, each column’s correction measured in units of its own noise.
The gap is in units too
The two numbers a caller has found that total least squares’ error is governed by the gap between the smallest singular value of and the smallest of : the error times the gap over the noise level held within 16 per cent across a sweep on which the error moved by a factor of 201. The gap is computable from the data, which made it the useful finding. But singular values are norms of the matrix, and the matrix is written in units.
Across the units sweep the gap runs from 0.000312 at to 8.02 at , a factor of , while the error moves by 1.51. As the first column’s smallest singular value shrinks with it, so the gap falls in proportion to , and a caller reading it would conclude the problem had become four decades less well conditioned. It had not: the answer converted back is the reverse regression, with an error 1.41 times the error at .
The earlier predictor did not fail; it was stated with the noise level in the denominator, and with the units changed the noise in the first column is times larger in its new units than in its old, so the noise level is no longer one number. The predictor needs one, and it has one only when every column’s noise is the same size — which is the weighted problem. In weighted units the noise is 1 in every column, the gap is a number about the experiment, and the error over the gap means what it meant. So the gap is not a diagnostic a caller can read off a plain fit in whatever units arrived. It is a diagnostic of the weighted fit, and computing it requires the same knowledge the weighting does. Two condition numbers of one matrix drew the same line for linear systems: a condition number describes a class of perturbations, and changing the units changes the class.
Two readings a column
All of this assumes the noise levels are known, and the two numbers a caller has ended on why they usually are not: the noise is an error that was never observed, and estimating it from the fit’s residual is circular. The honest way to know a column’s noise is to measure it, by reading the same quantity more than once. So the last measurement replaces the known levels by standard deviations estimated from replicate readings of each column, and repeats the fit on two hundred problems at each to steady the medians.
With the noise almost all in , two readings a column bring the error to 0.0528 against 0.1227 for the plain fit and 0.0447 with the levels known; by three readings it is 0.0462. A crude estimate is enough there, because the levels differ by a factor of twelve and even two readings put them in the right order of magnitude. With the noise largest in the first column, two readings give 0.1032, worse than the plain fit’s 0.0843; three match it, and ten, at 0.0748, approach the known-weight 0.0734. With equal noise, where the recorded units were already right, estimated weights only add error: 0.1713 with two readings and 0.0984 with thirty, against 0.0921 for both the plain fit and the known weights. The estimate cannot improve on weights that were correct by accident, and with two readings it misjudges each level by a factor drawn from a distribution with one degree of freedom.
So the price of the weighting is a handful of repeated readings, and it buys most where the recorded units were furthest from the noise. A caller who has no replicates and no instrument specification is choosing a unit for each column and, with it, a claim about the noise; choosing without knowing priced a guessed noise level in a neighbouring problem, and the price there was a factor of ninety thousand in the error.
What three placements do not show
Two regressors, sixty rows, one design and three placements of Gaussian noise that is independent between columns. With correlated noise the right weighting is a whole covariance and not a scaling per column, and a column-by-column weight would be wrong in a way none of these measurements can see. There is no intercept: a column of ones is exact by construction, and the right fit keeps it exact, which is the mixed fit applied to that column rather than a limit the units reach by accident. The errors are medians over forty fits, or two hundred for the replicates, and a median moves by a few per cent between seed sets: the known-weight error with equal noise is 0.1139 on the forty and 0.0921 on the two hundred, and every comparison above is made within one set. And the error is measured against true coefficients that a real fit does not have, which is the habit these essays keep because the residual is the one quantity that cannot tell these estimators apart.
Still open: weights from the data alone, and a column of ones
Weights from the data alone. A tempting shortcut is to iterate: fit, estimate each column’s noise from the correction the fit applied to it, reweight, and fit again. For Gaussian noise the ratio of the noise levels is known not to be identifiable from the data alone. The prediction with a sign is that on these fits the iteration converges, and to weights that depend on the units it starts from: started from and from on the same data, the converged weight on the first column differs by more than a factor of ten, and neither start reaches the error of the replicate weights with three readings.
A column of ones. Add an intercept to the model, exact by construction. The prediction is that plain total least squares, which corrects the column of ones as if it were noisy, has a median error at least 1.3 times that of the fit that keeps it exact and weights the rest by their noise, at every one of the three placements; and that the gap from the two numbers a caller has, computed with the column of ones treated as data, predicts the plain fit’s error worse than it predicts the mixed fit’s.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A backward-stable answer to a problem nobody asked — both name exact ground truth, scaling
- A correction that reads every row — both name exact ground truth, least-squares
- An estimate that does not move — both name exact ground truth, scaling
- Five precise points are five points — both name singular value decomposition, weighted least-squares
- The half of a problem a sketch may touch — both name exact ground truth, least-squares
- The residual the appended block cannot remove — both name exact ground truth, least-squares
Named objects
A flat tag is an object no other essay names yet.
Errors-in-variablesExact ground truthLeast-squaresNoise levelScalingSingular value decompositionTotal least-squaresWeighted least-squares