Five precise points are five points
Worth reading first: The nearest orthogonal matrix.
A rotation that comes back mirrored measured how often the nearest orthogonal matrix to a noisy alignment is a reflection. Twenty points on a disc of thickness t, rotated and blurred by noise σ: the rate depends on σ/t and not on the thickness itself, rises from nothing at a ratio of one towards a coin at a hundred, and falls as the point count grows. It offered a heuristic for why — the thin direction’s contribution to the smallest singular value grows like and the noise’s like , so the governing ratio should be — and it said plainly that the three-thickness agreement was checked and the was not.
It ended on the form every real application takes. Attitude determination weights each star sighting by its reliability; a molecular superposition weights atoms by mass or by how well they are resolved; a scan registration down-weights points near an edge. The cross-covariance becomes Σ wᵢbᵢaᵢᵀ, the same polar factor is taken, and the question left open was whether weighting a few precise points heavily behaves like having more points — raising the effective thickness — or like having fewer.
It behaves like fewer, and exactly as few as a formula for the weights says.
Two numbers a weighting carries
The heuristic generalises directly. With weights and per-point noise , the thin direction’s signal in the weighted cross-covariance is about ·Σwᵢ and the noise’s contribution about . Their ratio is
which is when every weight and every noise is equal. That is the first number.
The second is less obvious and it is the one that matters. The signal ·Σwᵢ is an average of each point’s thin coordinate squared, and each thin coordinate is a random draw. When the weights are equal that average has m terms and fluctuates little; when a few points carry most of the weight it has effectively as many terms as those few, and it fluctuates a great deal — and a small draw of the signal is exactly when the noise can flip the sign. The count of terms that matters is Kish’s effective sample size,
which is m for equal weights and approaches the number of heavy points when a few dominate.
Neither number is new. The first is a signal-to-noise ratio, and the second is the quantity survey statisticians use for how much a weighted sample is worth when the aim is an average. What is being tested is whether the two together are all a mirror depends on — whether an alignment’s handedness is decided by an average of a few random squares in exactly the way a sample mean of a few draws is. One draw in twenty is the reminder here of how badly a small sample’s extremes are described by its mean, and a thin direction’s signal from five points is a small sample. The dimension does not appear is the opposite reminder: for a random projection, how many points there are decides the distortion and the number of coordinates does not. Here too the count is the point count, weighted, and the next measurement asks whether the number of coordinates enters at all.
The claim to test is that a weighted alignment mirrors at the rate an unweighted alignment of points mirrors at the same z.
The square root of the point count, checked first
The generalisation rests on the unweighted heuristic, so that comes first. At each z from 0.1 to 4.5, a thousand unweighted alignments are run at 5, 20 and 80 points, with σ chosen so that is that z.
Twenty and eighty points give one curve: 0.1 and 0.0 per cent at z = 0.3, 1.8 and 1.3 at 0.45, 8.7 and 8.3 at 0.7, 17.6 and 18.0 at 1, 33.3 and 31.5 at 2.2. At no z do they differ by more than two and a half points, which is what a thousand trials a cell allow. So the earlier heuristic’s is right, from twenty points up: an alignment of eighty points at noise σ mirrors as often as one of twenty at noise σ/2.
Five points do not. At z = 0.45 five points mirror on 9.5 per cent of alignments, five times twenty points’ 1.8, and at z = 0.3 on 4.0 per cent against 0.1. The curves meet only at large z, 25.7 against 17.6 at z = 1 and nearly together by 2.2. That is the fluctuation of the signal the effective count describes: five thin coordinates squared and averaged are often small, and a small signal is a flipped one.
Precise points, weighted by their precision
Now the weighted case. Twenty points, five of them measured with a tenth of the noise of the others, weighted by 1/σ² — a hundred for the precise five, one for the rest — which is the weighting a least-squares derivation prescribes. The information those weights carry is Σ(σ/σᵢ)² = 5·100 + 15 = 515 equal points’ worth. The effective count is (500 + 15)²/(50,000 + 15) = 5.3.
At matched z the weighted alignment mirrors on 2.5 per cent at z = 0.3, 7.9 at 0.45, 17.6 at 0.7, 27.0 at 1 and 30.5 at 1.5. The five-point curve at the same z reads 4.0, 9.5, 17.8, 25.7 and 30.6; the twenty-point curve 0.1, 1.8, 8.7, 17.6 and 24.5. The weighted alignment is on the five-point curve, within three points everywhere from z = 0.45 up.
The second weighted alignment makes the point by contrast. Twenty equally noisy points, five of them weighted by a hundred for no reason — a mistake in the weights rather than a correct use of them. Its effective count is also 5.3, and it mirrors on 2.1, 6.8, 16.4, 25.1 and 30.2 per cent at the same five values of z. It too is on the five-point curve. The two alignments differ in whether the heavy points are precise, and at matched z they mirror alike, because at matched z and matched nothing else is left for the rate to depend on.
So the answer to the question the earlier measurement left is specific. Weighting a few precise points heavily lowers z — their low noise enters the numerator squared and weighted — and lowers the effective count to about the number of heavy points. The first helps and the second hurts, and the rate at the resulting z is the rate for that small number of points.
What each weighting does, in the units a user would read
A user does not choose z; a user has a noise, a thickness and a set of weights. At a fixed noise on the ordinary points, the five schemes separate cleanly.
At a noise ratio of ten, twenty equal points mirror on 32 per cent of alignments. Giving five of them a tenth of the noise but ignoring it changes almost nothing: 31 per cent, because the noise in the numerator is still dominated by the fifteen ordinary points and the effective count is still twenty. Weighting those five by ten brings it to 15 per cent, and weighting them by their precision, a hundred, to 9 per cent. Weighting five equally noisy points by a hundred instead raises it to 43 per cent, and at a noise ratio of three — where twenty equal points mirror on 7.8 per cent — to 26.
The correct weighting is a large improvement and it is still worse than the information count suggests. Five hundred and fifteen equal points at this noise would have z = 0.44, where the twenty- and eighty-point curves read under two per cent. The weighted alignment at the same z reads nine. The difference is the effective count, and it is the reason the claim that weighting makes an alignment worth its information has to fail.
Where the effective count stops being the answer
The collapse onto a curve for points is a measurement over a range, and the range has an edge.
At z = 0.7, with the weight concentrated on p of twenty equally noisy points, the mirror rate is 14.4 per cent at p = 5 ( 5.3), 12.9 at p = 10 ( 10.2) and 9.0 at p = 20. Unweighted alignments at the same z read 17.8 at five points, 11.1 at ten and 8.7 at twenty. From five heavy points up, the weighted rate is within a few points of the unweighted rate at the effective count.
Below five it breaks, and in the favourable direction. At p = 3 ( 3.3) the rate is 15.2 per cent; at p = 2 ( 2.4) it is 5.5; at p = 1 ( 1.4) it is 0.9 — lower than any unweighted alignment in the figure, where four points already mirror on 22. One heavy point is not one point. A likely reason, not separated here, is that the light points still contribute to the thin direction’s signal with weight one each, and when the heavy set is too small for its own signal to be reliable their combined signal is what the determinant rests on, which is steadier than the effective count assumes. Kish’s count describes how many points the average is effectively over when the weights are spread; when nearly all the weight is on one point, the rest of the average is small in weight and not in reliability, and the formula has no term for that.
So the rule has a stated domain. For weightings that spread their weight over five or more points, the mirror rate is the unweighted rate at and z. For weightings that put nearly all the weight on one or two points, the rule is pessimistic.
What a practitioner can take from it
The earlier measurement recommended the determinant correction and showed that the corrected rotation’s error does not grow as the disc thins. Nothing here changes that: the correction is the same for a weighted fit, and it is still the repair. What changes is the estimate of how often it fires, which is what an application that logs corrections, or treats a correction as a warning, needs.
Compute z and from the weights. Both are sums over the points, and both are available before any alignment is done. An alignment whose z is below 0.3 at an effective count of twenty or more will almost never need the correction; one whose effective count is five will need it on a few per cent of runs even there.
Weight precise points, and know what it buys. Weighting by precision lowers z by a large factor and lowers the effective count to the number of precise points, and the rate that results is the small set’s rate at the lower z. It is better than not weighting at every noise level measured here, and it is not as good as having that many more points.
Do not weight equally noisy points heavily. It lowers the effective count and does nothing for z, which is the worst of both: 26 per cent mirrored at a noise ratio where the unweighted fit gives 7.8.
The nearest orthogonal matrix is the essay that established the polar factor as the right answer before any handedness question arose, and the corrected rotation’s error, which that measurement and the one before this found independent of the thinness, is the reason a mirror rate is a rate of corrections rather than a rate of failures. What the weights change is how often the correction is doing work, not whether it works.
A count decided before the data arrives
The effective count has a property worth drawing out, because it is the one that makes it useful. It is a function of the weights alone. No point coordinate, no noise draw and no alignment enters it, so an application knows its — and therefore which curve its mirror rate will sit on — before it has taken a single measurement.
Influence is decided before the data found the same shape in least squares: the leverage of an observation, which decides how much one bad measurement can move a fit, depends on where the observation was taken and not on what it recorded. The effective count is the alignment’s version of it. Where leverage says which points the fit listens to, says how many points the thin direction’s handedness is decided by, and both are settled by the design of the experiment rather than by its outcome.
That makes the weights a design decision with a measurable side effect. A star tracker that assigns five bright stars a hundred times the weight of fifteen faint ones has chosen an effective count of five for its handedness, whatever the sky looks like on the day. A constraint is a weight at infinity followed a weight to its limit and found the limit to be a different problem; the limit here is an of one, and the measurement above says it is not the problem the formula predicts either.
And it says something about the determinant the correction reads. The number that decides nothing argued that a determinant is almost never the right quantity to base a decision on, and the earlier measurement noted that the sign of det(PQᵀ) is the exception, decided by a margin of one. What decides how often that sign is the wrong one is not in the determinant at all: it is z, which is set by the noise and the weights, and , which is set by the weights. A code that wants to know in advance how often it will be correcting a mirror can compute both and read the rate off one of the curves drawn above.
What this rests on
Three dimensions, a disc of thickness 0.01, a thousand alignments per point on the z curves and four hundred on the noise-ratio curves, one set of five precise or heavy points out of twenty. The noise on precise points is a tenth of the ordinary noise; other ratios are not measured. The weights are applied only in the cross-covariance, which is the Wahba form; a fit that also weighted the centring of the point sets would add a second effect that is not here.
The derivation of z and is a heuristic of the same standing as the earlier one: it predicts which curve a weighted alignment should sit on, and the measurement says that it does over the range drawn and that it stops doing so when one or two points carry the weight. It is not a bound.
The claim that has to fail
The claim is the natural reading of optimal weighting: points weighted by their precision behave like as many equal points as their information is worth. Five points with a tenth of the noise, weighted by 1/σ², are worth 515 equal points. At z = 0.45 they mirror on 7.9 per cent of alignments; 515 equal points at the same z mirror on under two. The refusal is fed the claim that the two are within three points and fails; the unweighted five-point curve beside it, at 9.5, is where the weighted alignment actually is.
Still open: more dimensions, and a weighting chosen for the handedness
Dimensions above three. The determinant correction flips the direction of the smallest singular value in any number of dimensions, and the rate at which it is needed should depend on the thin directions alone. How the rate changes with the dimension, and with the number of thin directions rather than just the thinnest, is measured in a mirror decided in the thin directions.
A weighting chosen for the handedness. The optimal weights for the rotation’s accuracy are 1/σ², and they are not the weights that minimise the mirror rate, since those would keep the effective count high. A weighting that trades a little accuracy in the rotation for a larger effective count — shrinking the heavy weights towards the light ones — would move the rate at a known cost in the rotation’s error, and where that trade is worth making is unmeasured.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A test with no answer in it — both name orthogonality, polar decomposition, singular value decomposition
- An iteration that only multiplies — both name orthogonality, polar decomposition
- The best approximation there is — both name orthogonality, singular value decomposition
Named objects
A flat tag is an object no other essay names yet.
DeterminantOrthogonalityPolar decompositionReflectionRotationSingular value decompositionWeighted least squares