Orthogonality, measured

Five precise points are five points

Weighting each sighting by its reliability is the standard form of an attitude or registration fit, and it changes how often the nearest orthogonal matrix comes back as a mirror. Measured, the rate is a function of two numbers: the weighted noise over thickness, and the effective count (Σw)²/Σw². Five points with a tenth of the noise, weighted by 1/σ², carry the information of 515 equal points and mirror like five — 7.9 per cent at a noise ratio where twenty points mirror 1.8 and five mirror 9.5. The √m the earlier measurement left unchecked is right, and it counts what carries the thin direction.

Worth reading first: The nearest orthogonal matrix.

A rotation that comes back mirrored measured how often the nearest orthogonal matrix to a noisy alignment is a reflection. Twenty points on a disc of thickness t, rotated and blurred by noise σ: the rate depends on σ/t and not on the thickness itself, rises from nothing at a ratio of one towards a coin at a hundred, and falls as the point count grows. It offered a heuristic for why — the thin direction’s contribution to the smallest singular value grows like mt2m t^2 and the noise’s like σtm\sigma t \sqrt{m}, so the governing ratio should be σ/(tm)\sigma/(t\sqrt{m}) — and it said plainly that the three-thickness agreement was checked and the m\sqrt{m} was not.

It ended on the form every real application takes. Attitude determination weights each star sighting by its reliability; a molecular superposition weights atoms by mass or by how well they are resolved; a scan registration down-weights points near an edge. The cross-covariance becomes Σ wᵢbᵢaᵢᵀ, the same polar factor is taken, and the question left open was whether weighting a few precise points heavily behaves like having more points — raising the effective thickness — or like having fewer.

It behaves like fewer, and exactly as few as a formula for the weights says.

How often a weighted alignment's polar factor is a reflection, against the effective noise ratio zPer cent of 1,000 trials in which the polar factor of a three-dimensional alignment is a reflection, against z — the square root of the sum of wᵢ²σᵢ², over t times the sum of the weights — which is σ over t times the square root of m for equal weights and noise. Lines are unweighted alignments of 5, 20 and 80 points; the points are two weighted alignments of twenty: five points with a tenth of the noise weighted by 1/σ², and twenty equally noisy points with five of them weighted by a hundred. Both weighted schemes have an effective count (Σw)²/Σw² of 5.3. At z = 0.45 the rates are 1.8% and 1.3% for twenty and eighty points, 9.5% for five, and 7.9% and 6.8% for the two weighted alignments — on the five-point curve, not the twenty-point one.10⁻¹101020304050z, the effective noise over thicknesstrials mirrored, %5 points, equal20 points, equal80 points, equal● five precise, 1/σ²■ five weighted ×100per cent mirrored at z = 0.45, 1,000 trialstwenty points1.8eighty points1.3five points9.5five precise, weighted 1/σ²7.9equal noise, five weighted ×1006.8both weighted schemes have five effective pointstwenty and eighty points share one curve
Fig. 1 The mirror rate against an effective noise ratio, for unweighted alignments of 5, 20 and 80 points as lines and for two weighted alignments of twenty as points. Both weighted alignments sit on the five-point line.

Two numbers a weighting carries

The heuristic generalises directly. With weights wiw_i and per-point noise σi\sigma_i, the thin direction’s signal in the weighted cross-covariance is about t2t^2·Σwᵢ and the noise’s contribution about tiwi2σi2t\sqrt{\sum_i w_i^2 \sigma_i^2}. Their ratio is

z  =  iwi2σi2tiwiz \;=\; \dfrac{\sqrt{\sum_i w_i^2 \sigma_i^2}}{t\,\sum_i w_i}

which is σ/(tm)\sigma/(t\sqrt{m}) when every weight and every noise is equal. That is the first number.

The second is less obvious and it is the one that matters. The signal t2t^2·Σwᵢ is an average of each point’s thin coordinate squared, and each thin coordinate is a random draw. When the weights are equal that average has m terms and fluctuates little; when a few points carry most of the weight it has effectively as many terms as those few, and it fluctuates a great deal — and a small draw of the signal is exactly when the noise can flip the sign. The count of terms that matters is Kish’s effective sample size,

neff  =  (iwi)2iwi2n_{\text{eff}} \;=\; \dfrac{\left(\sum_i w_i\right)^2}{\sum_i w_i^2}

which is m for equal weights and approaches the number of heavy points when a few dominate.

Neither number is new. The first is a signal-to-noise ratio, and the second is the quantity survey statisticians use for how much a weighted sample is worth when the aim is an average. What is being tested is whether the two together are all a mirror depends on — whether an alignment’s handedness is decided by an average of a few random squares in exactly the way a sample mean of a few draws is. One draw in twenty is the reminder here of how badly a small sample’s extremes are described by its mean, and a thin direction’s signal from five points is a small sample. The dimension does not appear is the opposite reminder: for a random projection, how many points there are decides the distortion and the number of coordinates does not. Here too the count is the point count, weighted, and the next measurement asks whether the number of coordinates enters at all.

The claim to test is that a weighted alignment mirrors at the rate an unweighted alignment of neffn_{\mathrm{eff}} points mirrors at the same z.

The square root of the point count, checked first

The generalisation rests on the unweighted heuristic, so that comes first. At each z from 0.1 to 4.5, a thousand unweighted alignments are run at 5, 20 and 80 points, with σ chosen so that σ/(tm)\sigma/(t\sqrt{m}) is that z.

Twenty and eighty points give one curve: 0.1 and 0.0 per cent at z = 0.3, 1.8 and 1.3 at 0.45, 8.7 and 8.3 at 0.7, 17.6 and 18.0 at 1, 33.3 and 31.5 at 2.2. At no z do they differ by more than two and a half points, which is what a thousand trials a cell allow. So the earlier heuristic’s m\sqrt{m} is right, from twenty points up: an alignment of eighty points at noise σ mirrors as often as one of twenty at noise σ/2.

How often the nearest orthogonal matrix to a noisy alignment is a reflection, 80 pointsPercentage of 400 trials in which the polar factor of the cross-covariance of 80 rotated points has determinant −1, against the noise divided by the thickness of the point set, for thicknesses 10⁻², 10⁻³ and 10⁻⁴. The three curves lie on one another: 1.0%, 0.3%, 0.3% at a ratio of 3 and 20%, 18%, 18% at 10, rising towards half.110¹10²01020304050noise ÷ thicknesstrials mirrored, %a coin: 50%t = 10⁻²t = 10⁻³t = 10⁻⁴per cent mirroredσ/t = 3, mean of three0.5σ/t = 10, mean of three18σ/t = 100, mean of three4880 points, 400 trials a stop, one seed per thicknessthe ratio decides, not the thinness
Fig. 2 The earlier measurement’s sweep at eighty points, against σ/t: 18.4 per cent mirrored at a ratio of ten, where twenty points gave 34.2. Against z instead of σ/t, the two sit on one curve.

Five points do not. At z = 0.45 five points mirror on 9.5 per cent of alignments, five times twenty points’ 1.8, and at z = 0.3 on 4.0 per cent against 0.1. The curves meet only at large z, 25.7 against 17.6 at z = 1 and nearly together by 2.2. That is the fluctuation of the signal the effective count describes: five thin coordinates squared and averaged are often small, and a small signal is a flipped one.

How often the nearest orthogonal matrix to a noisy alignment is a reflection, 5 pointsPercentage of 400 trials in which the polar factor of the cross-covariance of 5 rotated points has determinant −1, against the noise divided by the thickness of the point set, for thicknesses 10⁻², 10⁻³ and 10⁻⁴. The three curves lie on one another: 28%, 37%, 36% at a ratio of 3 and 45%, 43%, 42% at 10, rising towards half.110¹10²01020304050noise ÷ thicknesstrials mirrored, %a coin: 50%t = 10⁻²t = 10⁻³t = 10⁻⁴per cent mirroredσ/t = 3, mean of three34σ/t = 10, mean of three43σ/t = 100, mean of three475 points, 400 trials a stop, one seed per thicknessthe ratio decides, not the thinness
Fig. 3 The same sweep at five points, against σ/t: already 28 to 37 per cent mirrored at a ratio of three. Five points is too few for the signal in the thin direction to be close to its average.

Precise points, weighted by their precision

Now the weighted case. Twenty points, five of them measured with a tenth of the noise of the others, weighted by 1/σ² — a hundred for the precise five, one for the rest — which is the weighting a least-squares derivation prescribes. The information those weights carry is Σ(σ/σᵢ)² = 5·100 + 15 = 515 equal points’ worth. The effective count is (500 + 15)²/(50,000 + 15) = 5.3.

At matched z the weighted alignment mirrors on 2.5 per cent at z = 0.3, 7.9 at 0.45, 17.6 at 0.7, 27.0 at 1 and 30.5 at 1.5. The five-point curve at the same z reads 4.0, 9.5, 17.8, 25.7 and 30.6; the twenty-point curve 0.1, 1.8, 8.7, 17.6 and 24.5. The weighted alignment is on the five-point curve, within three points everywhere from z = 0.45 up.

The second weighted alignment makes the point by contrast. Twenty equally noisy points, five of them weighted by a hundred for no reason — a mistake in the weights rather than a correct use of them. Its effective count is also 5.3, and it mirrors on 2.1, 6.8, 16.4, 25.1 and 30.2 per cent at the same five values of z. It too is on the five-point curve. The two alignments differ in whether the heavy points are precise, and at matched z they mirror alike, because at matched z and matched neffn_{\mathrm{eff}} nothing else is left for the rate to depend on.

So the answer to the question the earlier measurement left is specific. Weighting a few precise points heavily lowers z — their low noise enters the numerator squared and weighted — and lowers the effective count to about the number of heavy points. The first helps and the second hurts, and the rate at the resulting z is the rate for that small number of points.

What each weighting does, in the units a user would read

A user does not choose z; a user has a noise, a thickness and a set of weights. At a fixed noise on the ordinary points, the five schemes separate cleanly.

How often a weighted alignment of twenty points is mirrored, against the base noise over thickness, for five weighting schemesPer cent of 400 trials mirrored against σ/t, where σ is the noise on the ordinary points of a twenty-point disc of thickness 0.01. twenty points, equal: 0% at 1, 2.8% at 2, 7.8% at 3, 21% at 5, 32% at 10, 43% at 20, 45% at 30, 53% at 100. five precise, unweighted: 0% at 1, 1.8% at 2, 5.0% at 3, 19% at 5, 31% at 10, 40% at 20, 43% at 30, 52% at 100. five precise, weighted 1/σ²: 0% at 1, 0% at 2, 0% at 3, 1.5% at 5, 9.0% at 10, 18% at 20, 33% at 30, 44% at 100. five precise, weighted ×10: 0% at 1, 0% at 2, 0.3% at 3, 4.8% at 5, 15% at 10, 30% at 20, 39% at 30, 49% at 100. equal noise, five weighted ×100: 6.5% at 1, 22% at 2, 26% at 3, 39% at 5, 43% at 10, 43% at 20, 54% at 30, 49% at 100.110¹10²01020304050noise on the ordinary points ÷ thicknesstrials mirrored, %twenty points, equalfive precise, unweightedfive precise, weighted 1/σ²five precise, weighted ×10equal noise, five weighted ×100per cent mirrored at σ/t = 10, 400 trialstwenty points, equal32five precise, unweighted31five precise, weighted 1/σ²9five precise, weighted ×1015equal noise, five weighted ×10043the same twenty points in every schemeonly the noise and the weights change
Fig. 4 The mirror rate against the noise on the ordinary points over thickness, for twenty equal points and four ways of treating five of them — with or without a tenth of the noise, weighted equally, by ten, by a hundred.

At a noise ratio of ten, twenty equal points mirror on 32 per cent of alignments. Giving five of them a tenth of the noise but ignoring it changes almost nothing: 31 per cent, because the noise in the numerator is still dominated by the fifteen ordinary points and the effective count is still twenty. Weighting those five by ten brings it to 15 per cent, and weighting them by their precision, a hundred, to 9 per cent. Weighting five equally noisy points by a hundred instead raises it to 43 per cent, and at a noise ratio of three — where twenty equal points mirror on 7.8 per cent — to 26.

The correct weighting is a large improvement and it is still worse than the information count suggests. Five hundred and fifteen equal points at this noise would have z = 0.44, where the twenty- and eighty-point curves read under two per cent. The weighted alignment at the same z reads nine. The difference is the effective count, and it is the reason the claim that weighting makes an alignment worth its information has to fail.

Where the effective count stops being the answer

The collapse onto a curve for neffn_{\mathrm{eff}} points is a measurement over a range, and the range has an edge.

Mirror rate at a fixed effective noise ratio, as the weight concentrates on fewer points, against Kish's effective countAt z = 0.7, twenty equally noisy points with p of them weighted by a hundred, for p = 1, 2, 3, 5, 10, 20, placed at the effective count (Σw)²/Σw² and the per cent of 1,000 trials mirrored, beside unweighted alignments of 4, 5, 8, 10, 20, 40, 80 points on a logarithmic count axis. p = 1: effective count 1.4, 0.9%; p = 2: effective count 2.4, 5.5%; p = 3: effective count 3.3, 15%; p = 5: effective count 5.3, 14%; p = 10: effective count 10.2, 13%; p = 20: effective count 20.0, 9.0%. Unweighted: 4 points 22%, 5 points 18%, 8 points 9.7%, 10 points 11%, 20 points 8.7%, 40 points 8.5%, 80 points 8.3%. From five heavy points the weighted rate sits near the unweighted rate at the same count; with one or two it falls far below any unweighted set.110¹10²0510152025points, or the effective count of a weighted settrials mirrored, %unweighted, m pointsweighted, effective countz = 0.7, per cent mirrored, 1,000 trials1 of 20 weighted ×100, effective count 1.40.92 of 20 weighted ×100, effective count 2.45.53 of 20 weighted ×100, effective count 3.3155 of 20 weighted ×100, effective count 5.31410 of 20 weighted ×100, effective count 10.21320 of 20 weighted ×100, effective count 20.09the count axis is logarithmicone heavy point is not one point
Fig. 5 At z = 0.7, twenty equally noisy points with 1, 2, 3, 5, 10 or 20 of them weighted by a hundred, placed at their effective count, beside unweighted alignments of 4 to 80 points.

At z = 0.7, with the weight concentrated on p of twenty equally noisy points, the mirror rate is 14.4 per cent at p = 5 (neffn_{\mathrm{eff}} 5.3), 12.9 at p = 10 (neffn_{\mathrm{eff}} 10.2) and 9.0 at p = 20. Unweighted alignments at the same z read 17.8 at five points, 11.1 at ten and 8.7 at twenty. From five heavy points up, the weighted rate is within a few points of the unweighted rate at the effective count.

Below five it breaks, and in the favourable direction. At p = 3 (neffn_{\mathrm{eff}} 3.3) the rate is 15.2 per cent; at p = 2 (neffn_{\mathrm{eff}} 2.4) it is 5.5; at p = 1 (neffn_{\mathrm{eff}} 1.4) it is 0.9 — lower than any unweighted alignment in the figure, where four points already mirror on 22. One heavy point is not one point. A likely reason, not separated here, is that the light points still contribute to the thin direction’s signal with weight one each, and when the heavy set is too small for its own signal to be reliable their combined signal is what the determinant rests on, which is steadier than the effective count assumes. Kish’s count describes how many points the average is effectively over when the weights are spread; when nearly all the weight is on one point, the rest of the average is small in weight and not in reliability, and the formula has no term for that.

So the rule has a stated domain. For weightings that spread their weight over five or more points, the mirror rate is the unweighted rate at neffn_{\mathrm{eff}} and z. For weightings that put nearly all the weight on one or two points, the rule is pessimistic.

What a practitioner can take from it

The earlier measurement recommended the determinant correction and showed that the corrected rotation’s error does not grow as the disc thins. Nothing here changes that: the correction is the same for a weighted fit, and it is still the repair. What changes is the estimate of how often it fires, which is what an application that logs corrections, or treats a correction as a warning, needs.

Compute z and neffn_{\mathrm{eff}} from the weights. Both are sums over the points, and both are available before any alignment is done. An alignment whose z is below 0.3 at an effective count of twenty or more will almost never need the correction; one whose effective count is five will need it on a few per cent of runs even there.

Weight precise points, and know what it buys. Weighting by precision lowers z by a large factor and lowers the effective count to the number of precise points, and the rate that results is the small set’s rate at the lower z. It is better than not weighting at every noise level measured here, and it is not as good as having that many more points.

Do not weight equally noisy points heavily. It lowers the effective count and does nothing for z, which is the worst of both: 26 per cent mirrored at a noise ratio where the unweighted fit gives 7.8.

Error of the polar factor and of the nearest rotation as the point set thins, noise 0.01Mean distance from the true rotation, divided by the noise, against the thickness of a set of 20 points on log axes. The rotation with its determinant fixed stays between 0.38 and 0.51 at every thickness. The polar factor matches it until the thickness falls below about the noise, then separates: 89.8 at thickness 3·10⁻⁴, where 45% of trials come back mirrored.10⁻⁴10⁻³10⁻²10⁻¹110⁻¹110¹10²10³thickness of the point set‖X − Q★‖ ÷ noisethickness = noisepolar factorrotation, det fixederror per unit of noiserotation, thickness 10.38rotation, thickness 3·10⁻⁴0.48polar, thickness 3·10⁻⁴90% mirrored there4520 points, 400 trials a stop; none mirrored down to 0.01thinner costs the rotation nothing
Fig. 6 The earlier measurement’s reason the correction is worth applying: once the determinant is fixed, the rotation’s error stays near half the noise however thin the point set becomes, while the uncorrected polar factor’s error leaves as soon as mirrors begin.

The nearest orthogonal matrix is the essay that established the polar factor as the right answer before any handedness question arose, and the corrected rotation’s error, which that measurement and the one before this found independent of the thinness, is the reason a mirror rate is a rate of corrections rather than a rate of failures. What the weights change is how often the correction is doing work, not whether it works.

A count decided before the data arrives

The effective count has a property worth drawing out, because it is the one that makes it useful. It is a function of the weights alone. No point coordinate, no noise draw and no alignment enters it, so an application knows its neffn_{\mathrm{eff}} — and therefore which curve its mirror rate will sit on — before it has taken a single measurement.

Influence is decided before the data found the same shape in least squares: the leverage of an observation, which decides how much one bad measurement can move a fit, depends on where the observation was taken and not on what it recorded. The effective count is the alignment’s version of it. Where leverage says which points the fit listens to, neffn_{\mathrm{eff}} says how many points the thin direction’s handedness is decided by, and both are settled by the design of the experiment rather than by its outcome.

That makes the weights a design decision with a measurable side effect. A star tracker that assigns five bright stars a hundred times the weight of fifteen faint ones has chosen an effective count of five for its handedness, whatever the sky looks like on the day. A constraint is a weight at infinity followed a weight to its limit and found the limit to be a different problem; the limit here is an neffn_{\mathrm{eff}} of one, and the measurement above says it is not the problem the formula predicts either.

And it says something about the determinant the correction reads. The number that decides nothing argued that a determinant is almost never the right quantity to base a decision on, and the earlier measurement noted that the sign of det(PQᵀ) is the exception, decided by a margin of one. What decides how often that sign is the wrong one is not in the determinant at all: it is z, which is set by the noise and the weights, and neffn_{\mathrm{eff}}, which is set by the weights. A code that wants to know in advance how often it will be correcting a mirror can compute both and read the rate off one of the curves drawn above.

What this rests on

Three dimensions, a disc of thickness 0.01, a thousand alignments per point on the z curves and four hundred on the noise-ratio curves, one set of five precise or heavy points out of twenty. The noise on precise points is a tenth of the ordinary noise; other ratios are not measured. The weights are applied only in the cross-covariance, which is the Wahba form; a fit that also weighted the centring of the point sets would add a second effect that is not here.

The derivation of z and neffn_{\mathrm{eff}} is a heuristic of the same standing as the earlier one: it predicts which curve a weighted alignment should sit on, and the measurement says that it does over the range drawn and that it stops doing so when one or two points carry the weight. It is not a bound.

The claim that has to fail

The claim is the natural reading of optimal weighting: points weighted by their precision behave like as many equal points as their information is worth. Five points with a tenth of the noise, weighted by 1/σ², are worth 515 equal points. At z = 0.45 they mirror on 7.9 per cent of alignments; 515 equal points at the same z mirror on under two. The refusal is fed the claim that the two are within three points and fails; the unweighted five-point curve beside it, at 9.5, is where the weighted alignment actually is.

Still open: more dimensions, and a weighting chosen for the handedness

Dimensions above three. The determinant correction flips the direction of the smallest singular value in any number of dimensions, and the rate at which it is needed should depend on the thin directions alone. How the rate changes with the dimension, and with the number of thin directions rather than just the thinnest, is measured in a mirror decided in the thin directions.

A weighting chosen for the handedness. The optimal weights for the rotation’s accuracy are 1/σ², and they are not the weights that minimise the mirror rate, since those would keep the effective count high. A weighting that trades a little accuracy in the rotation for a larger effective count — shrinking the heavy weights towards the light ones — would move the rate at a known cost in the rotation’s error, and where that trade is worth making is unmeasured.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DeterminantOrthogonalityPolar decompositionReflectionRotationSingular value decompositionWeighted least squares