Orthogonality, measured

A mirror decided in the thin directions

In n dimensions the nearest orthogonal matrix to a noisy alignment is still sometimes a reflection, and the rate at which it is does not depend on n. Three, five and ten dimensions with one thin direction mirror alike; two thin directions mirror like each other in five dimensions and in ten. The rate is the chance that a k × k matrix built from the k thin directions has a negative determinant — 21.7 per cent at a noise ratio of 0.7 for k = 2, measured at 23.0 — and it is well above k independent coin flips. With two or more thin directions the determinant correction still fires, and it no longer rescues the rotation: the answer is eleven noise-widths off whether or not it was mirrored.

Worth reading first: The nearest orthogonal matrix.

A rotation that comes back mirrored measured the reflection problem in three dimensions, where the application that motivates it — aligning a molecule, a spacecraft, a scan — lives. It closed on the dimensions where there is no picture to look at. The correction generalises unchanged: find the singular value decomposition of the cross-covariance, take the polar factor, and if its determinant is −1 flip the direction of the smallest singular value. The earlier measurement’s reasoning predicted that the probability of needing the flip should depend on the smallest singular value alone, and noted that in higher dimensions there are more directions a noisy estimate can be thin in. Registration of point clouds in feature space, alignment of two embeddings, Procrustes analysis of shapes with many landmarks — all are n-dimensional, and all can have a mirror nobody sees.

Five precise points are five points settled the three-dimensional rate’s dependence on the point count and on weights. This essay sets the dimension and the number of thin directions free.

How often the polar factor of an n-dimensional alignment with k thin directions is a reflection, against the k × k determinant modelPer cent of 600 trials mirrored, against z — the noise over the thickness, divided by the square root of the point count — for alignments in n = 3, 5, 8 and 10 dimensions with k = 1, 2 or 3 directions of thickness 0.01, drawn as points, beside the rate at which the k × k matrix W/m + zG has a negative determinant, W Wishart on m points and G Gaussian, drawn as lines from 20,000 draws each, on a linear axis. At z = 0.7: n = 3, k = 1: 9.7%; n = 5, k = 1: 13%; n = 10, k = 1: 12%; n = 5, k = 2: 23%; n = 10, k = 2: 24%; n = 8, k = 3: 32%; the models give 9.5% for k = 1, 22% for k = 2, 31% for k = 3.0123401020304050z, the effective noise ratiotrials mirrored, %1 thin: 1 × 1 model2 thin: 2 × 2 model3 thin: 3 × 3 modelper cent mirrored at z = 0.7, 600 trialsn = 3, 1 thin, 20 points9.7n = 5, 1 thin, 20 points13n = 10, 1 thin, 40 points12n = 5, 2 thin, 20 points23n = 10, 2 thin, 40 points24n = 8, 3 thin, 40 points32points: measured in n dimensionslines: a k × k determinant, no n in it
Fig. 1 Mirror rates measured in three to ten dimensions with one, two or three thin directions, as points, against the rate at which a k × k determinant model built from the thin directions alone is negative, as lines. The points group by k, not by n.

The experiment in n dimensions

m points in n dimensions, with n − k coordinates of unit spread and k coordinates of spread t = 0.01: a flat slab with k thin directions. A random rotation with determinant +1 is applied, independent noise σ is added to every coordinate, and the n × n cross-covariance is formed and decomposed. Its polar factor is the nearest orthogonal matrix to it, as in three dimensions, and it is a reflection when its determinant, computed by LU with pivoting, is negative; the corrected rotation flips the smallest singular direction.

The noise is set through z=σ/(tm)z = \sigma/(t\sqrt{m}), the ratio the previous two measurements established for three dimensions, so that a curve drawn against z can be compared across point counts. Six arrangements are measured with six hundred alignments at each of seven values of z: one thin direction in three, five and ten dimensions; two in five and in ten; three in eight.

The dimension drops out

With one thin direction, three dimensions mirror on 1.7 per cent of alignments at z = 0.45, 9.7 at 0.7, 18.0 at 1 and 26.3 at 1.5. Five dimensions give 2.7, 13.2, 21.3 and 27.7; ten dimensions with forty points give 3.2, 11.5, 17.8 and 26.8. The differences between the three are at most three and a half points, which is what six hundred trials a cell allow, and they are not ordered by the dimension.

With two thin directions, five dimensions give 9.5, 23.0, 34.2 and 40.0 at the same four values of z, and ten dimensions give 10.3, 23.8, 36.2 and 43.3. Again the dimension moves nothing.

The number of thin directions moves everything. At z = 0.7 one thin direction mirrors on 10 to 13 per cent of alignments, two on 23 to 24, and three on 32. At z = 1.5, one reaches about 27, two about 40 to 43, and three 49.5 — a coin, within sampling. More thin directions reach the coin sooner.

The handedness lives in a k × k block

The reason the dimension drops out is visible in the structure of the cross-covariance. In the basis of the point set’s own spread, the n − k thick directions carry signal of order m and noise of order σm\sigma\sqrt{m}, so their contribution to the polar factor is decided with a large margin and has determinant +1 in every trial that matters. The k thin directions carry signal of order t2mt^2 m and noise of order tσmt\sigma\sqrt{m}, and all the uncertainty about the determinant’s sign is in that k × k block. Divided by t²m, the block is

Wm  +  zG\dfrac{W}{m} \;+\; z\,G

where W is the k × k matrix of sums of the thin coordinates’ products — a Wishart matrix on m points — and G is a k × k matrix of standard Gaussians. There is no n in it. The model predicts that the mirror rate is the probability that this k × k matrix has a negative determinant, and that probability can be computed by drawing W and G directly, twenty thousand times at each z.

For one thin direction on twenty points the model gives 3.0, 9.5, 16.8 and 26.0 per cent at z = 0.45, 0.7, 1 and 1.5, against the three measured curves above. For two thin directions on twenty points it gives 9.9, 21.7, 32.7 and 40.7, against five dimensions’ 9.5, 23.0, 34.2 and 40.0. For three on forty points, 14.2, 30.8, 41.0 and 48.2, against eight dimensions’ 13.5, 31.8, 39.8 and 49.5. Every measured point in the range is within five points of its model, and most within two.

So the mirror in n dimensions is a statement about a k × k random matrix, and the three-dimensional case the earlier essays measured is the scalar case: one thin direction, and a determinant that is simply the sign of W/m + zG with both terms numbers.

Several thin directions are not several coins

The obvious simplification is that each thin direction flips independently, with the one-direction probability p, and the determinant is negative when an odd number of them flip. That gives (1 − (1 − 2p)ᵏ)/2, which is easy to compute and wrong.

Mirror rate with two and three thin directions: the determinant model against independent flipsFor forty points, the per cent of 20,000 draws in which W/m + zG has a negative determinant, for k × k blocks with k = 1, 2 and 3, against z, and the rate k independent flips at the one-direction rate p would give, (1 − (1 − 2p)ᵏ)/2, dashed, on linear axes. At z = 0.45 one direction mirrors 2.2%; two independent flips would give 4.2% and the determinant gives 7.0%; three would give 6.2% and the determinant gives 14%.00.511.5201020304050z, the effective noise ratiomirrored, %two, determinanttwo, independentthree, determinantthree, independentz = 0.45, forty points, per centone thin direction2.2two: independent flips4.2two: determinant7three: independent flips6.2three: determinant14dashed: k coinssolid: one determinant
Fig. 2 On forty points, the k × k determinant model for one, two and three thin directions, solid, against k independent flips at the one-direction rate, dashed. The determinant is negative more often than the coins would make it at every z below the coin.

On forty points at z = 0.45, one thin direction mirrors on 2.2 per cent of draws. Two independent flips at that rate would give 4.2 per cent; the two-direction determinant gives 7.0. Three independent flips would give 6.2; the three-direction determinant gives 14.0. At z = 0.7 the two curves are closer but still apart, and they meet only near the coin, where both are fifty.

The determinant is not a product of independent signs because the k × k block has off-diagonal entries. Its determinant can be driven negative by the diagonal signs, as the coins model says, and also by large off-diagonal noise when the diagonal signals are close to each other — a second route to a mirror that one thin direction does not have. The more thin directions, the more off-diagonal entries, and the further the determinant runs ahead of the coins.

That makes independence the wrong way round for a practitioner estimating how often a high-dimensional alignment will need correcting: it underestimates the rate by a factor of about two at small noise ratios, which is where a practitioner would hope the correction is rare.

What a mirror looks like in any dimension

One property of the three-dimensional mirror carries over exactly, and it is worth stating because it is what a code would see. The polar factor U and the corrected rotation R differ only in the sign of one singular direction, so U − R is the outer product of that direction’s left and right singular vectors, times two. Its Frobenius norm is 2 in every dimension, whatever n and k are.

Distance between the polar factor and the nearest rotation, 300 trials at noise ÷ thickness = 10Each trial is a dot: horizontally the smallest singular value of the cross-covariance, signed negative when the polar factor is a reflection; vertically the Frobenius distance between the polar factor and the rotation with its determinant fixed. 198 trials sit at exactly 0 and 102 at exactly 2, with nothing in between: the largest deviation from 2 is 8.9·10⁻¹⁶. In every mirrored trial the rotation's residual exceeds the reflection's by 4σ₃ to a relative 8.3·10⁻¹³.-12-10-8-6-4-202468101200.511.522.5signed σ₃ of BᵀA, ×10⁻³‖U − R‖, Frobenius102 reflections, all at 2198 rotations, all at 0no trial in betweenlargest |‖U − R‖ − 2|8.9·10⁻¹⁶largest ‖U − R‖ when not mirrored0largest |gap − 4σ₃| ÷ 4σ₃8.3·10⁻¹³20 points, thickness 10⁻²; beyond the axis a dot is drawn at its edgea flip, not a drift
Fig. 3 The three-dimensional measurement: three hundred alignments, each at a distance of exactly 0 or exactly 2 between the polar factor and the corrected rotation. The same algebra puts every n-dimensional mirror at exactly 2.

So a mirror is never a small error, in three dimensions or in ten. It is a discrete event of fixed size, and an application tracking the polar factor across a sequence of datasets sees either nothing or a jump of 2. What changes with k is not the size of the event but its frequency, and — the next section — what the corrected answer is worth once the event has been removed.

That separation is also why the dimension can drop out of the rate without dropping out of the problem. The rate belongs to the thin block; the jump belongs to the algebra of the correction; and neither sees the n − k thick directions, which contribute a well-determined rotation of their own and nothing else. The dimension does not appear found a random projection’s distortion depending on the number of points and not on the number of coordinates, and the mirror has the same indifference for a different reason: the coordinates that are thick are decided long before the ones that are thin.

Where these alignments come from

The three-dimensional cases motivate the correction and rarely have more than one thin direction; a molecule or a scan is either flat or it is not. The n-dimensional cases are the ones where k ≥ 2 is ordinary.

Aligning two learned embeddings of the same vocabulary is an orthogonal Procrustes problem in hundreds of dimensions, and the spectrum of an embedding’s covariance typically decays over many directions rather than dropping off after a few. Shape analysis with many landmarks aligns configurations whose spread in the directions of small variation is comparable across several of them. Subspace tracking rotates a basis to follow a slowly changing signal, and the directions near the noise floor are many and nearly equal.

In each of these the question the earlier measurements asked — is the answer mirrored? — has an answer from the k × k model, and it is the less important question. The more important one is how much of the alignment is determined at all. The gap decides the eigenvector is the collection’s statement that nearly equal spectral values leave their vectors to the noise, and an answer that changes with the seed is the same fact seen as run-to-run variation. An n-dimensional alignment with two or more nearly equally thin directions is both: a rotation within their span that the data cannot fix, reported with a determinant that makes it look fixed.

The correction still fires, and no longer rescues the rotation

The three-dimensional measurement’s strongest finding was that the corrected rotation’s error does not grow as the disc thins: the thin direction’s only job is to decide the handedness, and the constraint det = +1 supplies exactly that. With more than one thin direction the thin directions have a second job.

The determinant-fixed rotation's error with one, two and three thin directions, z = 0.7Median distance of the determinant-fixed rotation from the true one, divided by the noise, over 400 alignments at z = 0.7, split into trials whose polar factor was a rotation and trials where it was a reflection and was fixed. n = 5 with 1 thin direction: 7.5% mirrored, error 0.59σ unmirrored and 0.61σ fixed; n = 5 with 2 thin directions: 22% mirrored, error 10.93σ unmirrored and 15.77σ fixed; n = 8 with 3 thin directions: 30% mirrored, error 24.23σ unmirrored and 29.26σ fixed.median error of the rotation ÷ noisen = 5, 1 thin7.5% mirrored0.59 not mirrored0.61 mirrored, fixedn = 5, 2 thin22% mirrored10.93 not mirrored15.77 mirrored, fixedn = 8, 3 thin30% mirrored24.23 not mirrored29.26 mirrored, fixedone thin direction: the fix is exacttwo or more: a rotation the data cannot see
Fig. 4 The median distance of the corrected rotation from the true one, divided by the noise, at z = 0.7, for alignments with one, two and three thin directions, split into alignments whose polar factor was a rotation and those that were mirrored and corrected.

With one thin direction in five dimensions, the corrected rotation’s median error at z = 0.7 is 0.59 noise-widths on alignments that were not mirrored and 0.61 on those that were — the same, as in three dimensions. The correction is exact.

With two thin directions it is 10.9 noise-widths on alignments that were never mirrored and 15.8 on those that were corrected. With three, in eight dimensions, 24.2 and 29.3. The error is an order of magnitude past the noise before any mirror is involved.

What has happened is that k thin directions of equal thickness define a k-dimensional subspace inside which the data says almost nothing about orientation. The signal that would fix a rotation within that subspace is t2t^2 times the differences between the thin directions’ spreads, and those differences are sampling fluctuations of the same order as the spreads themselves. A rotation by an angle of order σ/t inside the thin subspace changes the fit by less than the noise, so the polar factor picks one at random, and the determinant constraint — which fixes a sign, not an angle — cannot pick it back. The plane survives what its vectors do not is the same fact from the eigenvector side: two nearly equal eigenvalues leave the plane they span well determined and the vectors inside it arbitrary, and the thin subspace here is well determined while the rotation inside it is not.

The determinant-fixed rotation's error with one, two and three thin directions, z = 0.45Median distance of the determinant-fixed rotation from the true one, divided by the noise, over 400 alignments at z = 0.45, split into trials whose polar factor was a rotation and trials where it was a reflection and was fixed. n = 5 with 1 thin direction: 1.5% mirrored, error 0.60σ unmirrored and 0.68σ fixed; n = 5 with 2 thin directions: 4.8% mirrored, error 10.98σ unmirrored and 17.93σ fixed; n = 8 with 3 thin directions: 13% mirrored, error 26.08σ unmirrored and 38.99σ fixed.median error of the rotation ÷ noisen = 5, 1 thin1.5% mirrored0.60 not mirrored0.68 mirrored, fixedn = 5, 2 thin4.8% mirrored10.98 not mirrored17.93 mirrored, fixedn = 8, 3 thin13% mirrored26.08 not mirrored38.99 mirrored, fixedone thin direction: the fix is exacttwo or more: a rotation the data cannot see
Fig. 5 The same comparison at z = 0.45, where mirrors are rarer: 1.5, 4.8 and 13 per cent. The corrected rotation’s error with two and three thin directions is 11 and 26 noise-widths when nothing was mirrored.

At z = 0.45 the rates fall to 1.5, 4.8 and 13 per cent, and the errors do not: 0.60, 11.0 and 26.1 noise-widths on unmirrored alignments. At z = 1.5 they are 0.58, 9.1 and 20.5. The error per unit of noise is set by the number of thin directions and barely moves with z, which is the signature of a quantity the data does not determine rather than one it determines badly.

The rod, in any number of dimensions

The earlier measurement found one place the correction was badly conditioned in three dimensions: a rod, whose two short axes are nearly equal, where the flipped rotation’s sensitivity grew like (σ2\sigma_2 + σ3\sigma_3)/(σ2\sigma_2σ3\sigma_3).

Sensitivity of the nearest rotation to a perturbation of the cross-covariance, with and without the flipThe largest change in the rotation per unit change in the cross-covariance, found over forty random directions, against σ₃/σ₂ on a logarithmic vertical axis, with σ₂ = 0.1. Without the flip it follows √2/(σ₂ + σ₃): 13.9 at σ₃/σ₂ = 0.01 and 7.4 at 0.98. With the flip it follows √2/(σ₂ − σ₃): 15.7 and then 724, so the correction is well conditioned for a flat point set and badly conditioned for a rod.00.20.40.60.81110¹10²10³σ₃ ÷ σ₂ — from a disc (left) to a rod (right)‖ΔR‖ ÷ ‖ΔM‖flipped, measured√2/(σ₂ − σ₃)not flipped, measured√2/(σ₂ + σ₃)near a rodnot flipped, σ₃/σ₂ = 0.987.4flipped, σ₃/σ₂ = 0.98724ratio97noise-free shapes, forty random perturbations of relative size 10⁻⁹the fix is cheap on a disc and dear on a rod
Fig. 6 The three-dimensional measurement of the corrected rotation’s sensitivity against the ratio of the two smallest singular values. On a disc the correction is well conditioned; on a rod, with two nearly equal short axes, it is 97 times more sensitive.

The k ≥ 2 case in n dimensions is that rod, generalised. A three-dimensional rod has one long axis and two equal short ones; an n-dimensional slab with two thin directions has n − 2 long axes and two equal short ones. The sensitivity that measurement found — a small change in the cross-covariance rotating the answer by a large angle about the long axis — is here a rotation within the thin plane, and its size is the eleven noise-widths above.

So the three-dimensional picture had two cases that looked like opposite edge cases, a disc and a rod, and the n-dimensional picture says which one is typical. A point cloud in feature space is rarely thin in exactly one direction; its spectrum of spreads trails off over many. Those whose k smallest spreads are nearly equal are rods in the sense that matters, and for them the determinant correction repairs the handedness of an answer whose orientation within the thin subspace was never determined.

What an n-dimensional alignment should report

The measurements suggest three quantities, each available from the singular value decomposition the alignment already computes.

The number of thin directions and their spread. How many of the cross-covariance’s singular values are small compared with the rest, and how close to each other they are. One small singular value well separated from the next is the three-dimensional disc, and the correction is exact. Two or more close together is the rod.

The noise ratio for the thin block. z, from the noise estimate and the thin spread, reads the mirror rate off the k × k model’s curve for that k — which needs only k and the point count, and a table of such curves is small.

The rotation’s own uncertainty within the thin subspace. When k ≥ 2 this is the quantity to report, and the mirror is secondary: a user comparing two embeddings needs to know that the alignment’s rotation within the thin subspace is arbitrary to about σ/t, and a correct determinant does not tell them. A test with no answer in it proposed permuting the columns and checking whether the answer moved; here the equivalent test is to perturb the data by its noise and check whether the rotation within the thin subspace moves by more than the noise.

What this rests on

Dimensions three to ten, one to three thin directions of equal thickness 0.01, twenty or forty points, six hundred alignments per measured point and four hundred per error comparison; the k × k model from twenty thousand draws. Unequal thin spreads, heavier-tailed noise and weighted fits are not measured. The model’s derivation drops the coupling between the thick and thin blocks, which is second order in σ/t times t; the agreement within five points says that is safe over the range drawn, and it is not a bound.

The error comparison uses the median, because the error within a thin subspace is heavy-tailed: an arbitrary angle is sometimes a large one. The means would be larger and would say the same thing.

The claim that has to fail

The claim is the independent-coins model: k thin directions mirror at the rate k independent flips, each at the one-direction rate, would give. At z = 0.45 on twenty points, one thin direction’s determinant is negative on 3.0 per cent of draws, and two coins at that rate give 5.8; the two-direction determinant gives 9.9. The refusal is fed the claim that the two are within two points and fails. The k = 1 model’s agreement with three, five and ten dimensions beside it is what makes the failure the coins’ rather than the model’s.

Still open: unequal thin spreads, and a correction for the subspace

Thin directions of different thickness. Every thin direction here has the same spread, which is the worst case for the rotation inside the subspace. Two thin directions a factor of ten apart should behave like one thin direction for the rotation and like something between one and two for the mirror, and where between is the interpolation that would make the k × k model usable on a real spectrum.

A correction that reports what it cannot fix. The determinant correction repairs a sign. A correction that also returned the angular uncertainty of the rotation within the thin subspace — from the gaps between the small singular values and the noise estimate — would turn an answer that is silently arbitrary into one that says so, and whether that uncertainty estimate is accurate enough to trust is the measurement it would need.

Iterating to the rotation. The earlier measurement named an inverse-free iteration that converges to the rotation rather than to the polar factor. In n dimensions with k ≥ 2 thin directions it would converge to one of a continuum of nearly equally good rotations, and which one it chooses, and whether its choice is stable from one noisy dataset to the next, is the version of this question for an iteration.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DeterminantOrthogonalityPolar decompositionRotationSensitivitySingular value decomposition