A mirror decided in the thin directions
Worth reading first: The nearest orthogonal matrix.
A rotation that comes back mirrored measured the reflection problem in three dimensions, where the application that motivates it — aligning a molecule, a spacecraft, a scan — lives. It closed on the dimensions where there is no picture to look at. The correction generalises unchanged: find the singular value decomposition of the cross-covariance, take the polar factor, and if its determinant is −1 flip the direction of the smallest singular value. The earlier measurement’s reasoning predicted that the probability of needing the flip should depend on the smallest singular value alone, and noted that in higher dimensions there are more directions a noisy estimate can be thin in. Registration of point clouds in feature space, alignment of two embeddings, Procrustes analysis of shapes with many landmarks — all are n-dimensional, and all can have a mirror nobody sees.
Five precise points are five points settled the three-dimensional rate’s dependence on the point count and on weights. This essay sets the dimension and the number of thin directions free.
The experiment in n dimensions
m points in n dimensions, with n − k coordinates of unit spread and k coordinates of spread t = 0.01: a flat slab with k thin directions. A random rotation with determinant +1 is applied, independent noise σ is added to every coordinate, and the n × n cross-covariance is formed and decomposed. Its polar factor is the nearest orthogonal matrix to it, as in three dimensions, and it is a reflection when its determinant, computed by LU with pivoting, is negative; the corrected rotation flips the smallest singular direction.
The noise is set through , the ratio the previous two measurements established for three dimensions, so that a curve drawn against z can be compared across point counts. Six arrangements are measured with six hundred alignments at each of seven values of z: one thin direction in three, five and ten dimensions; two in five and in ten; three in eight.
The dimension drops out
With one thin direction, three dimensions mirror on 1.7 per cent of alignments at z = 0.45, 9.7 at 0.7, 18.0 at 1 and 26.3 at 1.5. Five dimensions give 2.7, 13.2, 21.3 and 27.7; ten dimensions with forty points give 3.2, 11.5, 17.8 and 26.8. The differences between the three are at most three and a half points, which is what six hundred trials a cell allow, and they are not ordered by the dimension.
With two thin directions, five dimensions give 9.5, 23.0, 34.2 and 40.0 at the same four values of z, and ten dimensions give 10.3, 23.8, 36.2 and 43.3. Again the dimension moves nothing.
The number of thin directions moves everything. At z = 0.7 one thin direction mirrors on 10 to 13 per cent of alignments, two on 23 to 24, and three on 32. At z = 1.5, one reaches about 27, two about 40 to 43, and three 49.5 — a coin, within sampling. More thin directions reach the coin sooner.
The handedness lives in a k × k block
The reason the dimension drops out is visible in the structure of the cross-covariance. In the basis of the point set’s own spread, the n − k thick directions carry signal of order m and noise of order , so their contribution to the polar factor is decided with a large margin and has determinant +1 in every trial that matters. The k thin directions carry signal of order and noise of order , and all the uncertainty about the determinant’s sign is in that k × k block. Divided by t²m, the block is
where W is the k × k matrix of sums of the thin coordinates’ products — a Wishart matrix on m points — and G is a k × k matrix of standard Gaussians. There is no n in it. The model predicts that the mirror rate is the probability that this k × k matrix has a negative determinant, and that probability can be computed by drawing W and G directly, twenty thousand times at each z.
For one thin direction on twenty points the model gives 3.0, 9.5, 16.8 and 26.0 per cent at z = 0.45, 0.7, 1 and 1.5, against the three measured curves above. For two thin directions on twenty points it gives 9.9, 21.7, 32.7 and 40.7, against five dimensions’ 9.5, 23.0, 34.2 and 40.0. For three on forty points, 14.2, 30.8, 41.0 and 48.2, against eight dimensions’ 13.5, 31.8, 39.8 and 49.5. Every measured point in the range is within five points of its model, and most within two.
So the mirror in n dimensions is a statement about a k × k random matrix, and the three-dimensional case the earlier essays measured is the scalar case: one thin direction, and a determinant that is simply the sign of W/m + zG with both terms numbers.
Several thin directions are not several coins
The obvious simplification is that each thin direction flips independently, with the one-direction probability p, and the determinant is negative when an odd number of them flip. That gives (1 − (1 − 2p)ᵏ)/2, which is easy to compute and wrong.
On forty points at z = 0.45, one thin direction mirrors on 2.2 per cent of draws. Two independent flips at that rate would give 4.2 per cent; the two-direction determinant gives 7.0. Three independent flips would give 6.2; the three-direction determinant gives 14.0. At z = 0.7 the two curves are closer but still apart, and they meet only near the coin, where both are fifty.
The determinant is not a product of independent signs because the k × k block has off-diagonal entries. Its determinant can be driven negative by the diagonal signs, as the coins model says, and also by large off-diagonal noise when the diagonal signals are close to each other — a second route to a mirror that one thin direction does not have. The more thin directions, the more off-diagonal entries, and the further the determinant runs ahead of the coins.
That makes independence the wrong way round for a practitioner estimating how often a high-dimensional alignment will need correcting: it underestimates the rate by a factor of about two at small noise ratios, which is where a practitioner would hope the correction is rare.
What a mirror looks like in any dimension
One property of the three-dimensional mirror carries over exactly, and it is worth stating because it is what a code would see. The polar factor U and the corrected rotation R differ only in the sign of one singular direction, so U − R is the outer product of that direction’s left and right singular vectors, times two. Its Frobenius norm is 2 in every dimension, whatever n and k are.
So a mirror is never a small error, in three dimensions or in ten. It is a discrete event of fixed size, and an application tracking the polar factor across a sequence of datasets sees either nothing or a jump of 2. What changes with k is not the size of the event but its frequency, and — the next section — what the corrected answer is worth once the event has been removed.
That separation is also why the dimension can drop out of the rate without dropping out of the problem. The rate belongs to the thin block; the jump belongs to the algebra of the correction; and neither sees the n − k thick directions, which contribute a well-determined rotation of their own and nothing else. The dimension does not appear found a random projection’s distortion depending on the number of points and not on the number of coordinates, and the mirror has the same indifference for a different reason: the coordinates that are thick are decided long before the ones that are thin.
Where these alignments come from
The three-dimensional cases motivate the correction and rarely have more than one thin direction; a molecule or a scan is either flat or it is not. The n-dimensional cases are the ones where k ≥ 2 is ordinary.
Aligning two learned embeddings of the same vocabulary is an orthogonal Procrustes problem in hundreds of dimensions, and the spectrum of an embedding’s covariance typically decays over many directions rather than dropping off after a few. Shape analysis with many landmarks aligns configurations whose spread in the directions of small variation is comparable across several of them. Subspace tracking rotates a basis to follow a slowly changing signal, and the directions near the noise floor are many and nearly equal.
In each of these the question the earlier measurements asked — is the answer mirrored? — has an answer from the k × k model, and it is the less important question. The more important one is how much of the alignment is determined at all. The gap decides the eigenvector is the collection’s statement that nearly equal spectral values leave their vectors to the noise, and an answer that changes with the seed is the same fact seen as run-to-run variation. An n-dimensional alignment with two or more nearly equally thin directions is both: a rotation within their span that the data cannot fix, reported with a determinant that makes it look fixed.
The correction still fires, and no longer rescues the rotation
The three-dimensional measurement’s strongest finding was that the corrected rotation’s error does not grow as the disc thins: the thin direction’s only job is to decide the handedness, and the constraint det = +1 supplies exactly that. With more than one thin direction the thin directions have a second job.
With one thin direction in five dimensions, the corrected rotation’s median error at z = 0.7 is 0.59 noise-widths on alignments that were not mirrored and 0.61 on those that were — the same, as in three dimensions. The correction is exact.
With two thin directions it is 10.9 noise-widths on alignments that were never mirrored and 15.8 on those that were corrected. With three, in eight dimensions, 24.2 and 29.3. The error is an order of magnitude past the noise before any mirror is involved.
What has happened is that k thin directions of equal thickness define a k-dimensional subspace inside which the data says almost nothing about orientation. The signal that would fix a rotation within that subspace is times the differences between the thin directions’ spreads, and those differences are sampling fluctuations of the same order as the spreads themselves. A rotation by an angle of order σ/t inside the thin subspace changes the fit by less than the noise, so the polar factor picks one at random, and the determinant constraint — which fixes a sign, not an angle — cannot pick it back. The plane survives what its vectors do not is the same fact from the eigenvector side: two nearly equal eigenvalues leave the plane they span well determined and the vectors inside it arbitrary, and the thin subspace here is well determined while the rotation inside it is not.
At z = 0.45 the rates fall to 1.5, 4.8 and 13 per cent, and the errors do not: 0.60, 11.0 and 26.1 noise-widths on unmirrored alignments. At z = 1.5 they are 0.58, 9.1 and 20.5. The error per unit of noise is set by the number of thin directions and barely moves with z, which is the signature of a quantity the data does not determine rather than one it determines badly.
The rod, in any number of dimensions
The earlier measurement found one place the correction was badly conditioned in three dimensions: a rod, whose two short axes are nearly equal, where the flipped rotation’s sensitivity grew like ( + )/( − ).
The k ≥ 2 case in n dimensions is that rod, generalised. A three-dimensional rod has one long axis and two equal short ones; an n-dimensional slab with two thin directions has n − 2 long axes and two equal short ones. The sensitivity that measurement found — a small change in the cross-covariance rotating the answer by a large angle about the long axis — is here a rotation within the thin plane, and its size is the eleven noise-widths above.
So the three-dimensional picture had two cases that looked like opposite edge cases, a disc and a rod, and the n-dimensional picture says which one is typical. A point cloud in feature space is rarely thin in exactly one direction; its spectrum of spreads trails off over many. Those whose k smallest spreads are nearly equal are rods in the sense that matters, and for them the determinant correction repairs the handedness of an answer whose orientation within the thin subspace was never determined.
What an n-dimensional alignment should report
The measurements suggest three quantities, each available from the singular value decomposition the alignment already computes.
The number of thin directions and their spread. How many of the cross-covariance’s singular values are small compared with the rest, and how close to each other they are. One small singular value well separated from the next is the three-dimensional disc, and the correction is exact. Two or more close together is the rod.
The noise ratio for the thin block. z, from the noise estimate and the thin spread, reads the mirror rate off the k × k model’s curve for that k — which needs only k and the point count, and a table of such curves is small.
The rotation’s own uncertainty within the thin subspace. When k ≥ 2 this is the quantity to report, and the mirror is secondary: a user comparing two embeddings needs to know that the alignment’s rotation within the thin subspace is arbitrary to about σ/t, and a correct determinant does not tell them. A test with no answer in it proposed permuting the columns and checking whether the answer moved; here the equivalent test is to perturb the data by its noise and check whether the rotation within the thin subspace moves by more than the noise.
What this rests on
Dimensions three to ten, one to three thin directions of equal thickness 0.01, twenty or forty points, six hundred alignments per measured point and four hundred per error comparison; the k × k model from twenty thousand draws. Unequal thin spreads, heavier-tailed noise and weighted fits are not measured. The model’s derivation drops the coupling between the thick and thin blocks, which is second order in σ/t times t; the agreement within five points says that is safe over the range drawn, and it is not a bound.
The error comparison uses the median, because the error within a thin subspace is heavy-tailed: an arbitrary angle is sometimes a large one. The means would be larger and would say the same thing.
The claim that has to fail
The claim is the independent-coins model: k thin directions mirror at the rate k independent flips, each at the one-direction rate, would give. At z = 0.45 on twenty points, one thin direction’s determinant is negative on 3.0 per cent of draws, and two coins at that rate give 5.8; the two-direction determinant gives 9.9. The refusal is fed the claim that the two are within two points and fails. The k = 1 model’s agreement with three, five and ten dimensions beside it is what makes the failure the coins’ rather than the model’s.
Still open: unequal thin spreads, and a correction for the subspace
Thin directions of different thickness. Every thin direction here has the same spread, which is the worst case for the rotation inside the subspace. Two thin directions a factor of ten apart should behave like one thin direction for the rotation and like something between one and two for the mirror, and where between is the interpolation that would make the k × k model usable on a real spectrum.
A correction that reports what it cannot fix. The determinant correction repairs a sign. A correction that also returned the angular uncertainty of the rotation within the thin subspace — from the gaps between the small singular values and the noise estimate — would turn an answer that is silently arbitrary into one that says so, and whether that uncertainty estimate is accurate enough to trust is the measurement it would need.
Iterating to the rotation. The earlier measurement named an inverse-free iteration that converges to the rotation rather than to the polar factor. In n dimensions with k ≥ 2 thin directions it would converge to one of a continuum of nearly equally good rotations, and which one it chooses, and whether its choice is stable from one noisy dataset to the next, is the version of this question for an iteration.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A condition number for one eigenvalue — both name orthogonality, sensitivity
- An iteration that only multiplies — both name orthogonality, polar decomposition
- The best approximation there is — both name orthogonality, singular value decomposition
- The condition number is an amplifier — both name orthogonality, sensitivity
- The number that decides nothing — both name determinant, orthogonality
Named objects
A flat tag is an object no other essay names yet.
DeterminantOrthogonalityPolar decompositionRotationSensitivitySingular value decomposition