Concept

Swamp — where it appears

A long plateau in an alternating fit during which the objective falls and the factors grow without bound. It is the geometry of a non-closed set arriving as an iteration count, and the size of the terms is the only thing that tells it from slow convergence.

Named by 9 essays across one field — each of them below, with the objects they name alongside it.

110¹10²10³10⁴10⁻¹⁴10⁻¹¹10⁻⁸10⁻⁵10⁻²10¹sweeprelative error, and the largest term's sizerising: the swamp's largest rank-one termfalling, slowly: its errorfalling, once: a fit with an answera plateau with a rising floorswamp error0.0014swamp term10term growth2.2benign error9.7·10⁻¹⁵benign term growth1the error alone cannot tellthe size of the terms can

An iteration that walks out of the set

Every sweep of alternating least squares is the exact minimiser of its own subproblem, so the objective can only fall. What it cannot do is converge, when the target's nearest rank-r point is not in the rank-r set — and a plateau at a small residual looks identical to slow convergence unless the size of the terms is plotted beside it.

tensor · Alternating least-squares
6 × 6 × 6, rank 3 · 9 ≥ 81.00002 × 2 × 2, rank 3 · 6 < 80.02076 × 6 matrix, rank 3 · no condition0.1089worst agreement between two runs about the factors, 0 to 1every run fits to 3·10⁻¹³every run fits to 1.2·10⁻¹¹every run factorises to 2.4·10⁻¹⁵ and none agrees with anotherthe only thing that improvestensor, Kruskal holds1tensor, Kruskal fails0.021matrix0.11worst residual1.2·10⁻¹¹three successful fitsone recoverable answer

A factorisation that is unique for once

A rank-r factorisation of a matrix is never unique — AB is (AM)(M⁻¹B) for any invertible M, so no factor means anything on its own. For three indices a checkable condition on the factors' k-ranks makes the decomposition unique up to permuting and scaling the terms, and it holds generically.

tensor · Uniqueness
01210⁻¹⁶10⁻¹³10⁻¹⁰10⁻⁷10⁻⁴10⁻¹terms fitted1 − the largest cosine between two fitted terms3 — the tensor's rank45below this line, two terms are the same termthe residual says nothing3 terms, collinearity0.524 terms, collinearity15 terms, collinearity13 terms, error2.1·10⁻¹⁴4 terms, error10⁻¹³4 terms, sweeps76one term too manyand two of them cancel

One term too many

A tensor built from three rank-one terms, fitted with four, reaches a relative error of 1.0·10⁻¹³ in seventy-six sweeps. Two of the four terms come back with a cosine of 0.99927 between them, subtracting each other, and the smallest weight is a fifth of the largest rather than a rounding. Nothing in the residual says any of it.

tensor · Alternating least-squares
10⁻¹⁰10⁻⁸10⁻⁶10⁻⁴10⁻²10⁻¹⁰10⁻⁸10⁻⁶10⁻⁴10⁻²1λ, the ridge on every subproblemerror, and the largest rank-one termthe terms a swamp was growingthe error it costs a fit that has an answerthe error it costs one that has nota bound, at its own priceterms, no ridge6.7terms, heaviest1.2sweeps, no ridge3000sweeps, heaviest69benign error at λ0.089terms still the right terms1the plateau is boundedand the answer is biased by exactly λ

The repair that costs exactly itself

A ridge on every subproblem is the standard cure for a swamp and it works — the terms stop at 1.61 instead of climbing past 6.7, and a run that never finished finishes in 798 sweeps. On a tensor that does have an answer the error it costs is the ridge itself, to within a factor of two, at every setting from 10⁻¹⁰ to 10⁻¹. And the terms it returns are still the right terms.

tensor · Alternating least-squares
01210¹10²10³the three kinds of runsweep at which the test firesno answer to reachbadly conditionedordinaryhollow: never fired, drawn where the run endedwhat it fires on, and whenthreshold0.25boundary, fired6boundary, median sweep22collinear, fired6collinear, median sweep759ordinary, fired0the value does not separate themand the sweep does

The test that is a deadline

The quantity that separates a boundary from slow convergence fires on every run that has nothing to reach, at sweep 21 of four thousand, and on no ordinary fit. It also fires on every badly conditioned fit that does converge — at sweep 758. No threshold between 0.10 and 0.40 separates those two, and the sweep it fires at separates them by a factor of thirty-six.

tensor · Alternating least-squares
at rank threelowest across tensors1highest across tensors1at rank fourlowest across tensors0.056highest across tensors0.25234500.20.40.60.81rank fittedagreement among the starts that reach itsix planted tensorsdashed: the true rankeach line one tensor, fitted at ranks two to fivefive starting points at every rank

The rank a sweep can vouch for

To decide a tensor's rank, fit it at one rank after another and stop where something changes. Three things can be read off each rank's fits — the error, how nearly parallel the terms are, and whether fits from different starting points agree — and over six planted rank-three tensors at four noise levels they decide it 24, 12 and 15 times out of 24; the ratio of consecutive errors, which needs nothing, decides it 24 times. The error is right every time and has to be told the noise level. The collinearity is a coin toss. Agreement among starts is never wrong and, under noise, cannot decide on nine of the twenty-four — and the rank past the answer, where every decision is made, costs ten times the answer.

tensor · Alternating least-squares
no noisesweeps to stop49final error2.5·10⁻¹⁴1% noisesweeps to stop6000largest term, ÷ data0.67110¹10²10³10⁻¹⁸10⁻¹⁵10⁻¹²10⁻⁹10⁻⁶10⁻³1sweeprelative change in the error, per sweepthe stopping test1% noiseno noise: stops at 10⁻¹⁴the clean run stops; the noisy one crawlsits terms stay the size they were

A fit that has an answer and cannot stop

Fit a noisy rank-three tensor with four terms and the prediction was that the spare term would find real structure in the noise and stop pairing off with the others. It does not: the largest cosine between two fitted terms has a median of 0.977 at 1% noise against 0.975 without noise. What changes is the solver. Without noise the four-term fit stops in fifty sweeps; with noise it never stops: its error keeps moving by a part in a million a sweep for six thousand sweeps, while on most tensors its terms stand still. The extra term takes exactly its share of the noise, and the error is settled by sweep 150.

tensor · Alternating least-squares
start 1, errorreaches2.8·10⁻¹³stopped by δ = 0.010.0099stopped by δ = 0.0010.0098stopped by δ = 10⁻⁴2.6·10⁻¹³110¹10²10³10⁻¹²10⁻¹⁰10⁻⁸10⁻⁶10⁻⁴10⁻²sweeprelative errorδ = 0.01δ = 0.001δ = 10⁻⁴δ = 10⁻⁵a plateau is a still errorand a still error is what the test asks for

A still error is not a settled one

A CP fit with one term too many on a noisy tensor never meets the usual stopping test, and its error has settled by sweep 150. A test that asks only whether the error has moved by less than δ of itself over ten sweeps stops those fits by sweep 25, and over twenty-four noisy tensors it names the same rank as the usual test every time, at δ = 10⁻² for 30,250 sweeps against 311,458. But δ = 10⁻² is the tolerance a rank decision states, and on a slow fit that converges it stops five starts of six on a plateau ten orders above the answer and names the wrong rank for four tensors of six. A tenth of the decision's tolerance keeps every decision on both families.

tensor · Alternating least-squares
per unit of perturbationtrue move, θ = 10⁻³200ALS off by, θ = 10⁻³200ALS off by, θ = 17.9·10⁻⁴10⁻³10⁻²10⁻¹110¹10²10³angle θ between two columns of the third factorchange ÷ perturbation110⁻¹10⁻²10⁻³true decomposition's moveALS's distance from itKruskal: unique at every θ hereθ closes to the leftunique, and not determined

Unique, and not determined

Kruskal's condition says a CP decomposition is unique, and it says it as a yes or a no. Turn two columns of one factor of a 4 × 4 × 4 rank-three tensor toward each other and the condition keeps saying yes at every angle above zero, by one to spare. What it does not say is how well the unique decomposition is determined. The Jacobian of the map from factors to tensor loses its smallest singular value in proportion to the angle, and the true decomposition of a tensor perturbed by one part in 10⁸ moves by 1.6 parts at θ = 0.1 and by 200 at θ = 10⁻³. Alternating least squares, started at the exact answer, finds the moved decomposition while the angle is large, slows by ten times or more on the way to θ = 0.1, and below θ = 0.03 never arrives: at θ = 10⁻³ its own stopping test fires after 11 to 65 sweeps with its terms 200 to 420 perturbations from the answer and its residual within a per cent of the answer's.

tensor · Uniqueness

Named alongside it

The objects these essays reach for when they reach for this one.

Alternating least-squaresCP decompositionBorder-rankIll-posed problemTensor rankUniquenessDegeneracyStopping criterionCondition squaringConvergence rateKruskal's conditionLow-rank approximation

All concepts