Concept

Spectral decay — where it appears

How fast a matrix's singular values fall away, which decides whether a low-rank approximation, a deflation or a sketch buys anything at all. It is what every low-rank method's benefit is a function of, and a matrix whose singular values are flat gives a randomised method nothing to find.

Named by 17 essays across 5 fields — each of them below, with the objects they name alongside it.

051015202530354010⁻¹⁷10⁻¹⁴10⁻¹¹10⁻⁸10⁻⁵10⁻²10¹singular value, in orderσ ⁄ σ₁eight digitsan independent draw per entry1 ⁄ r, falling to the floora cliff, and a control1/r rank at 10⁻⁸5log r rank at 10⁻⁸5noise rank at 10⁻⁸96σ₂ ⁄ σ₁0.024σ₆ ⁄ σ₁2.8·10⁻⁹the block has full rankand five useful columns

A block nobody can call sparse

A 96 × 96 block of a kernel matrix has ninety-six nonzero singular values and five that matter. It has no zero entries, it is not described by fewer numbers than it contains, and neither of the two ways this collection already knows to make a large matrix affordable applies to it.

hierarchy · Off-diagonal rank
0369121501122334455digits asked for, −log₁₀ εcolumns keptwhat the geometry promiseswhat the matrix costsa rank is a number of digitscolumns a decade0.55bound, a decade3.3rank at 10⁻⁸5bound at 10⁻⁸28q0.5the shape is rightand the constant is not

A rank that is a number of digits

Ask a kernel block for two digits and it costs two columns; ask for fourteen and it costs nine. The curve is a straight line at 0.55 columns a decade, and the bound the geometry gives is a straight line too — at 3.32, which is the same shape and six times the price.

hierarchy · Off-diagonal rank
04812162010⁻¹10⁻⁰.⁵1target rank k‖A − Aₖ‖₂published boundrandomisedσₖ₊₁, optimalhow far apart the three areworst seed spread1.1bound / median at k = 1211median / optimum at k = 12160×60, 6 seeds, oversampling p = 5band is best to worst

Randomisation does not create structure

On a matrix whose singular values are all equal, a rank-ten randomised approximation has error 1.0 — and so does the optimal deterministic one. Neither achieved anything, and only one of them is usually sold with the implication that it might.

randomised · Randomised
-1012301020304050log₂ of the segment lengthcolumns above 10⁻⁸cos(κr) ⁄ r at κ = 401 ⁄ r, same pointsthe case with no answerq, held fixed0.51/r at L = 0.561/r at L = 86cos(κr)/r at L = 0.512cos(κr)/r at L = 853the geometry did not moveand the rank did

The kernel with nothing to compress

Hold the geometry fixed at q = ½, fix the wavelength, and scale the picture up by sixteen. A smooth kernel needs six columns at every scale. An oscillatory one needs twelve, sixteen, twenty-two, thirty-three, fifty-three, and there is no scale at which it stops.

hierarchy · Off-diagonal rank
1112131415193111.365129.731148.096166.461probes takenrunning estimate of the tracenormal±1one probe, no errorthe exact trace99±1 variance, this matrix0±1 variance, rotated57normal variance545the same spectrum in a general basiscosts the ±1 probe its whole advantage

Counting what cannot be looked at

The trace is n additions and one of the most expensive quantities in the subject to estimate, because the matrices whose trace is wanted are never stored. Hutchinson's estimator is unbiased with one line of algebra — and its variance depends on which random vector is used, by a factor that is a property of the matrix, and on a diagonal matrix one choice is exact from the first probe and the other is not.

randomised · Trace estimation
10¹10²10⁻⁴10⁻³10⁻²10⁻¹1products with Arelative errorHutchinsonHutch++measured at equal costfitted rate, Hutchinson-0.5fitted rate, Hutch++-1.7error at 96, Hutchinson0.017error at 96, Hutch++0.0022both axes count products with Aso the sketch is paid for in the picture

A rate that belongs to the matrix

Hutchinson's fitted exponent sits near a half on every spectrum measured. Hutch++'s runs from −7.15 to −0.67 across the same four budgets, decided entirely by how fast the singular values fall — so one of the two methods has a convergence rate and the other has a rate per matrix. The ±1 probe's advantage moves the same way, from 1.56× at n = 10 to 1.09× at n = 120.

randomised · Trace estimation
-13-11-9-7-5-3-102468log₁₀ of the accuracy asked forlargest train rank0.30 of rank per decadea cost that is typed inentries216stored at 10⁻²90stored at 10⁻¹²288rank per decade0.3worst error ⁄ tolerance0.83the storage is chosena constant of rank a decade

The digit that costs more than the tensor

Ask a three-index reciprocal tensor on six points a side for seven digits and its train is 288 numbers against 216 entries. The break-even rank is n − 1 at all four grids measured, and a train that reaches it fits with exactly n numbers to spare.

tensor · Tensor train
13579111310⁻¹⁴10⁻¹¹10⁻⁸10⁻⁵10⁻²10¹index ksingular valueσ₁ · 10⁻¹⁴, the thresholdone matrix, three rankspartitionings7lowest rank10highest rank12threshold7.3·10⁻¹⁴σ₁7.3the curves separate in the noiseand the threshold is drawn through it

A rank that depends on the thread count

One 60 × 14 matrix, one threshold, seven partitionings of the inner products that build its Gram matrix — and numerical ranks of 12, 12, 12, 10, 10, 11 and 11. Not a digit of an answer: the number of columns a model built from this matrix would have.

machine · Rank
48 products, 24 drawsbest share, decay 0.70.45best share, decay 0.950.2worst cost of a third2.800.10.20.30.410⁻³10⁻²10⁻¹share spent on the sketchmedian relative errorthe published thirddecay 0.7decay 0.85decay 0.95large dots: the best split on each curveand the dashed line is the one a library picks

The split nobody is in a position to choose

Hutch++ spends two thirds of its budget on a sketch and a third on probes, and the third is published as a constant. Swept across six rates of spectral decay at a fixed budget of 48 products, the best share is 0.45 on the fastest and 0.00 on the slowest — sketch nothing at all — and the published third costs between 1.09 and 4.11 times the best error. The decay that decides it is readable from the sketch's own singular values, for products the estimator was going to spend anyway.

randomised · Trace estimation
40 drawsproducts at target 0.01696deflated277draws inside the target0.810⁻²10⁻¹10⁻²10⁻¹target standard errorrelative error reachedworst draw, no deflationmedian error, no deflationmedian error, deflatedthe grey diagonal is a calibrated rulethe medians are on it and the worst draws are not

A rule that reads only its own probes

A trace estimator is a mean of independent samples, so its own standard error is estimable from the samples and a stopping rule needs nothing the estimator does not already have. Over forty draws it is calibrated in the middle and not at the edge: at a target relative standard error of 1% the median error reached is 5.3·10⁻³ and the worst of forty is 2.9·10⁻² — three times the target. And the cost of the target is the estimator's own square root: tightening it from 3% to 1% takes the median probe count from 75 to 696.

randomised · Trace estimation
051015202530354010⁻⁴10⁻³10⁻²10⁻¹110¹rank of the basiserror, and what the rule readsprobe × 7.98largest probetrue error, Frobeniustrue error, spectralσₖ₊₁, the floorgeometric 0.8, one run, ten probesoptimal rank for ε = 0.111basis reaches ε at rank15probe below ε at rank20published estimate below ε at rank34ε = 0.1the dotted line across is ε = 0.1each rule stops where its curve crosses it

The rank a certificate charges

A randomised range finder can choose its own rank: grow the basis a column at a time and stop when ten fresh probes all come back short. With the published safety factor it never stopped early in any draw measured, and on a matrix whose singular values fall by 0.8 a step it stopped at rank 30 for a tolerance the best rank-11 approximation already meets. The nineteen extra columns are three separate prices — four for building the basis from random vectors, five because a probe reads more than the spectral norm, and ten for the constant — and the spectrum decides which of them dominates.

randomised · Randomised
111315171921232510⁻¹110¹sketch width l, for rank 10‖(I − QQᵀ)A‖ ÷ σ₁₁GaussianHadamardone nonzero a rowthree nonzeros a rowcoherent, geometric 0.8, width 20, twenty seedsGaussian: median 0.41, worst0.78Hadamard: median 0.37, worst0.94one nonzero a row: median 3.3, worst7.1three nonzeros a row: median 0.41, worst0.99solid: median · dashed: worst of twentybelow one: better than the best rank-10 error

A sketch that finds the columns it can see

A sparse sketch with one nonzero in each row is twenty times cheaper to apply than a Gaussian one, and on a matrix whose important directions are spread across its columns it finds the same range: a median error of 0.45 against 0.43. Put the same ten directions into ten particular columns and it is eight times worse — 3.33 against 0.41, with a worst draw of 7.1 — because two important columns hashed to one bucket are one direction. Three nonzeros a row repair it at a sixth of the Gaussian's cost, and a randomised Hadamard transform never had the problem.

randomised · Randomised
decay 0.8, 400 draws3% target, c = 10.663% target, c = 1.960.9410% target, c = 10.6611.251.51.7522.252.5556065707580859095100margin c on the standard errordraws inside the target, %3% target10% targeta normal tablethe dashed curve is 2Φ(c) − 1what a margin of c promises if the error is normal

The miss a normal table already priced

A trace estimator that stops when its own standard error reaches a target misses the target on about a third of draws, and the essay that measured it read the loose targets as the worst calibrated. Over 400 draws the loose target is the better covered — 76% at 10% against 64% at 3% — because the warm-up stops most of its runs with probes to spare. Where the criterion decides, the misses are a normal distribution's: a margin of c on the standard error buys what a normal table says, 94.8% at 1.96, and costs c² in probes, 3.86 times.

randomised · Trace estimation
median ratioone nonzero, t = 03.5one nonzero, t = 0.011.210⁻⁴10⁻³10⁻²10⁻¹11mixture t (t = 0 at the left edge)error ÷ σ₁₁one nonzero a rowheavy columns reservedthree nonzeros a rowGaussianthe left edge is exact coherencethe failure fades over four decades of mixing

The leverage that did not move

A one-nonzero sketch fails on a matrix whose leading directions sit on ten particular columns, and coherence — the largest column leverage — is the statistic that names the failure. Turn the directions away from their columns by a hundredth of a radian and the sketch's median error falls from 3.47 to 1.23 times σ₁₁ while the coherence stays at 6.40 to three figures. Giving the heaviest columns buckets of their own repairs the rest, but only when it reserves more buckets than the rank: ten reserved leave 1.16, sixteen reach 0.36, below the Gaussian's 0.41.

randomised · Randomised
points of coveragegain at c = 1, fast decays, best5.3worst0.750.811.21.41.61.822.22.42.6-30-20-10010products ÷ the sequential rule'scoverage gained, pointsdecay 0.8decay 0.9decay 0.97dashed: no gaina few points, bought with a fifth to double the probes

A spread measured on probes it does not average

A trace estimator that stops on its own standard error misses its target a few points more often than a normal table says, because the runs that stop earliest are the ones that underestimated their noise. Spend a pilot of probes only on the spread, fix the number of probes to average in advance, and the selection is gone: with Student's margin the two-stage rule covers 64.8 to 73.3 per cent at one standard error where the table says 68.3. With the normal's margin and a pilot of four it covers 58.5. At one standard error it recovers one to five points for a fifth to two fifths more probes; at 1.96 there was nothing to recover, and the guarantee costs a tenth to double.

randomised · Trace estimation
sixteenth trace, c = 1, %a new pilot every trace, fastest drift63the first pilot, frozen, fastest drift34the last trace's probes, carried, fastest drift65304050607080decay lost per traceestimates inside the target, %00.0050.010.02a new pilot every tracethe first pilot, frozenthe last trace's probes, carrieddashed: the normal tablea frozen spread fails; a carried one holds

A spread carried from the trace before

A two-stage trace estimator spends a pilot of probes learning its spread, and a computation that needs many traces of a slowly changing operator would rather pay for that once. Frozen at the first trace, the pilot is wrong by the sixteenth: on a spectrum that drifts from decay 0.9 to 0.86 the last estimate is inside a 3% target 52 times in a hundred against the table's 68, and at a faster drift 34. Carry instead the spread of the previous trace's own averaged probes — free, independent of this trace's, one step stale — and its sixteenth estimate covers between 64.5 and 69.5 per cent at every drift, for up to a quarter fewer products than a fresh pilot every time.

randomised · Trace estimation
1611162126313610⁻¹³10⁻¹¹10⁻⁹10⁻⁷10⁻⁵10⁻³10⁻¹10¹indexmagnitude|rₖₖ|σₖone factorisation, two verdicts‖AP − QR‖/‖A‖10⁻¹⁵|rₙₙ|1.1·10⁻¹²σₘᵢₙ10·10⁻¹³column interchanges33|rₙₙ| is never below σₘᵢₙso the cheap verdict errs one way only

A good curve and a bad verdict

The diagonal of a column-pivoted R is famous for the one matrix it is wrong about. On that matrix it is right about thirty-nine of its forty entries — every |rₖₖ| within a factor of six of the σₖ it stands for — and wrong by 4·10⁶ at the fortieth, which is the only one a rank verdict ever reads.

spectra · Rank

Named alongside it

The objects these essays reach for when they reach for this one.

Probabilistic boundsLow-rank approximationMatrix-freeHutchinson's estimatorTrace estimationFlop countNumerical rankRandom probeExact ground truthRandom projectionSingular valuesStopping criterion

All concepts