Randomised, and the guarantee that changes kind

A spread measured on probes it does not average

A trace estimator that stops on its own standard error misses its target a few points more often than a normal table says, because the runs that stop earliest are the ones that underestimated their noise. Spend a pilot of probes only on the spread, fix the number of probes to average in advance, and the selection is gone: with Student's margin the two-stage rule covers 64.8 to 73.3 per cent at one standard error where the table says 68.3. With the normal's margin and a pilot of four it covers 58.5. At one standard error it recovers one to five points for a fifth to two fifths more probes; at 1.96 there was nothing to recover, and the guarantee costs a tenth to double.

Worth reading first: Counting what cannot be looked at · A parameter that counts steps.

The miss a normal table already priced found that a trace estimator which stops when its own standard error reaches a target misses the target about as often as a normal distribution says it should — where the target decides the stop rather than the warm-up. At a 3% target and one standard error it covered 64 per cent of draws against the table’s 68, at 1.96 standard errors 94.8 against 95, and a margin of c on the standard error cost c2c^2 in probes, as the table predicts. The remaining gap, a few points at one standard error, it traced to selection. A run stops the first time its estimated standard error is small enough, so the runs that stop earliest are disproportionately the ones whose first probes happened to agree with each other — the ones that underestimated their own noise — and those are the runs most likely to miss.

The essay’s last section proposed the classical repair. Use a small, separate set of probes only to estimate the spread, and the rest only for the mean. Charles Stein’s two-stage procedure does exactly this: a pilot of p samples gives a sample standard deviation s; that fixes, before another sample is drawn, how many more are needed to reach the target; those are drawn and averaged. The number of samples averaged no longer depends on the samples being averaged, so there is nothing for the stopping to select on. The essay asked whether the three points of coverage this recovers at one standard error are worth the separate probes.

They are recovered, and they are not quite three. What the measurement adds is the part the proposal left out: the pilot’s spread is an estimate from p − 1 degrees of freedom, and the margin has to know that.

The rule, and the two margins it can use

The two-stage rule has one parameter besides the target and the margin, the pilot’s size p. It draws p Rademacher probes ziz_i and records ziTAziz_i^{\mathsf T}Az_i; their mean mˉ\bar m and standard deviation s estimate the trace and the spread of a single sample. For a target T on the relative standard error and a margin k, the number of probes to average is

t=⌈(k sT ∣mˉ∣)2⌉,t = \left\lceil \left(\frac{k\,s}{T\,|\bar m|}\right)^{2} \right\rceil,

and t fresh probes are drawn and averaged; the pilot’s probes are not reused. The margin k is either the normal table’s c — one for 68.3 per cent, 1.96 for 95 — or Student’s t quantile at the same coverage with p − 1 degrees of freedom. Student’s is larger, because the spread was estimated: with four pilot probes it is 1.197 instead of one and 3.18 instead of 1.96, with eight 1.077 and 2.36, with thirty-two 1.016 and 2.04. Under a normal model of the samples, Stein’s argument makes the second choice exact.

The matrices, the draws and the scoring are the previous essay’s, and the estimator is the one counting what cannot be looked at introduced, whose single-probe variance is a property of the matrix: a 60 × 60 symmetric matrix with eigenvalues 0.8k0.8^k, 0.9k0.9^k or 0.97k0.97^k, 400 seeded draws of each rule, and the share of draws whose true relative error is inside the target. The sequential rule it is compared with uses a warm-up equal to the pilot.

On the table, with Student’s margin

How often a trace estimate lands inside its target, against the margin on its standard error, at decay 0.9Against the margin c by which the target is divided before the rule is run: the share of 400 draws whose relative error is inside the target, for targets of 3% and 10%, with a warm-up of eight probes, beside the share a normal distribution puts within c standard deviations. At the 3% target the rule covers 63.7%, 90.3%, 94.8%, 98.5% at margins 1, 1.645, 1.96, 2.576, against the normal's 68.3%, 90.0%, 95.0%, 99.0%. At the 10% target it covers 75.8%, 88.0%, 93.5%, 98.3%.decay 0.9, 400 draws3% target, c = 10.643% target, c = 1.960.9510% target, c = 10.7611.251.51.7522.252.5556065707580859095100margin c on the standard errordraws inside the target, %3% target10% targeta normal tablethe dashed curve is 2Φ(c) − 1what a margin of c promises if the error is normal
Fig. 1 The sequential rule’s coverage against the margin on its standard error, for 3% and 10% targets at decay 0.9, beside the share a normal distribution puts within that many standard deviations.

The previous essay’s picture, redrawn for reference: at the 3% target, where the criterion decides, the sequential rule’s dots sit on the normal curve a point or four under it, and at the 10% target the warm-up lifts them above.

How often a trace estimate lands inside a 3% target at 1 standard error, against the pilot's size, decay 0.9Over 400 draws of a 60 × 60 matrix with eigenvalues 0.9 to the power k: the share whose relative error is inside 3%, for a two-stage rule that spends a pilot of 4, 8, 16 or 32 probes on estimating the spread and then fixes the number of probes it averages, with the margin from Student's t and from the normal table, and for the sequential rule with a warm-up of the same length. The dashed line is the table's 68.3%. two-stage, Student: 68.0%, 64.8%, 66.3%, 65.0%; two-stage, normal: 58.5%, 63.0%, 65.0%, 63.7%; sequential: 62.7%, 63.7%, 64.0%, 64.3%.decay 0.9, 3%, c = 1two-stage, Student, pilot 865sequential, warm-up 86450556065707580pilot probesdraws inside the target, %481632two-stage, Studenttwo-stage, normalsequentialdashed: the normal tableStudent's margin puts the rule on the table
Fig. 2 At a 3% target and one standard error: the share of 400 draws inside the target against the pilot’s size, for the two-stage rule with Student’s margin, with the normal margin, and for the sequential rule with a warm-up of the same length; the table’s 68.3% dashed. The dial sets the spectral decay.

At decay 0.9 the two-stage rule with Student’s margin covers 68.0, 64.8, 66.3 and 65.0 per cent with pilots of 4, 8, 16 and 32. The standard error of a share estimated from 400 draws is about 2.3 points at these values, and every one of those is within one and a half of them of the table’s 68.3. The sequential rule covers 62.7, 63.7, 64.0 and 64.3. Turn the dial: at decay 0.8 the two-stage rule covers 67.0 to 70.5 against the sequential rule’s 65.8 to 66.0; at 0.97, 66.8 to 73.3. Over all twelve settings at one standard error, the two-stage rule with Student’s margin lands within five and a half points of the table, most of them within two.

The normal margin does not. With a pilot of four it covers 58.5 per cent at decay 0.8 and 0.9 — ten points short, worse than the sequential rule it was meant to fix — and 64.5 at 0.97. With eight probes it is 63.0 to 68.3, and with sixteen or thirty-two 63.7 to 68.3: the normal margin’s shortfall shrinks as the pilot grows, because Student’s quantile shrinks towards it. The pilot’s spread is an estimate with three degrees of freedom at p = 4, and a standard deviation estimated from four numbers is too small about as often as it is too large, but when it is too small the rule averages too few probes and misses. Student’s quantile is exactly the correction for that, and at four probes it is a fifth larger than the normal’s.

The sequential rule’s shortfall barely moves with its warm-up — 62.7 per cent with four probes of warm-up, 64.3 with thirty-two — which is the signature of selection rather than of a small sample. A longer warm-up makes the first stopping decision better informed, but every later decision is still taken on the same stream whose luck is being judged, and the stream that happens to run quiet for a while is the one that stops. The two-stage rule never takes a decision on the probes it averages, and its coverage does not depend on the pilot’s size once Student’s margin is used. It is the same distinction a stopping rule that follows the run it is given met in an iterative regularisation: a rule that watches the quantity it is about to report inherits that quantity’s fluctuations as a bias.

At 1.96, nothing to recover

How often a trace estimate lands inside a 3% target at 1.96 standard errors, against the pilot's size, decay 0.9Over 400 draws of a 60 × 60 matrix with eigenvalues 0.9 to the power k: the share whose relative error is inside 3%, for a two-stage rule that spends a pilot of 4, 8, 16 or 32 probes on estimating the spread and then fixes the number of probes it averages, with the margin from Student's t and from the normal table, and for the sequential rule with a warm-up of the same length. The dashed line is the table's 95.0%. two-stage, Student: 94.3%, 93.8%, 93.8%, 96.0%; two-stage, normal: 87.3%, 92.5%, 92.0%, 95.0%; sequential: 94.8%, 94.8%, 94.8%, 94.8%.decay 0.9, 3%, c = 1.96two-stage, Student, pilot 894sequential, warm-up 8957883889398pilot probesdraws inside the target, %481632two-stage, Studenttwo-stage, normalsequentialdashed: the normal tableat 1.96 all three are near the table
Fig. 3 The same three rules at a 3% target and 1.96 standard errors, decay 0.9, against the table’s 95%.

At 1.96 standard errors the sequential rule was already on the table: 94.8 per cent at decay 0.9 for every warm-up, 94.0 to 94.3 at 0.8. Its selection costs little there because it stops late. Asked for 1.96 standard errors, the rule needs about 1.9621.96^2 times as many probes as at one, and by the time it stops its spread has been estimated from hundreds of samples rather than a handful; an early stop on an unlucky stream is rare. The two-stage rule with Student’s margin covers 93.8 to 96.0 at decay 0.9 and 93.0 to 94.3 at 0.8, the same coverage within the draws’ noise. With the normal margin and a pilot of four it covers 87.3 and 81.5.

What the guarantee costs

Probes a 3% target costs, two-stage against sequential, at one and at 1.96 standard errorsThe median number of matrix–vector products over 400 draws at decay 0.9, against the pilot or warm-up length, on a logarithmic axis. two-stage, 1.96: 635, 409, 337, 338; sequential, 1.96: 293, 293, 293, 293; two-stage, 1: 94, 92, 92, 108; sequential, 1: 76, 76, 76, 76. Student's margin at 1.96 standard errors is 3.18, 2.36, 2.13, 2.04 for pilots of 4, 8, 16, 32.the price of a guarantee1.96, pilot 4, two-stage ÷ sequential2.21.96, pilot 161.210²10³pilot probesproducts, median481632two-stage, 1.96sequential, 1.96two-stage, 1sequential, 1Student's margin on a small pilot is largethe pilot pays most where it is smallest
Fig. 4 Median matrix–vector products to reach a 3% target at decay 0.9, against the pilot or warm-up length, for both rules at one and at 1.96 standard errors.

The two-stage rule pays for its pilot twice: the pilot’s probes are spent on the spread and not averaged, and Student’s margin, being larger, asks for more fresh probes than the normal would. At decay 0.9 and one standard error the sequential rule’s median is 76 products at every warm-up; the two-stage rule’s is 94, 92, 92 and 108 for pilots of 4, 8, 16 and 32 — a fifth to two fifths more. At 1.96 the sequential rule takes 293; the two-stage rule 635 with a pilot of four, 409 with eight, and 337 and 338 with sixteen and thirty-two. A small pilot is the expensive one, because (3.18/1.96)2(3.18/1.96)^2 is 2.6 — Student’s quantile on three degrees of freedom more than doubles the probes a normal margin would ask for. A pilot of sixteen costs 15 per cent more than the sequential rule, and that is about the floor: the fresh probes are what the target needs, and the pilot is on top.

At decay 0.8 the numbers are larger and the ratios the same: 202 sequential against 212 to 235 at one standard error, 771 against 846 to 1,505 at 1.96.

How large a pilot should be

The cost curves have a minimum, and it can be predicted before any run. If a normal margin with a perfectly known spread would need N0N_0 probes, the two-stage rule with a pilot of p needs about

p+(tp−1c)2N0,p + \left(\frac{t_{p-1}}{c}\right)^{2} N_0,

the pilot plus the fresh probes, inflated by the square of Student’s quantile over the normal’s. A small pilot is cheap and inflates the rest; a large one inflates nothing and is itself the cost. Take N0N_0 from the sequential rule’s median, 76 products at one standard error and 293 at 1.96 for decay 0.9. At one standard error the formula gives 113, 96, 97 and 110 for pilots of 4, 8, 16 and 32; the measured medians are 94, 92, 92 and 108. At 1.96 it gives 776, 435, 362 and 349; measured, 635, 409, 337 and 338. The formula overestimates at the smallest pilot — Student’s margin covers the rare tiny spread, and the median run does not meet one — and otherwise tracks the measurement to within a tenth, and it puts the minimum where the measurement does: a pilot of eight to sixteen at one standard error, sixteen to thirty-two at 1.96.

The rule of thumb that follows is to spend on the pilot about as many probes as it takes for Student’s quantile to come within a few per cent of the normal’s at the coverage wanted — around ten degrees of freedom at one standard error and twenty or thirty at 1.96 — and never four. It is the same shape of argument as the split nobody is in a position to choose made for Hutch++'s sketch and probes: a fixed budget divided between learning something about the matrix and using what was learned, with the division decided by a number the first part estimates.

Points bought against probes paid

Coverage the two-stage rule gains over the sequential one, against the probes it costs, at the 3% targetOne dot per decay, margin and pilot: horizontally the two-stage rule's median products over the sequential rule's, vertically its coverage minus the sequential rule's, in points. At one standard error and the two faster decays the gain is 0.8 to 5.3 points; at decay 0.97 the sequential rule's warm-up covers more, and the gain is negative, down to -28.0. Filled dots are one standard error and open dots 1.96.points of coveragegain at c = 1, fast decays, best5.3worst0.750.811.21.41.61.822.22.42.6-30-20-10010products ÷ the sequential rule'scoverage gained, pointsdecay 0.8decay 0.9decay 0.97dashed: no gaina few points, bought with a fifth to double the probes
Fig. 5 For every decay, margin and pilot at the 3% target: the two-stage rule’s coverage minus the sequential rule’s, in points, against its median products over the sequential rule’s. Filled dots are one standard error, open dots 1.96.

Put the two together and the trade is plain. At one standard error, at the two faster decays, the two-stage rule buys between 0.75 and 5.3 points of coverage for between 1.05 and 1.42 times the probes. At 1.96, at the same decays, it buys between minus 1.3 and plus 1.2 points — nothing the draws can distinguish from zero — for between 1.10 and 2.2 times the probes. At decay 0.97 the comparison is not about selection at all: the variance is small enough that the sequential rule’s warm-up forces probes the target did not ask for, and it covers 81 to 95 per cent with a warm-up of sixteen or thirty-two where the two-stage rule, which averages only what the target asks for, covers 67 to 69 — the table’s number. The two filled dots far below the line are those two settings: the two-stage rule is not worse there, it is correctly calibrated where the sequential rule over-covers by accident.

So the selection the proposal set out to remove is real, and it is small. It is worth a few points of coverage at one standard error and none at 1.96, and removing it costs a fifth or more of the probes.

What the pilot actually buys

The coverage is not the only thing the two-stage rule changes, and the other thing is arguably the more useful.

Its coverage is a theorem. Under a normal model of the samples zTAzz^{\mathsf T}Az, Stein’s procedure with Student’s margin has exactly the nominal coverage, for any p and any target; the measurements above are that theorem checked on samples that are not normal — a sum of Rademacher quadratic forms has a distribution with a finite, lattice-valued support — and holding within the draws’ noise. The sequential rule’s coverage has no such statement behind it. It is near the table on these matrices because its stops are late at the tight target, and a rule that reads only its own probes found it three times the target at the worst of forty draws. A code that must state a confidence — that must print “the trace is within 3% with probability 95%” and mean it — can print the two-stage rule’s without qualification and the sequential rule’s only with a footnote about how early it stopped. One draw in twenty made the same distinction for parameter-choice rules, whose median draws and worst draws answer different questions: a coverage is a statement about the whole distribution of draws, and only a rule whose sample size does not depend on its samples has one that can be derived rather than measured.

And its cost is known before the averaging starts. After the pilot, t is fixed; a code can report how many products it will spend and stop to ask whether that is affordable, which a sequential rule cannot. For a matrix whose products are expensive — a trace of a matrix function, each product an inner solve, as counting what is inside a circle computes — that is not a convenience. It is the difference between a budget and a hope.

Why the pooled mean is not the fix it looks like

A natural economy is to keep the pilot’s probes in the final mean, averaging all p + t. It is cheaper in coverage terms — the pilot’s samples add information — and on these matrices it over-covers: at decay 0.9, one standard error and the 3% target, 70.5 to 74.8 per cent against the fresh-only rule’s 64.8 to 68.0. That looks like a gain and is a loss of calibration in the other direction. The pooled mean’s variance is smaller than the one the margin was computed for, so the rule is conservative by an amount that depends on p and t, and the stated confidence is again not the achieved one. Stein’s result is for the fresh mean; the pooled one is safe, not calibrated. At the loose 10% target, where t is small, pooling covers 82 to 100 per cent at one standard error — the pilot dominates the mean.

Three matrices and Rademacher probes, four hundred draws each

Three matrices of one size, Rademacher probes, and 400 draws per setting, so each coverage carries about two points of noise at one standard error and one at 1.96. The two-stage rule is scored with its fresh mean. Stein’s exactness is for normal samples; the samples here are sums of products of ±1 entries, and the agreement with the table is measured rather than derived.

The sequential rule is scored with a warm-up equal to the pilot, which is the fair comparison of what each rule knows at the start and not the only one; the miss a normal table already priced measured the warm-up’s own effect. Hutch++'s deflated estimator, which a rate that belongs to the matrix found far better on fast-decaying spectra, is not given a two-stage version here; its stochastic part is a mean of samples like these, and the same construction applies to it.

Still open: a pilot shared across traces, and the contour

A pilot shared across related traces. The pilot is the two-stage rule’s overhead, and in a computation that estimates many traces of related operators — the log-determinants inside an optimisation — the spread changes slowly from one to the next. A pilot run once and carried forward would remove the overhead after the first trace, at the cost of the pilot’s spread being stale. Whether the coverage survives a drift in the operator, and how much drift a stale pilot tolerates before Student’s margin stops covering it, is the measurement the previous essay proposed for the sequential rule and which the two-stage rule makes cleaner, since its dependence on the spread is a single number.

A rule inside the contour estimator. Counting what is inside a circle estimates a trace whose exact value is an integer. A two-stage rule there has a target with no tolerance to choose — half an integer — and a pilot that could be spent on the quadrature as well as the probes. Whether the contour’s samples are close enough to independent for Student’s margin to hold is the measurement, and it is the one place in this field where the stated confidence would be about an exact answer.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Flop countHutchinson's estimatorMatrix-freeProbabilistic boundsRandom probeSpectral decayStopping criterionTrace estimation