The miss a normal table already priced
Worth reading first: Counting what cannot be looked at · A parameter that counts steps.
A rule that reads only its own probes stopped Hutchinson’s estimator when the standard error of its own running mean, divided by the mean, reached a target. It found the rule calibrated in the median and not in the tail, and it ordered the two ways a run can miss. The larger was the one any rule built on one standard error must have: the error of a mean exceeds one standard error on about a third of draws. The smaller was specific to estimating the spread from the same few samples, and it placed that one at the loose targets, where the rule stops after eight probes on a standard deviation it has barely measured. “So the rule is worst calibrated exactly where it is cheapest, and the loose targets are the ones to distrust.”
That was read off forty draws a cell, which the same essay said could not tell 70% from 80%. It left the margin and the warm-up unswept. Swept, with 400 draws a cell on the same 60×60 matrices, the ordering turns over: the loose target is the better covered, and the reason is the warm-up doing the opposite of what it was blamed for.
Where the criterion decides, a normal table is the calibration
The cleanest cell is a tight target on a spectrum that decays quickly enough for the rule to need many probes. At a 3% target and a decay of 0.9 per eigenvalue, the median run takes 76 probes and one run of 400 stops on the eighth probe, the first it is allowed to. The criterion decides when the other 399 end.
There the rule covers 63.7% of draws with no margin, 90.3% with a margin of 1.645, 94.8% at 1.96 and 98.5% at 2.576. A normal table says 68.3%, 90.0%, 95.0% and 99.0%. At a decay of 0.8, where the median run takes 202 probes, the same four margins give 66.0%, 88.0%, 94.3% and 97.5%. Nothing about the rule was fitted to those numbers; they are what falls out of dividing a target by a constant and counting.
The dial moves the spectrum, and the three settings are the three regimes of this essay. At 0.8 the rule needs hundreds of probes at the tight target and dozens at the loose one, the warm-up rarely binds, and both lines sit on the curve. At 0.9 the tight line stays on the curve and the loose line leaves it upward at the left, where no margin is applied and the warm-up stops two runs in three. At 0.97 the variance is small against the trace, the warm-up stops every run at the loose target and half of them at the tight one. The loose line sits at 99.5% whatever the margin, because eight probes already exceed every margin’s request, and the tight line starts three points above the curve at c = 1 and falls a point or two under it once the margin asks for more than eight probes and the criterion takes over again. The curve is the same in every frame, because it is a statement about means of independent samples and nothing else, and the matrix decides only how many of the rule’s stops the criterion made.
The miss at one standard error is therefore not a defect with a remedy. It is the normal distribution’s third, and the rule delivers it within a few points. The few points are nearly all on one side — the rule covers a little less than the table in seven of the eight cells where the warm-up does not bind, by 2 to 5 points at c = 1 and by under one point at 1.96, and the eighth is 0.3 points over — and that has a cause worth naming. A run that stops is a run whose estimated standard error just fell below a line, and among the ways to fall below it is for the estimate of the spread to dip. The rule preferentially stops on streams that are temporarily underestimating their own noise. That is a real bias, it is what the preceding essay’s “second failure mode” was reaching for, and it costs about three points of coverage at one standard error, not the tail.
What a margin costs
The obvious objection to a margin is that it is paid for, and the question is at what rate.
The standard error of a mean of t independent samples is their standard deviation over , so asking for a standard error c times smaller asks for times the samples. At a 3% target the measured ratios are 2.68, 3.82 and 6.61 at a decay of 0.8 and 2.74, 3.86 and 6.71 at 0.9, against squares of 2.71, 3.84 and 6.64. The margin costs its square to within 1.4%, at every margin, on both spectra.
The square is the estimator’s own rate read backwards. A rate that belongs to the matrix fitted Hutchinson’s error against its budget and found an exponent near a half on every spectrum it tried, which is the same statement: error falls as the inverse square root of the probes, so probes rise as the inverse square of the error asked for. The margin is simply a smaller error asked for. What the spectrum sets is the constant in front — 202 probes against 76 at the same target, for the two decays — and the margin multiplies whatever that constant is.
So the price list a user needs is a normal table and a square. A 95% promise costs 3.84 times a one-standard-error promise, and a 99% promise costs 6.6 times. The rule’s first measurement said a margin of two would put 95% inside where one puts 80%, at four times the probes. The four is right; the 80 was the 1% row of a forty-draw table, and the one-standard-error coverage where the criterion decides is nearer 64 to 66.
The third line on that figure is where the arithmetic stops working, and it is the thread the rest of this essay pulls. At a 10% target and a decay of 0.9, the margin’s cost falls short of its square: 2.13, 3.13 and 5.50 times, against 2.71, 3.84 and 6.64. The first count, at no margin, is eight probes, and eight is not what the criterion asked for. It is what the warm-up allowed.
The warm-up stops the loose target, and pays for its coverage
At a loose target the criterion is often satisfied before the rule is permitted to look at it. At a 10% target and a decay of 0.9, 66% of runs stop on the eighth probe exactly — the first the warm-up allows — which means that by then their standard error was already below the target, and on many of them it had been for several probes.
A run the warm-up stops has more probes than its target needed, and so a smaller error than its target allowed. That is why the loose target is covered better, not worse. At a decay of 0.9, 75.8% of draws are inside a 10% target and 63.7% inside a 3% one. At 0.97, where the spectrum is flat enough that the variance is small against the trace, the warm-up stops every run at the loose target and 99.5% of them are inside it; the tight target, which the warm-up stops on 47% of runs, is at 71.3%. At 0.8 the warm-up stops only 14% of the loose runs and the two targets are covered equally, at 66.0%. The ordering follows the share the warm-up stopped, on all three decays, and it never runs the way the forty-draw table read it.
The preceding essay’s second failure mode — a stream whose first eight probes happen to agree, reporting a small standard error and stopping on it — is real. It is outweighed. A run the warm-up stops has usually overshot its target by enough that an underestimated standard error still leaves it inside, and the ones the underestimate matters for are the runs that stop near their target, which at a loose target are the minority.
The spread the early runs report is too small, and it does not matter yet
The standard error each run reports can be scored directly, because the trace is known: divide the run’s actual relative error by the relative standard error it stopped with. If the reported standard error is honest, that ratio is the size of a standard normal variable, whose median is 0.674 and whose 95th percentile is 1.96.
Runs stopped by the criterion, at a 3% target with a margin of 1.96 and a decay of 0.9, lie on the diagonal: median 0.67, 95th percentile 1.99. Their reported standard errors are honest. Runs stopped by the warm-up, at a 10% target and a decay of 0.97, lie on it to the median — 0.69 — and then climb above it, to a 95th percentile of 2.36 and a largest value near four. Those are the streams whose eight probes agreed by chance: their reported spread is too small by about a sixth at the 95th percentile.
That is the preceding essay’s mechanism, measured, and it is the right size for what eight samples can know about a standard deviation. What it does not do is cost coverage, because those runs stopped on the warm-up, which means their standard error had fallen below the target before the rule was allowed to look, often well below it. An error of 2.36 standard errors, where a standard error is a small fraction of the target, is still inside it. The under-reporting is a fact about the runs’ statements of their own precision, and it becomes a fact about their accuracy only on a matrix where the warm-up stops runs close to their target — which is the tight target on a flat spectrum, the 71.3% cell, and not the loose target at all.
What a longer warm-up buys
A longer warm-up is the other knob, and at a loose target it acts as a budget rather than as better statistics.
At a decay of 0.9 and a 10% target, a warm-up of 4, 8, 16 and 30 probes gives 67.8%, 75.8%, 89.0% and 97.3%, and the median run takes 6, 8, 16 and 30 probes — that is, from eight on, exactly the warm-up. At 0.97 the coverage is 96.8% already at four probes and 100% from sixteen. At 0.8, where most runs need more probes than any of the warm-ups, it rises only from 61.5% to 83.0%, and the median cost from 15 to 30.
A warm-up of thirty buys 97% coverage at a 10% target on the middle spectrum, and it does so by spending thirty probes where the criterion wanted six. The same thirty probes, spent through a margin instead, sit between the 25 that a margin of 1.96 costs at that target and the 44 that 2.576 costs — a margin of about 2.2, which a normal table reads as a 97% promise — and would have bought it deliberately, stated in advance, and at every decay. The warm-up buys the same coverage only on the spectra where the criterion would have stopped earlier, which is to say it buys coverage exactly where the rule was already cheap, and says nothing about it.
With the margin at 1.96, the warm-up hardly matters. Across three decays and two targets, a warm-up of four covers 90.8% to 98.0%, a warm-up of eight 92.8% to 99.5%, and sixteen or more 94.3% or better. Four knobs and one floor asked of every rule scored against an oracle which knob decides its result. Here it is the margin, and the warm-up’s job is only to stop the rule reading a standard deviation from two samples.
The early stops, taken alone
There is one place where the preceding essay’s mechanism does cost coverage, and it is worth stating so the picture is not tidier than the measurement. With a warm-up of four, the runs that stop within eight probes are the ones whose first few samples agreed, and they are inside the target less often than the rest. At a decay of 0.8 and a 10% target with no margin, 103 of 400 runs stop that early and 49% of them are inside; with a margin of 1.96, eleven stop that early and three are inside. With a warm-up of thirty, no run can stop that early, and the early-stop penalty disappears.
So the mechanism exists, it is confined to warm-ups shorter than eight, and it is the reason eight rather than four is the smallest warm-up worth having. At eight and above, the loose target’s supposed fragility is covered by the probes the warm-up forces.
Against the other rules in this field
A stopping test is a race described every stopping rule in this collection as a proxy reaching a threshold before the quantity it stands for. This rule’s proxy — a standard error — stands for the error up to a distribution, and the distribution turns out to be the normal one to within a few points. That is the unusual position the preceding essay claimed for it, confirmed with a table: the rule’s promise can be stated as a coverage and priced as a square, which no residual-based test in the iterative field can offer.
It also sits beside the rank a certificate charges, where a randomised range finder certified its own rank with ten Gaussian probes and a safety factor of ten, and never stopped early in any draw. There the price of never missing was nineteen columns beyond the answer’s rank. Here the price of a 95% promise is 3.84 times the probes of a 68% one, and of a 99% promise 6.6 times. A certificate that is never wrong and a margin that is wrong one time in twenty are different contracts, and this rule can offer either at a known cost; the range finder’s constant cannot be turned down.
The constant is also where deflation acts. The split nobody is in a position to choose measured Hutch++ spending part of its budget on a sketch that removes the spectrum’s head, which lowers the variance of every remaining probe. A margin applied to the deflated remainder costs the same square, of a smaller count, so the two are independent choices: deflation lowers the price of a probe, and the margin decides how many standard errors of promise to buy with them.
And the whole of this is a second instance of the contract a bound that holds with probability introduced for randomised low-rank approximation: an answer that is right with a stated probability rather than always. There the probability came from a theorem and arrived with constants nobody tunes. Here it comes from a normal table and arrives with a price anyone can compute, and a user can move along it.
The estimator’s variance depends on which random vector it uses, and every probe here is ±1. The coverage should not move under Gaussian probes, since the argument is about means of independent samples and not about their distribution; the cost column would, by the ratio of the two variances.
What this does not settle
Three geometric spectra on one size, 400 draws a cell. A binomial share near 0.95 has a standard error of about 1.1 points at 400 draws, so the gaps of under one point between the rule and the table at c = 1.96 are not measured; the gaps of 2 to 5 points at c = 1 are.
The sequential-stopping bias — the rule covering a little less than a normal table because it stops when its spread is underestimated — is inferred from the sign of the gap and from the early stops, not isolated. A rule that estimated the spread from one stream and the mean from another would remove it, and would say how much of the gap it is.
The z-score figure compares two cells chosen to show the two regimes, not a sweep. Where the transition between them sits, as a function of how far below the target a run’s standard error is when the warm-up releases it, is not measured.
Still open: a rule for the contour, and a spread carried between traces
A rule inside the contour estimator. Counting what is inside a circle integrates a trace around a contour, and its answer is an integer. A rule there could stop when the estimate is within half an integer with a stated margin — a normal table gives the margin, a square gives its price, and the target has no tolerance to choose. Whether the contour’s samples are close enough to independent for the table to hold is the measurement.
A spread carried between traces. Every run here estimates its spread from scratch, which is where both the early-stop penalty and the sequential bias come from. A computation that estimates many traces of related operators — the log-determinants inside an optimisation — could carry the spread from one to the next, stop without a warm-up and state its margin from the start. Whether the carried spread is close enough to the current one for the coverage to survive a drift in the operator is unmeasured.
Separating the two uses of the samples. A rule that estimated the spread from a separate, small set of probes and the mean from the rest would not select streams that underestimate their own noise. It would cost the separate probes. Whether the three points of coverage it recovers at one standard error are worth that is a measurement this essay’s sweep can already make.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A sketch that finds the columns it can see — both name flop count, probabilistic bounds, spectral decay
- A straight path has nothing for a parabola to fit — both name exact ground truth, flop count, stopping criterion
- A test with no tolerance in it — both name exact ground truth, flop count, stopping criterion
- One line that buys a quarter of the run — both name exact ground truth, flop count, stopping criterion
- The degree the history chooses — both name exact ground truth, flop count, stopping criterion
- The vector was what was wanted — both name exact ground truth, flop count, matrix-free
Named objects
A flat tag is an object no other essay names yet.
Exact ground truthFlop countHutchinson's estimatorMatrix-freeProbabilistic boundsRandom probeSpectral decayStopping criterionTrace estimation