A drift put back at the table's price
Worth reading first: Counting what cannot be looked at · A parameter that counts steps.
A spread carried from the trace before ran Stein’s two-stage trace estimator along sixteen operators whose spectrum drifts, and found the cheapest honest way to know its spread: take it from the previous trace’s own averaged probes. Those probes are already paid for, they are independent of this trace’s, and they are one step stale. The staleness was the whole of the loss — two or three points of coverage at one standard error at the fastest drift — and the essay named the repair. “The drift is visible: two consecutive traces’ spreads give its rate. Multiplying the carried spread by the ratio of the last two — extrapolating one step — would remove the loss to first order.” Its prediction had a sign on each side: at the fastest drift the sixteenth trace’s coverage would rise from 64.5 per cent to within a point of the table, and on a sequence that does not drift the repair would cost no coverage but inflate the probe counts by the noise of a ratio of two estimated spreads.
The figure above is the first half of that, measured in the way it can be measured, and the repair works. The rest of the essay is about the second half, which does not go the way it was predicted, and about what the repair costs, which is the part that decides whether to use it.
The same sequences, four ways to know the spread
The operators are the earlier essay’s: sixty by sixty, fixed eigenvectors, eigenvalues for , the decay starting at 0.9 and falling by a fixed step each trace — 0, 0.0025, 0.005, 0.01 or 0.02 — so that at the fastest drift the decay reaches 0.6 by the sixteenth operator and the relative spread of one Rademacher probe grows from 0.264 to 0.670, by about 6.4 per cent a trace. Each trace is a two-stage estimate with a target of 3 per cent at one standard error: a spread, from somewhere, fixes how many fresh probes to average, and the average is scored inside or outside the target against the exact trace. Four places to get the spread:
- fresh, a pilot of eight probes for every trace, discarded after use, which a spread measured on probes it does not average showed covers what Student’s distribution says;
- the plain carry, the previous trace’s averaged probes, the earlier essay’s rule;
- the trend carry, the same spread multiplied by the ratio of the last two carried spreads, the proposal;
- a four-trace trend, the same spread multiplied by where is the slope of a least-squares line through the logarithms of the last four carried spreads — the obvious way to quiet a ratio of two noisy numbers.
The first trace of every carry needs a pilot, the second has one carried spread and no ratio, so the carries differ only from the third trace on. Everything is over 400 seeded sequences.
A single trace is noisier than the effect
The prediction was stated about one number, the sixteenth trace’s coverage at the fastest drift, and that is where to start, because it cannot carry the weight. A share of 400 draws has a standard error of 2.3 points at a coverage near 68 per cent. The earlier essay measured the plain carry’s sixteenth trace at 64.5 per cent; run again here, with the same rule on the same operators but its random probes drawn in a different order because four rules now share each sequence’s stream, it reads 66.0. The trend carry’s sixteenth trace reads 72.3 — four points above the table, a better result than was predicted and no more believable for it.
Every value in the figure is inside two standard errors of the table, both rules, every drift. A single trace’s coverage cannot tell these rules apart, and the honest reading of the earlier essay’s 64.5 per cent is that it was a draw from a distribution centred perhaps two points below the table. The measurement that can see a difference of two or three points is the average over many traces.
Averaged, the drift comes back
Trace by trace at the fastest drift the plain carry sits below the table from the third trace to the sixteenth, thirteen traces of fourteen under it, and the trend carry scatters around it. Averaged over traces three to sixteen — 5,600 estimates — the plain carry covers 65.6 per cent and the trend carry 68.9, where the table says 68.3. The first figure of this essay shows the same average at every drift. The plain carry is on the table without drift (68.0) and falls as the drift grows, to 65.4 at 0.01 and 65.6 at 0.02. The trend carry stays between 67.8 and 68.9 at every drift.
Turn the dial to no drift and the two rules are one rule with a noisy factor, both inside the band. That is the prediction’s first half holding, measured on the quantity that can show it: the trend carry restores the table at the fastest drift, to within six tenths of a point.
The four-trace trend does the same on average — 68.7 at the fastest drift — but it over-covers without drift, 70.2, two points above the table, which a quieter estimate of a zero trend should not do. It is not nearer the table than the two-trace ratio at every drift, which was the reason to fit a line in the first place.
The bias is gone and the noise is not
Why it works is visible in the spread each rule actually used. The plain carry’s is low by about one step of drift — 0.96 of the true spread at the sixteenth trace, having climbed there from 0.86 at the second, where it was still carrying the pilot’s downward-biased estimate on seven degrees of freedom. The trend carry’s median is 0.98 to 1.01 from the third trace on, and the four-trace trend’s about 1.00 to 1.04. One ratio removes the lag.
What it adds is scatter. A ratio of two estimated spreads is noisier than either, and the trend carry’s spread scatters with an interquartile range of 0.27 of its median without drift, against the plain carry’s 0.105 — two and a half times as wide. The four-trace line is in between at 0.154. All three narrow as the drift grows, from 0.27 to 0.12 for the trend carry, because faster drift means harder operators, more probes per trace and so more degrees of freedom in every carried spread: the sixteenth trace averages about 75 probes without drift and about 500 at 0.02, as and say it should.
The noise costs nothing
The second half of the prediction was that this noise would cost probes where there is no drift to correct: a spread that is sometimes high and sometimes low asks for probes, and a square is convex, so a symmetric scatter in should raise the average count. Without drift the trend carry spends 1,225 products for the sixteen traces and the plain carry 1,246. The noisier rule is cheaper, by 1.7 per cent, and covers 68.1 per cent against 68.0.
The convexity argument assumes the scatter is symmetric in around the right value. A ratio of two spreads with similar degrees of freedom is closer to symmetric in its logarithm, which puts its median below its mean and its typical value slightly low; the median spread the trend carry uses without drift is 0.974 of the true one against the plain carry’s 0.992. A slightly low typical spread spends slightly fewer probes, and the occasional high one spends more, and on this sequence they net to a saving. The coverage cost a noisy margin should have — coverage is concave in the margin — is real and too small to see: a scatter of this width in the margin is worth a few tenths of a point at one standard error, inside the average’s own noise of about six tenths.
So the factor’s noise is not the price of the repair. The price is somewhere else.
It costs what the table charges
At the fastest drift the trend carry spends 4,412 products for the sixteen traces, the plain carry 3,994, and a fresh pilot every trace 4,377. The repair costs 10.5 per cent more than the rule it repairs, and it costs more than re-measuring from scratch.
That is not waste, and the normal table says how much it should be. The plain carry covers 65.6 per cent, which is what a normal estimate covers at a margin of 0.946 standard errors instead of 1; to cover 68.3 it would need a margin of 1, and probes go as the square of the margin, so the price of the gap is — 11.9 per cent more products. The trend carry pays 10.5. At every drift it pays no more than that price: 1.5 per cent at 0.0025 against a price of 4.0, 3.4 against 3.8 at 0.005, 7.2 against 12.7 at 0.01.
So the plain carry’s saving over a fresh pilot, which the earlier essay counted at seven per cent at this drift, was partly not a saving. It was coverage the rule was not buying, and buying it back costs about what it saved. The earlier essay said this of the frozen rule — “the cheapest on every drift, and the reason is that it is wrong” — and the same sentence applies, in a smaller way, to the rule that replaced it. What the trend carry offers is aim: the margin that is stated is the margin that is delivered, at whatever drift, for the probes that margin actually requires.
What one ratio assumes
The trend carry is a one-line change to the plain carry, and the line encodes a model: the relative spread of one probe changes by a constant factor per trace. On these sequences that is nearly true by construction — the decay moves by a fixed step, and the relative spread of a Rademacher probe, which is proportional to the Frobenius norm of the off-diagonal part over the trace, responds smoothly — so the measured success is a statement about operators that drift smoothly, not about drift in general.
What makes the model cheap to trust is that it fails visibly. A factor far from one is evidence of a change the plain carry would also have mishandled. On these sequences the true one-step ratio averages 1.064 at the fastest drift, and the spread the factor produces has a median within two per cent of the true spread from the third trace on, though any one draw’s factor scatters widely around it. A computation whose operators drift smoothly — the log-determinants inside an optimisation, which motivated the carried spread in the first place, change by a step whose size the optimiser controls — is the case this model describes. A factorisation kept past its date found the same structure for a Cholesky factor reused along a drifting sequence: what can be carried depends on how fast the operator moves per step, and a correction that knows the rate can carry further than one that does not. The trace estimator has the advantage that its rate is measured for free, by the spreads it was going to compute anyway.
Where this leaves the carried spread
The fresh pilot still has something the carries do not: it needs no model of how the operator changes. It is not exact either — averaged over the same traces it covers 67.7 per cent without drift and 66.5 at the fastest, a point or two under, for reasons the earlier essays traced to the pilot’s seven degrees of freedom meeting a spread that is not normally distributed — but its error does not depend on the sequence. The trend carry assumes the relative spread changes by a constant factor from one trace to the next, which on these sequences it nearly does; on a sequence whose drift reverses, or jumps, the ratio of the last two traces is the wrong extrapolation for exactly one trace after each turn. The plain carry assumes only that one step is small. The order of preference that comes out of the measurements is therefore by what is known about the sequence. With no drift, carry plainly: the trend adds noise for nothing. With a steady drift, carry with the ratio: the table’s coverage at the table’s price. With drift of unknown shape, a fresh pilot every trace is the rule that cannot be fooled, and at the fastest drift here it is no more expensive than the trend carry and within two points of the table.
The same accounting runs through the field. The miss a normal table already priced found the margin’s price to be its square when the estimator is free to choose its own stopping point; a rule that reads only its own probes found the cost of a target to be the estimator’s own square root. Here the same square appears as the price of coverage that drift had quietly stopped buying, and it is paid in full.
What the four-trace line was for
A longer window is the standard answer to a noisy ratio, and it halves the scatter here. It costs one more thing the measurements show: a lag of its own. The line through four spreads estimates the average drift over the last three steps, which on a steady drift is the drift, and it starts later — it has a full window only from the fifth trace. Without drift it over-covers by two points, for a reason these measurements do not isolate; whatever it is shows where there is nothing to correct, which is the one place a correction should be invisible. Its products run 1.8 to 3.1 per cent above the trend carry’s at every drift. Nothing here argues for it over the ratio, which is cheaper, covers as well and needs no window.
Counting what cannot be looked at began this field with Hutchinson’s estimator and its variance; a rate that belongs to the matrix and the split nobody is in a position to choose made the estimator’s cost a property of the spectrum. Everything since has been about knowing that cost honestly, and the result here is that knowing it under drift is cheap to compute and costs exactly what it should.
What these sequences do not show
One family of operators, a drift that is steady by construction, one target and one margin. A drift that changes speed, reverses or jumps is the case the trend carry is built not to handle, and it is unmeasured; so is a carried sketch for Hutch++, and so are margins other than one standard error, where the earlier essay found the same shape at 1.96 but the trend carry has not been run.
Still open: a drift that turns, and the contour
A drift that turns. A sequence whose decay falls for eight traces and then rises would make the ratio of the last two spreads the wrong extrapolation for one trace at the turn. The prediction with a sign is that the trend carry’s coverage dips at the trace after the turn by about twice the plain carry’s lag — four or five points — and recovers within one trace, so that its average over the sequence stays within a point of the table; and that a guard which caps the factor at the ratio of the last three spreads removes the dip for no measurable cost.
The contour. Counting what is inside a circle estimates a trace whose exact value is an integer. A spectrum drifting across the contour changes that integer between traces, and both the plain and the trend carry would carry a spread across the jump. Whether a jump is visible in the ratio — a factor far from one — and could trigger a fresh pilot is where this essay’s factor would become a detector rather than a correction.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rank a certificate charges — both name probabilistic bounds, random probe, spectral decay, stopping criterion
- A block nobody can call sparse — both name matrix-free, spectral decay
- A sketch that finds the columns it can see — both name probabilistic bounds, spectral decay
- Columns that steer a random start — both name probabilistic bounds, random probe
- The leverage that did not move — both name probabilistic bounds, spectral decay
Named objects
A flat tag is an object no other essay names yet.
Hutchinson's estimatorMatrix-freeProbabilistic boundsRandom probeSpectral decayStopping criterionTrace estimation