Words moved against the block size, n = 48, M = 100
At its defaults it draws words moved against the block size, n = 48, m = 100. A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.
block-scan is one function in lib/figures/cost.js —
cost — where the flop count stopped predicting the time. Everything below came out of it during this build, at
arguments taken from the essays rather than invented for this page. A figure here is the
figure a reader meets in an essay, and if the generator changes, this page changes with it.
At its defaults
Drawn even though every essay passes arguments — which on this site is every essay, at 100% of placements since the standard pass. A default nothing exercises is a trap for the next essay to call this with none, and this is the page where a default that has drifted from the figures around it becomes visible.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.
fast: 100
The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.
fast: 100, n: 96
The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.
fast: 64, n: 96
The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 6, which is √M − 2. The three-line count says √(M/3) = 4, which has the right scaling and the wrong constant.
fast: 144, n: 96
The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 10, which is √M − 2. The three-line count says √(M/3) = 6, which has the right scaling and the wrong constant.
fast: 196, n: 96
The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.
A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 12, which is √M − 2. The three-line count says √(M/3) = 8, which has the right scaling and the wrong constant.
What it checked while drawing
Every figure above checked its own claims on the way to being drawn, and a claim that failed
would have stopped the picture rather than shipped a wrong one. Those checks used to leave
no trace at all: a passing one returned true and the only evidence the figure had
checked anything was that nothing crashed. The list below is what they actually said, collected
by running this generator with an observer installed — not a description of
what it is believed to check.
155 distinct claims across 6 sets of arguments, grouped below by shape — because most of them are one sentence with a different number in it, and how many separate times that sentence was put to the test is the informative part.
the reordered recursion is the same elimination on 96 columns at M = 144 — checked 96 times
a base case of 1 performs the same elimination — checked 12 times
a base case between one column and the whole matrix
a block chosen too large gives most of the saving back
a clear corner for the badge
a fast memory the edge sweep is drawn at
a fast memory the edge sweep measures: 64, 144 or 256 words
a fast memory the matrix does not fit in
a fast memory the width sweep measures: 64, 144 or 256 words
a matrix well outside the outer cache
a matrix well outside the outer cache, at a size the scan is affordable at
a matrix well outside the outermost cache, at a size the scan is affordable at
a middle cache can only make the best compromise worse
a reading of the block scan this generator draws
a size the recursion is affordable at
a size the scan and the recursion are both affordable at
a size the scan is affordable at
a size the ten scans are affordable at
a way of splitting a dimension the recursion knows
an outer cache larger than the inner one
and chooses the same pivots
and each level's best costs the other level
and far fewer than the unblocked elimination
and moves within a third of the best block's words without ever reading M
and neither is the largest
and one the matrix is well outside
and returns a factorisation at rounding
and the best block does not move when the problem grows
and the first panel past it falls off that cache's cliff
both cache legends fit on the axis label's line
caches each larger than the one inside it
caches from 16 to 1,024 words, each at least twice the one inside it
each wider panel removes calls
matmul shapes agree
panels no wider than √M − 2 for the innermost cache cost the recursion little there
the band under the curve holds the best block's label
the band under the line at 1 holds the two best-block labels
the block fits the matrix
the inner cache wants the smaller block
the optimum sits at √M − 2
the recursion is never more than 1% behind the block tuned to all three
the recursion performs exactly the unblocked elimination's operations
the recursion's column fits right of the scan
the recursion's label fits right of the scan
the smallest block is not the best
the three cache legends fit on the axis label's line
three caches, each at least twice the one inside it
two caches, the outer at least twice the inner
while the recursion is within 40% of each level's best
Against the rule
It draws a decomposition and prints its residual. It calls
blockedLU,
and every figure above carries the badge — which residualcheck verifies by looking
for it in the emitted SVG rather than by finding the call that builds one. A badge that is
constructed and then left out of the body is the failure that check exists for.
Across the library: the rule bites on 217
of 397 generators —
199 print a residual and
18 are exempt with a published reason;
180 factorise nothing.
Read from lib/residual-rule.js, which is the same body the gate enforces from,
and the gate's last check fails the build if this page and it disagree about any generator.
Where it is called
Changing this generator changes every figure on this list. That is what makes the list worth publishing rather than keeping in a check script.
A block size is a property of the machine
Three lines of counting say the best block size is √(M/3). Scanned over every integer at five fast memories, the measured optimum is √M − 2 — exactly, at all five. The count has the right scaling and the wrong constant, low by a factor of 1.56, and the wrong form: the answer is affine in √M rather than proportional to it.
Where the flop count stopped predicting the timeLeaves cut to the edge on purpose
A recursive elimination's cheapest leaf is exactly the square root of M, less 2, columns wide, and the rule drawn from it was to set the base case there and let the halvings put the leaves at or below it. On thirty-two sizes from 96 to 127 columns, halving to that base case moves more words than the pure recursion on sixteen of them at 144 words of fast memory, because the halvings stop at six and seven columns, not ten. Cut every dimension at a multiple of the edge instead and every size keeps the saving: 0.79 to 0.86 of the pure recursion's words, against halving's 0.91 to 1.07, with the same arithmetic and the same pivots. What it cannot make full is the one leftover leaf, and a leftover of one column is where it loses.
Where the flop count stopped predicting the timeThe block size a recursion still has
A recursive elimination is sold as having no block size, and every real one switches to plain loops below some width. Swept over that width, the traffic is a staircase with its steps at the halvings of n, and its cliff sits where the blocked elimination's does — the first panel wider than √M − 2 moves 1.53 to 4.47 times the words, on six memories of six. On three caches the innermost decides, and a third cache costs every tuned block up to 14 per cent and the recursion nothing.
Where the flop count stopped predicting the timeThe leaf that sits on the edge
A recursive elimination's base case was found to be a block size in disguise, with a cliff where the blocked elimination's is, and the choice read as a trade: the processor wants wide leaves, the cache wants narrow ones. Measured at every width rather than at the halvings of 96, there is no trade inside the edge. Leaves exactly √M − 2 wide are the cheapest the recursion can have in words as well as calls — 0.81, 0.80 and 0.84 of the pure recursion's traffic at 64, 144 and 256 words — and one column wider moves 1.77 to 3.15 times it. And a matrix of 100 columns, which halves unevenly, meets the cliff in two steps rather than one.
Where the flop count stopped predicting the timeThe recursion that was never told the memory
A blocked elimination has to be tuned to its fast memory, and tuned to one memory it costs up to 2.9 times the best at another. A recursive elimination splits the columns in half down to one and reads no memory size at all. On eight fast memories from 36 to 576 words it moves between 0.94 and 1.28 times the words of the best tuned block, with the same 585,200 operations and the same pivots — and on a machine with two caches it beats the block tuned to either cache on six machines of seven.
Where the flop count stopped predicting the timeThe same arithmetic at a different price
A blocked and an unblocked elimination perform 72,568 operations each — the same operations, associated differently — choose the same pivots, and return a factorisation identical to the last bit: ‖PA − LU‖/‖A‖ = 4.487946226420872·10⁻¹⁶ in both. One of them moves 41,332 words between fast and slow memory and the other moves 19,476.