Generator

Words moved against the block size, n = 48, M = 100

One function in the cost library, called 36 times across 6 essays. Below: what it draws at its defaults, what it draws at every value an essay asks for, the 155 claims it put to the test while drawing them, and where it stands against the rule this site is named for.

At its defaults it draws words moved against the block size, n = 48, m = 100. A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.

block-scan is one function in lib/figures/cost.js — cost — where the flop count stopped predicting the time. Everything below came out of it during this build, at arguments taken from the essays rather than invented for this page. A figure here is the figure a reader meets in an essay, and if the generator changes, this page changes with it.

At its defaults

Drawn even though every essay passes arguments — which on this site is every essay, at 100% of placements since the standard pass. A default nothing exercises is a trap for the next essay to call this with none, and this is the page where a default that has drifted from the figures around it becomes visible.

Words moved against the block size, n = 48, M = 100A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.110¹10⁴10⁵block size bwords movedthe count: √(M/3) = 5measured best: b = 8words movedat the best block1.6·10⁴at b = 13.9·10⁴at b = 243.9·10⁴derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.

fast: 100

The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.

Words moved against the block size, n = 48, M = 100A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.110¹10⁴10⁵block size bwords movedthe count: √(M/3) = 5measured best: b = 8words movedat the best block1.6·10⁴at b = 13.9·10⁴at b = 243.9·10⁴derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.

fast: 100, n: 96

The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.

Words moved against the block size, n = 96, M = 100A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.110¹10²10⁵block size bwords movedthe count: √(M/3) = 5measured best: b = 8words movedat the best block1.2·10⁵at b = 15.5·10⁵at b = 243.2·10⁵derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 8, which is √M − 2. The three-line count says √(M/3) = 5, which has the right scaling and the wrong constant.

fast: 64, n: 96

The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.

Words moved against the block size, n = 96, M = 64A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 6, which is √M − 2. The three-line count says √(M/3) = 4, which has the right scaling and the wrong constant.110¹10²10⁵10⁶block size bwords movedthe count: √(M/3) = 4measured best: b = 6words movedat the best block1.6·10⁵at b = 15.9·10⁵at b = 243.2·10⁵derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 6, which is √M − 2. The three-line count says √(M/3) = 4, which has the right scaling and the wrong constant.

fast: 144, n: 96

The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.

Words moved against the block size, n = 96, M = 144A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 10, which is √M − 2. The three-line count says √(M/3) = 6, which has the right scaling and the wrong constant.110¹10²10⁵block size bwords movedthe count: √(M/3) = 6measured best: b = 10words movedat the best block1.1·10⁵at b = 14.8·10⁵at b = 243.2·10⁵derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 10, which is √M − 2. The three-line count says √(M/3) = 6, which has the right scaling and the wrong constant.

fast: 196, n: 96

The arguments are the ones A block size is a property of the machine passes. A value nobody placed would be a picture no essay asked for and no claim was ever checked against.

Words moved against the block size, n = 96, M = 196A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 12, which is √M − 2. The three-line count says √(M/3) = 8, which has the right scaling and the wrong constant.110¹10²10⁵block size bwords movedthe count: √(M/3) = 8measured best: b = 12words movedat the best block9.5·10⁴at b = 13·10⁵at b = 243.2·10⁵derived from M with no measurement, and scannedthe two agree

A U-shaped curve on logarithmic axes. Too small a block refactors the panel too often; too large a block does not fit in the fast memory and the tiled update evicts what it is about to read. The minimum is at b = 12, which is √M − 2. The three-line count says √(M/3) = 8, which has the right scaling and the wrong constant.

What it checked while drawing

Every figure above checked its own claims on the way to being drawn, and a claim that failed would have stopped the picture rather than shipped a wrong one. Those checks used to leave no trace at all: a passing one returned true and the only evidence the figure had checked anything was that nothing crashed. The list below is what they actually said, collected by running this generator with an observer installed — not a description of what it is believed to check.

155 distinct claims across 6 sets of arguments, grouped below by shape — because most of them are one sentence with a different number in it, and how many separate times that sentence was put to the test is the informative part.

the reordered recursion is the same elimination on 96 columns at M = 144 — checked 96 times

a base case of 1 performs the same elimination — checked 12 times

a base case between one column and the whole matrix

a block chosen too large gives most of the saving back

a clear corner for the badge

a fast memory the edge sweep is drawn at

a fast memory the edge sweep measures: 64, 144 or 256 words

a fast memory the matrix does not fit in

a fast memory the width sweep measures: 64, 144 or 256 words

a matrix well outside the outer cache

a matrix well outside the outer cache, at a size the scan is affordable at

a matrix well outside the outermost cache, at a size the scan is affordable at

a middle cache can only make the best compromise worse

a reading of the block scan this generator draws

a size the recursion is affordable at

a size the scan and the recursion are both affordable at

a size the scan is affordable at

a size the ten scans are affordable at

a way of splitting a dimension the recursion knows

an outer cache larger than the inner one

and chooses the same pivots

and each level's best costs the other level

and far fewer than the unblocked elimination

and moves within a third of the best block's words without ever reading M

and neither is the largest

and one the matrix is well outside

and returns a factorisation at rounding

and the best block does not move when the problem grows

and the first panel past it falls off that cache's cliff

both cache legends fit on the axis label's line

caches each larger than the one inside it

caches from 16 to 1,024 words, each at least twice the one inside it

each wider panel removes calls

matmul shapes agree

panels no wider than √M − 2 for the innermost cache cost the recursion little there

the band under the curve holds the best block's label

the band under the line at 1 holds the two best-block labels

the block fits the matrix

the inner cache wants the smaller block

the optimum sits at √M − 2

the recursion is never more than 1% behind the block tuned to all three

the recursion performs exactly the unblocked elimination's operations

the recursion's column fits right of the scan

the recursion's label fits right of the scan

the smallest block is not the best

the three cache legends fit on the axis label's line

three caches, each at least twice the one inside it

two caches, the outer at least twice the inner

while the recursion is within 40% of each level's best

Against the rule

It draws a decomposition and prints its residual. It calls blockedLU, and every figure above carries the badge — which residualcheck verifies by looking for it in the emitted SVG rather than by finding the call that builds one. A badge that is constructed and then left out of the body is the failure that check exists for.

Across the library: the rule bites on 217 of 397 generators — 199 print a residual and 18 are exempt with a published reason; 180 factorise nothing. Read from lib/residual-rule.js, which is the same body the gate enforces from, and the gate's last check fails the build if this page and it disagree about any generator.

Where it is called

Changing this generator changes every figure on this list. That is what makes the list worth publishing rather than keeping in a check script.

Where the flop count stopped predicting the time

A block size is a property of the machine

Three lines of counting say the best block size is √(M/3). Scanned over every integer at five fast memories, the measured optimum is √M − 2 — exactly, at all five. The count has the right scaling and the wrong constant, low by a factor of 1.56, and the wrong form: the answer is affine in √M rather than proportional to it.

Where the flop count stopped predicting the time

Leaves cut to the edge on purpose

A recursive elimination's cheapest leaf is exactly the square root of M, less 2, columns wide, and the rule drawn from it was to set the base case there and let the halvings put the leaves at or below it. On thirty-two sizes from 96 to 127 columns, halving to that base case moves more words than the pure recursion on sixteen of them at 144 words of fast memory, because the halvings stop at six and seven columns, not ten. Cut every dimension at a multiple of the edge instead and every size keeps the saving: 0.79 to 0.86 of the pure recursion's words, against halving's 0.91 to 1.07, with the same arithmetic and the same pivots. What it cannot make full is the one leftover leaf, and a leftover of one column is where it loses.

Where the flop count stopped predicting the time

The block size a recursion still has

A recursive elimination is sold as having no block size, and every real one switches to plain loops below some width. Swept over that width, the traffic is a staircase with its steps at the halvings of n, and its cliff sits where the blocked elimination's does — the first panel wider than √M − 2 moves 1.53 to 4.47 times the words, on six memories of six. On three caches the innermost decides, and a third cache costs every tuned block up to 14 per cent and the recursion nothing.

Where the flop count stopped predicting the time

The leaf that sits on the edge

A recursive elimination's base case was found to be a block size in disguise, with a cliff where the blocked elimination's is, and the choice read as a trade: the processor wants wide leaves, the cache wants narrow ones. Measured at every width rather than at the halvings of 96, there is no trade inside the edge. Leaves exactly √M − 2 wide are the cheapest the recursion can have in words as well as calls — 0.81, 0.80 and 0.84 of the pure recursion's traffic at 64, 144 and 256 words — and one column wider moves 1.77 to 3.15 times it. And a matrix of 100 columns, which halves unevenly, meets the cliff in two steps rather than one.

Where the flop count stopped predicting the time

The recursion that was never told the memory

A blocked elimination has to be tuned to its fast memory, and tuned to one memory it costs up to 2.9 times the best at another. A recursive elimination splits the columns in half down to one and reads no memory size at all. On eight fast memories from 36 to 576 words it moves between 0.94 and 1.28 times the words of the best tuned block, with the same 585,200 operations and the same pivots — and on a machine with two caches it beats the block tuned to either cache on six machines of seven.

Where the flop count stopped predicting the time

The same arithmetic at a different price

A blocked and an unblocked elimination perform 72,568 operations each — the same operations, associated differently — choose the same pivots, and return a factorisation identical to the last bit: ‖PA − LU‖/‖A‖ = 4.487946226420872·10⁻¹⁶ in both. One of them moves 41,332 words between fast and slow memory and the other moves 19,476.

The whole library · All essays · What must fail