bfloat16 — where it appears
Also named here as dynamic range — the same set of essays touches all of them, so they are one junction rather than several.
Where the hardware went
bfloat16 carries eight mantissa bits, which puts its refinement threshold at a condition number of 256. That is not an exotic matrix. It is an ordinary one, and past it the method still improves the answer by a factor of four hundred while getting nowhere near a usable one.
The other half of a format
fp16 and tf32 have the same eleven significand bits and their largest numbers are 65,504 and 3.4·10³⁸. For two phases this site simulated the significand alone, so it was obliged to report them as the same format — which is a claim, and a false one.
A norm that overflows before it is a norm
The vector of sixteen thousands has a Euclidean norm of 4,000, which fp16 represents exactly. Written as the square root of the sum of squares it returns infinity, because squaring doubles the exponent — and the expression costs half the format's range on the one computation every iterative method performs at every step.
Named alongside it
The objects these essays reach for when they reach for this one.
Dynamic rangeHalf precisionOverflowSignificandUnit roundoffCatastrophic cancellationCondition numberExponent rangehypotIEEE 754Iterative refinementScaled norm