Two errors, and whose fault they are

The number that decides nothing

The determinant is the first scalar anybody attaches to a matrix and the last one worth consulting. A tenth of the identity has a determinant of 10⁻⁶⁰ and a condition number of exactly one. The Hilbert matrix's determinant stops being right at n = 13 and stops being a number at n = 29, and nothing in between reports either.

Worth reading first: The condition number is an amplifier · An answer that is known.

Ask a room of people who have taken one linear algebra course how to tell whether a matrix can be inverted, and the answer comes back immediately: check whether the determinant is zero. It is the first scalar the subject attaches to a matrix and it arrives with a theorem — A is invertible if and only if det A ≠ 0 — that is correct, complete, and useless.

Useless in a specific sense, which this essay is about. The theorem is a statement about the real numbers, where a quantity is zero or it is not. A floating-point matrix is never exactly singular and never exactly anything else, so the question a computation actually asks is how close to singular, and the determinant answers that question wrongly in both directions, by any margin one cares to name.

Four candidate singularity tests on three matrices, at n = 40A table of four quantities on three matrices. 0.1·I at n = 40 has a condition number of exactly 1 and a determinant of 10⁻⁴⁰; a matrix with κ = 10¹⁰ normalised to |det| = 1 has a determinant of one. The determinant and its nth root give the wrong verdict on both, the smallest singular value gives the wrong verdict on the normalised matrix because it carries the units, and only σₘᵢₙ/σₘₐₓ is right on all three. A test is read as saying *singular* when its number is below 10⁻⁶; green cells agree with the true verdict and red ones do not.does this matrix look nearly singular?green: the test agrees with the truth · red: it does not · the bar under each number is its magnitude, over sixty-two decades|det A||det A|^(1/n)σₘᵢₙ1/κ = σₘᵢₙ/σₘₐₓ0.1·I at n = 40perfectly conditioned10⁻⁴⁰0.10.11κ = 10¹⁰, |det| = 1nearly singular1110·10⁻⁶10·10⁻¹¹Hilbert at n = 8nearly singular2.7·10⁻³³8.5·10⁻⁵1.1·10⁻¹⁰6.6·10⁻¹¹the two counterexamplesκ of the scaled identity1its determinant10⁻⁴⁰κ of the normalised matrix10¹⁰its determinant1det(cA) = cⁿ det(A)so a determinant carries the units n times over
Fig. 1 Four candidate answers to is this matrix nearly singular, on three matrices chosen so that they disagree. Green means the number gives the right verdict and red means it does not. Drag the size of the scaled identity: only the first column moves, and only its top two rows.

The counterexample is one line long

Take a tenth of the identity, at n = 60.

Its determinant is 10⁻⁶⁰. Its condition number is exactly 1 — not approximately, exactly: every singular value is 0.1, so the ratio of the largest to the smallest is one. Its columns are orthogonal. A solve against it is a division of each component by a tenth, which loses nothing at all, and its inverse is ten times the identity and is exact.

There is no sense in which this matrix is nearly singular. The scalar a reader was taught to consult is 10⁻⁶⁰, which is below every threshold anybody has ever proposed.

The mechanism is not subtle and it is worth naming, because it generalises. det(cA) = cⁿ det(A). The determinant carries the units of the matrix raised to the power of the dimension. Multiply a 60×60 matrix by ten — change nothing about the problem, since the same scaling applied to both sides of Ax = b leaves x alone — and the determinant moves by sixty orders of magnitude.

The previous phase on this site established that the condition number is a choice of units: κ moves under a row scaling that leaves the answer alone. The determinant is worse by a whole dimension. κ is a ratio of two singular values and a uniform scaling divides out of it exactly; det is a product of all n of them and a uniform scaling multiplies it n times over.

Four candidate singularity tests on three matrices, at n = 120A table of four quantities on three matrices. 0.1·I at n = 120 has a condition number of exactly 1 and a determinant of 10⁻¹²⁰; a matrix with κ = 10¹⁰ normalised to |det| = 1 has a determinant of one. The determinant and its nth root give the wrong verdict on both, the smallest singular value gives the wrong verdict on the normalised matrix because it carries the units, and only σₘᵢₙ/σₘₐₓ is right on all three. A test is read as saying *singular* when its number is below 10⁻⁶; green cells agree with the true verdict and red ones do not.does this matrix look nearly singular?green: the test agrees with the truth · red: it does not · the bar under each number is its magnitude, over sixty-two decades|det A||det A|^(1/n)σₘᵢₙ1/κ = σₘᵢₙ/σₘₐₓ0.1·I at n = 120perfectly conditioned10⁻¹²⁰0.10.11κ = 10¹⁰, |det| = 1nearly singular1110·10⁻⁶10·10⁻¹¹Hilbert at n = 8nearly singular2.7·10⁻³³8.5·10⁻⁵1.1·10⁻¹⁰6.6·10⁻¹¹the two counterexamplesκ of the scaled identity1its determinant10⁻¹²⁰κ of the normalised matrix10¹⁰its determinant1det(cA) = cⁿ det(A)so a determinant carries the units n times over
Fig. 2 The same table at n = 120, where a tenth of the identity has a determinant of 10⁻¹²⁰. Nothing else on the row has moved: the condition number is still one and the matrix is still trivial to solve against.

Two sizes make that a pair of readings, and a pair of readings is a coincidence somebody might explain away. Five make it a law with an exponent in it, and the exponent is the essay’s whole complaint about the determinant.

Four candidate singularity tests on three matrices, at n = 10A table of four quantities on three matrices. 0.1·I at n = 10 has a condition number of exactly 1 and a determinant of 10⁻¹⁰; a matrix with κ = 10¹⁰ normalised to |det| = 1 has a determinant of one. The determinant and its nth root give the wrong verdict on both, the smallest singular value gives the wrong verdict on the normalised matrix because it carries the units, and only σₘᵢₙ/σₘₐₓ is right on all three. A test is read as saying *singular* when its number is below 10⁻⁶; green cells agree with the true verdict and red ones do not.does this matrix look nearly singular?green: the test agrees with the truth · red: it does not · the bar under each number is its magnitude, over sixty-two decades|det A||det A|^(1/n)σₘᵢₙ1/κ = σₘᵢₙ/σₘₐₓ0.1·I at n = 10perfectly conditioned10⁻¹⁰0.10.11κ = 10¹⁰, |det| = 1nearly singular1110·10⁻⁶10·10⁻¹¹Hilbert at n = 8nearly singular2.7·10⁻³³8.5·10⁻⁵1.1·10⁻¹⁰6.6·10⁻¹¹the two counterexamplesκ of the scaled identity1its determinant10⁻¹⁰κ of the normalised matrix10¹⁰its determinant1det(cA) = cⁿ det(A)so a determinant carries the units n times over
Fig. 3 Ten unknowns: |det| = 10⁻¹⁰ with κ = 1.
Four candidate singularity tests on three matrices, at n = 20A table of four quantities on three matrices. 0.1·I at n = 20 has a condition number of exactly 1 and a determinant of 10⁻²⁰; a matrix with κ = 10¹⁰ normalised to |det| = 1 has a determinant of one. The determinant and its nth root give the wrong verdict on both, the smallest singular value gives the wrong verdict on the normalised matrix because it carries the units, and only σₘᵢₙ/σₘₐₓ is right on all three. A test is read as saying *singular* when its number is below 10⁻⁶; green cells agree with the true verdict and red ones do not.does this matrix look nearly singular?green: the test agrees with the truth · red: it does not · the bar under each number is its magnitude, over sixty-two decades|det A||det A|^(1/n)σₘᵢₙ1/κ = σₘᵢₙ/σₘₐₓ0.1·I at n = 20perfectly conditioned10⁻²⁰0.10.11κ = 10¹⁰, |det| = 1nearly singular1110·10⁻⁶10·10⁻¹¹Hilbert at n = 8nearly singular2.7·10⁻³³8.5·10⁻⁵1.1·10⁻¹⁰6.6·10⁻¹¹the two counterexamplesκ of the scaled identity1its determinant10⁻²⁰κ of the normalised matrix10¹⁰its determinant1det(cA) = cⁿ det(A)so a determinant carries the units n times over
Fig. 4 Twenty: |det| = 10⁻²⁰, κ = 1.

Each doubling of the size squares the determinant and leaves every other entry in the table alone, which is already more than a pair of frames can establish — a squaring is a claim about a rate and needs three points before it is one.

Four candidate singularity tests on three matrices, at n = 80A table of four quantities on three matrices. 0.1·I at n = 80 has a condition number of exactly 1 and a determinant of 10⁻⁸⁰; a matrix with κ = 10¹⁰ normalised to |det| = 1 has a determinant of one. The determinant and its nth root give the wrong verdict on both, the smallest singular value gives the wrong verdict on the normalised matrix because it carries the units, and only σₘᵢₙ/σₘₐₓ is right on all three. A test is read as saying *singular* when its number is below 10⁻⁶; green cells agree with the true verdict and red ones do not.does this matrix look nearly singular?green: the test agrees with the truth · red: it does not · the bar under each number is its magnitude, over sixty-two decades|det A||det A|^(1/n)σₘᵢₙ1/κ = σₘᵢₙ/σₘₐₓ0.1·I at n = 80perfectly conditioned10⁻⁸⁰0.10.11κ = 10¹⁰, |det| = 1nearly singular1110·10⁻⁶10·10⁻¹¹Hilbert at n = 8nearly singular2.7·10⁻³³8.5·10⁻⁵1.1·10⁻¹⁰6.6·10⁻¹¹the two counterexamplesκ of the scaled identity1its determinant10⁻⁸⁰κ of the normalised matrix10¹⁰its determinant1det(cA) = cⁿ det(A)so a determinant carries the units n times over
Fig. 5 Eighty: |det| = 10⁻⁸⁰, κ = 1, and the other two matrices’ rows are the same rows they were at ten.
n κ(0.1·I) |det(0.1·I)| |det| of the κ = 10¹⁰ matrix κ(Hilbert 8)
10 1 10⁻¹⁰ 1 1.5·10¹⁰
20 1 10⁻²⁰ 1 1.5·10¹⁰
40 1 10⁻⁴⁰ 1 1.5·10¹⁰
80 1 10⁻⁸⁰ 1 1.5·10¹⁰
120 1 10⁻¹²⁰ 1 1.5·10¹⁰

The determinant is exactly 10⁻ⁿ and the condition number is exactly 1, at all five sizes. One column of that table moves by a hundred and ten orders of magnitude across the slider and every other column is constant. There is no size at which the two quantities come into agreement, no regime where the determinant becomes informative, and no rate at which the disagreement grows or shrinks — it is 0.1ⁿ against 1, and n is the only thing that changed.

And the failure runs in both directions on the same page. The first matrix has κ = 1 and a determinant of 10⁻¹²⁰: as far below any plausible singularity threshold as arithmetic permits, and trivially solvable. The second has κ = 10¹⁰ and a determinant of exactly 1: at the threshold’s comfortable end by that measure, and genuinely hard. A test can be fooled in one direction by bad luck; a test fooled in both directions by matrices this simple is not measuring the thing it is being asked about.

The third row is the control. Hilbert 8’s condition number is 1.5·10¹⁰ at every stop, because it is an 8×8 matrix and the slider does not touch it — which is what makes the first row’s movement attributable to the size and not to anything the figure is doing.

And the other direction, which is the one that hurts

The converse is worse, because it is the failure mode that lets something through.

Take a matrix with a condition number of 10¹⁰ and divide it by its own geometric mean singular value. That is a uniform scaling, so the condition number is untouched — every singular value moves by the same factor and the ratio does not care. And the determinant is now exactly 1.

A matrix that will destroy ten digits of any answer computed from it, wearing a determinant of one.

The two counterexamples sit at opposite ends of one fact. The determinant is not scale-invariant, and singular is a scale-invariant property. A quantity that moves under a transformation the question does not move under cannot be an answer to the question — which is the previous phase’s finding, arriving in the first scalar the subject introduces.

The obvious repair, and why it is not enough

The scale problem is easy to fix. Take the nth root:

|det A|^(1/n)  =  (σ₁ σ₂ … σₙ)^(1/n)

which is the geometric mean of the singular values, and is invariant under exactly the transformation that broke the determinant. A tenth of the identity has |det|^(1/n) = 0.1 at every size. The scaled identity’s row in the table goes green.

And the repair stops there, because the geometric mean is a mean. One singular value at 10⁻¹⁶ among ninety-nine at 1 moves it by 10^(−16/100), which is 0.69. A matrix whose smallest singular value has vanished — the exact case the question is about — has a scale-free determinant of about seven tenths of its typical singular value, and looks entirely healthy.

A quantity that averages over the spectrum cannot report the smallest member of it, and the smallest member is the question. That sentence is the whole objection and it applies to every determinant-derived quantity there is, because they are all products.

The number that does answer the question is σmin/σmax — scale-free, and a statement about the worst direction rather than the typical one. It costs an SVD, or an estimate of one, and an estimate that can be fooled is about what that estimate costs in honesty.

The determinant is also not a number

Everything above is about what the determinant means. There is a separate and equally serious problem, which is that at ordinary sizes it cannot be represented at all.

A determinant is a sum of n! products of n entries. It therefore scales like the nth power of the entries, and n is not small. A 200×200 matrix of entries around 10 has a determinant of about 10²⁰⁰ times whatever the structure contributes, and binary64 stops at 1.8·10³⁰⁸.

The Hilbert matrix goes the other way, and it goes there fast. Its determinant is known in closed form — the entries are 1/(i+j−1), which is a Cauchy matrix, and Cauchy determinants have a formula — so this site can compute the exact answer in BigInt rationals and compare.

At n = 28 the pivot product is 1.04·10⁻³²¹, which is a subnormal: it is inside the format only because binary64 degrades gracefully below 2.2·10⁻³⁰⁸, and it has already lost most of its significand. The numbers below the smallest one is about what that costs.

At n = 29 it is exactly zero.

And every pivot at n = 29 is a perfectly ordinary number. The smallest is 1.2·10⁻¹⁷, the largest is 1, the factorisation completed, no assertion is false and nothing in the computation went wrong. What left the format is the answer, not the arithmetic.

So accumulate the logarithm, which repairs the wrong thing

The fix is standard and it is the shape of this whole phase: do not form the object. The pivots are in hand; take their logarithms and add.

log|det A|  =  Σ log|uᵢᵢ|,   sign = parity of the interchanges × signs of the pivots

One logarithm per pivot on top of a factorisation that was going to happen anyway, and the result cannot overflow until the logarithm does, which is never. At n = 30 it returns −1,194, and −1,194 is an entirely ordinary double.

Except that at n = 30 the true value is −1,194 and the computed value is −811.

log₁₀|det Hₙ| by three routes, to n = 22The exact rational determinant of the Hilbert matrix falls to 10^-274 at n = 22. The accumulated logarithm of the pivots follows it to n = 13 and then departs, reaching 10^-225 — wrong by 49 orders of magnitude and still an ordinary-looking number. The product of the pivots is a separate curve that ends at n = 22, where it underflows to exactly zero.25811141720-300-250-200-150-100-500nlog₁₀|det Hₙ|exact rationalΣ log|uᵢᵢ|Π uᵢᵢthree routes, one theoremexact at the largest n-274accumulated logarithm-225decades of disagreement49smallest pivot at that n4·10⁻¹⁷every pivot is a normal numberat every size on this axis
Fig. 6 The same sweep stopped at n = 22, where the two float curves have already separated from the exact one by more than a hundred orders of magnitude in the determinant while still looking like two smooth curves.

The accumulated logarithm departs from the truth at n = 13, and the product underflows at n = 29. Sixteen sizes apart, and the accuracy limit comes first. At n = 20 the computed determinant is 10³⁰·⁶ times the true one — a perfectly ordinary-looking negative number, with nothing on the page to say it is wrong by thirty orders of magnitude.

This is the finding this essay was written to reach and it was not the one expected. The substitution repairs representability and does nothing for conditioning, and the conditioning gives out first by a wide margin. A reader who takes “accumulate the logarithm” as the fix has bought sixteen sizes of an answer that is absent and inherited sixteen sizes of an answer that is wrong.

And the two float routes agree with each other

The trap has a second floor. The site’s standing habit is two routes to every number, and the natural second route to a log-determinant is Σ log σᵢ from the singular values — a completely different algorithm, sharing no arithmetic with the pivots.

At n = 20 the pivots give −448.55 and the singular values give −449.09. They agree to three digits. The exact answer is −518.95.

Two independent computations, agreeing with each other and not with the truth, because both are computing the determinant of the same rounded matrix and the rounded matrix is the problem. The Hilbert matrix’s entries are 1/3, 1/5, 1/7 — none of them representable — so what the routine is handed at n = 20 is a matrix whose determinant genuinely is about 10⁻⁴⁴⁹, and it computes it correctly.

That is the sharpest version of this site’s own sentence. The algorithm is blameless. The problem is sensitive by 10¹⁸. And no second route can find it, because the two routes share the input.

The one exact part, which is also wrong

Something in a determinant does not round. The sign is a parity of row interchanges multiplied by the signs of the pivots: comparisons and an integer count, with no arithmetic in them at all.

The Hilbert matrix is positive definite. Every pivot of its exact unpivoted elimination — which is Cholesky’s pivot sequence — is positive, and det Hₙ > 0 at every n. That is a theorem.

In binary64 the sign that comes back from the factorisation is negative at n = 15, and wrong again at 18, 19, 20, 21, 24, 26 and 29.

No rounding happened in the sign. What happened is that an exact function was applied to an inexact input, and it returned a wrong answer with no error term attached to it. An inaccurate number carries its own warning, since one can ask how inaccurate. A wrong sign is a clean, confident, integer falsehood.

It does not degrade, it randomises

The list above is not a threshold and that is the point of printing all of it.

Eight of the seventeen sizes from n = 14 to n = 30 report the wrong sign, and the other nine report the right one — right at 14, 16, 17, 22, 23, 25, 27, 28 and 30, wrong at 15, 18, 19, 20, 21, 24, 26 and 29. Below n = 13 it is right at every size. So there is no n past which the sign is reliably wrong, and no trend to extrapolate from a run at one size to a run at another.

That is a worse property than being wrong. A quantity that fails from some size onwards can be distrusted past that size; a quantity that is right 53 per cent of the time past it cannot be distrusted at all, because distrust would be as unfounded as trust. And unlike the magnitude, it offers nothing to calibrate against: an inaccurate number invites the question how inaccurate, and the answer is a number. A sign has no such answer. It is a clean, confident integer with no error term that could exist.

The randomisation is the rounding rather than the routine, which the early rows establish: below the accuracy limit every size is right, so nothing about the algorithm is choosing a sign badly. Once the pivots are noise, their product’s sign is the parity of however many of them happened to land negative, and that is a coin toss with no memory of the size before it.

And the error only ever points one way

The magnitude has the opposite property and it is equally worth having.

log|det Hₙ| computed from the pivots, minus the exact value, runs 0.03 decades at n = 12, then 0.45, 0.53, 4.13, 7.16, 11.63, 15.16, 22.23, 30.57 … and 166.09 at n = 30. Every one of those is positive. At no size in the sweep does the computed determinant understate the true one.

So the error is a bias and not noise, and its direction is the harmful one. The computed determinant consistently reports more volume than the matrix has — which is to say it reports a matrix as further from singular than it is. At n = 20 the computed |det| is 10³⁰·⁶ times the truth, in exactly the direction that would carry a matrix past a singularity threshold it should have failed.

That closes the argument the essay opens with. The first objection is that the determinant answers a question about volume when it was asked one about conditioning. The second is that making it scale-free does not help, because a mean cannot see its smallest member. The third is this one: where it is computed anyway, past n = 13 on this family, the number that comes back is wrong by an unbounded amount, in one direction, and the direction is reassuring. A quantity that fails towards alarm is a nuisance. One that fails towards reassurance is a hazard, and it is the same hazard the scaled identity at the top of the page has — a matrix that will destroy ten digits, wearing a determinant of one.

What the determinant is exactly right for

This is not an argument that the determinant is a bad idea, and two of its jobs are untouched.

Volume. |det A| is the factor by which A multiplies volume, and that is not an approximation to anything — it is what the determinant is. Hadamard’s inequality says the volume of a parallelepiped is at most the product of its edge lengths,

|det A|  ≤  ‖a₁‖ ‖a₂‖ … ‖aₙ‖

with equality exactly when the edges are orthogonal. Written as a ratio it is scale-free — both sides carry the same units — and it is a genuinely useful number: a volume efficiency between 0 and 1 that is 1 for an orthogonal matrix and falls as the columns line up. On random 8×8 Gaussian matrices it averages 0.037.

Orientation. The sign, where the input is exact. A permutation matrix, an integer matrix, a rotation given by its entries: for those the sign is a fact and the whole objection above evaporates, because the objection was about an exact function of an inexact input.

One consequence for the volume reading is worth stating, since that is the use this essay defends. The Hadamard ratio is scale-free — both sides of the inequality carry the same units — so it escapes the first objection entirely. It does not escape the second. It is still built from a product over every column, so a matrix with ninety-nine well-spread columns and one that has quietly aligned with the others has a ratio barely distinguishable from one with no alignment at all. Volume efficiency is an honest answer to how orthogonal is this basis on average, and it is the same wrong answer to is this basis about to become dependent that the determinant itself gives. The scale-free repair fixes the units and not the averaging, and the averaging is what the singularity question is about.

Where it is used anyway, and what to use instead

Four places where a determinant appears in code, and what belongs there:

A singularity test. Use σmin/σmax, or an estimate of it — every library’s rcond — and never abs(det(A)) < tol. LAPACK does not compute a determinant at all; xGECON estimates the reciprocal condition number and that is the number a driver checks.

A multivariate normal density. The log-likelihood needs log|det Σ|, not det Σ, and Σ is positive definite, so the honest route is 2 Σ log Lᵢᵢ from a Cholesky factor. Every statistics library that computes a Gaussian density this way is doing exactly the substitution above — and inherits exactly the accuracy limit, on covariance matrices that are routinely ill-conditioned.

A characteristic polynomial. det(λI − A) as a route to eigenvalues is the classic bad idea and the site has drawn what happens: the polynomial’s coefficients are catastrophically worse conditioned than the eigenvalues themselves, and the roots of a computed polynomial are not the eigenvalues of the matrix.

A change of variables. Here the determinant is the answer, it is usually 2×2 or 3×3, the entries are exact, and none of this applies — which is the same exemption exact arithmetic buys everywhere on this site.

What is worth carrying

The determinant answers a question about volume and is asked a question about conditioning. Those are different questions, and the second one is scale-free while the determinant is not, by a whole power of the dimension.

Making it scale-free does not make it right, because it is still a mean. The geometric mean of the singular values cannot see one small singular value among ninety-nine ordinary ones, and one small singular value among ninety-nine ordinary ones is the entire subject.

And two limits arrive, sixteen sizes apart, with the wrong one first. The product underflows at n = 29 and the answer stopped being right at n = 13. The repair everybody knows — accumulate the logarithm — buys the second limit and nothing at all for the first, and the number in between is wrong by thirty orders of magnitude while looking entirely ordinary.

The next essay takes the same subject into the one place a determinant is still taught as a method: a rule that is correct and unusable.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Condition numberDeterminantExact ground truthHilbert matrixLU factorisationOrthogonalityRankSingular valuesSubnormalUnderflow