The eigenvalue problem that is not linear

The number that moves when the problem does

Two quantities are offered as the condition number of one eigenvalue. One is unmoved to eight digits by a change of variable that is exact in both directions, and grows like the square root of the chain's length. The other is inflated by ten orders by that change of variable, and is ten times too large before anything has been done at all.

Worth reading first: An index that is a pair · A matrix that depends on its own eigenvalue · The exact answer to a nearby problem.

The spine of this site is an inequality with three quantities in it: a forward error, a backward error, and a condition number between them that says how much the problem magnifies a perturbation. Two of the three are measurements of something that happened. The middle one is a property claimed about the question itself, and the condition number is an amplifier is where that claim was first written down.

For one eigenvalue of one quadratic, this site computes two numbers that are both offered as that middle factor. One is the condition number of the eigenvalue as an eigenvalue of the quadratic — the three coefficient norms in the numerator, and the derivative Q′(λ) between the left and right null vectors in the denominator. The other is the ordinary eigenvalue condition number of the 2n × 2n matrix the solver was actually handed, which is the quantity every error analysis the solver’s authors wrote is stated against. They are both defensible, they are computed from the same problem, and at four masses one of them is 3.596 and the other is 36.5.

So there is a question with an answer, and it is not a matter of preference: which of the two belongs in the inequality. The way to settle it is to notice that “a property of the problem” is a testable sentence rather than a definition, and that it takes two tests rather than one. The first is invariance: apply a change of variable that is exact in both directions, so that the description changes and the problem does not, and an honest condition number cannot move. The second is response: change the problem, and an honest condition number must move with it. A quantity that fails the first is measuring the route. A quantity that passes the first and fails the second is measuring nothing.

The condition number the problem has, and the one the solver's error analysis is written againstThe same eigenvalue of the same overdamped chain of 4 masses, in seven systems of units. Its condition number as an eigenvalue of the QUADRATIC — Tisseur's, with the three coefficient norms in the numerator and yᵀQ′(λ)x in the denominator — is 3.596 at γ = 1 and 3.596 at γ = 10⁶, a spread of 1 over six decades: it cannot move, because a change of units is not a change of problem. Its condition number as an eigenvalue of the LINEARISED MATRIX runs 36.5 to 5.929·10¹¹, a factor of 1.62·10¹⁰. The forward error follows the second one, and the first one is the honest description of the problem — so the substitution has manufactured an ill conditioning that belongs to the algorithm rather than to the question.012345610⁻¹10²10⁵10⁸10¹¹log₁₀ γ, the change of unitscondition numberthe linearisationthe quadraticone problem, two amplifiersκ(quadratic), first3.6κ(quadratic), last3.6κ(linearisation), last5.9·10¹¹how far the first moved1the problem is as well conditioned as everand the method is not
Fig. 1 An overdamped chain of four masses, in seven systems of units. The flat line is the eigenvalue’s condition number for the quadratic, 3.596 at γ = 1 and 3.596 at γ = 10⁶. The climbing one is its condition number for the linearised matrix, 36.5 to 5.929·10¹¹. The slider is the number of masses.

A change of description is not a change of problem

The transformation is the one a backward-stable answer to a problem nobody asked is built on. Replacing λ by γμ in Q(λ) = λ²M + λC + K turns the coefficients into

(γ²M, γC, K),

and divides the spectrum by γ exactly. Nothing is approximated and nothing is chosen. If time is measured in milliseconds rather than seconds, γ is a thousand; the chain is the same chain, its modes are the same modes, and the map back is a multiplication. Every stop of the sweep across the figure above is one physical problem written down in a different system of units, which is the property that makes it a test rather than an experiment: whatever the right answer is at one stop, it is the right answer at all of them.

The quadratic’s condition number does not move. At four masses it is 3.596 at both ends of six decades, and the invariance is stronger than the four figures a caption can print. Measured at eight masses, the seven stops return 4.980463300628977, 4.980463300628978, 4.980463300628978, 4.980463300628514, 4.980463300719729, 4.980463308545053 and 4.980463401308598 — sixteen digits of agreement at the first three stops, and 2.02·10⁻⁸ of relative spread across all seven. What departs at the far end is the computed eigenvalue that the quantity is evaluated at, not the quantity. This is invariance by construction: the numerator and the denominator both carry one factor of γ for every power of λ, and the ratio has none.

The linearised matrix’s condition number climbs by a factor of 1.62·10¹⁰ across the same six decades at four masses, and by 1.57·10¹⁰ at fourteen. The climb has a clean shape once it is past the balanced region: from γ = 10² onward it is exactly two decades per decade of units, 0.593γ² at four masses and 0.784γ² at eight, which is what the identity blocks in the companion form do to the norm of a matrix whose other blocks have been scaled. That is the same diagnostic the units the matrix is measured in gives for a row scaling, arriving here from the coefficients rather than from the rows.

The condition number the problem has, and the one the solver's error analysis is written againstThe same eigenvalue of the same overdamped chain of 6 masses, in seven systems of units. Its condition number as an eigenvalue of the QUADRATIC — Tisseur's, with the three coefficient norms in the numerator and yᵀQ′(λ)x in the denominator — is 4.338 at γ = 1 and 4.338 at γ = 10⁶, a spread of 1 over six decades: it cannot move, because a change of units is not a change of problem. Its condition number as an eigenvalue of the LINEARISED MATRIX runs 43.48 to 6.925·10¹¹, a factor of 1.59·10¹⁰. The forward error follows the second one, and the first one is the honest description of the problem — so the substitution has manufactured an ill conditioning that belongs to the algorithm rather than to the question.012345610⁻¹10²10⁵10⁸10¹¹log₁₀ γ, the change of unitscondition numberthe linearisationthe quadraticone problem, two amplifiersκ(quadratic), first4.3κ(quadratic), last4.3κ(linearisation), last6.9·10¹¹how far the first moved1the problem is as well conditioned as everand the method is not
Fig. 2 Six masses. The flat line has moved up to 4.338 and is still flat; the climbing one starts at 43.48 and ends at 6.925·10¹¹, a factor of 1.59·10¹⁰ across the same six decades.

Invariance alone certifies nothing

The first test disqualifies one candidate, and it is tempting to stop there. It should not be, because invariance is cheap: a constant passes it, and so does any quantity built to have no net power of γ in it. The interesting failure is not a constant but a real number that looks like an amplifier, is genuinely invariant, and is useless.

There is one in this family, and it costs three norms. The quantity ‖C‖²/(‖M‖‖K‖) is exactly invariant under the change of variable — the numerator picks up γ² and so does the denominator — and it is measured at 21.0536 at γ = 1 and 21.0536 at γ = 10⁶, identical to six figures. It is not a constant either. It is essentially the chain’s overdamping margin, and it responds strongly to the thing it is about: at eight masses it reads 6.855 with the damping coefficient at three, 20.62 at six, 70.66 at twelve, and 44.62 when the stiffness-proportional part of the damping is raised fourfold.

And it does not respond to the size of the problem at all. Across the range this essay measures — four masses to fourteen — it reads 21.0536, 20.7592, 20.6169, 20.5330, 20.4777 and 20.4385, a change of 3.0%, while the quadratic’s condition number over the same six problems moves by 82%. A code that reported it as the amplifier in the inequality would be reporting a number that is invariant, dimensionless, cheap, and blind to the one axis along which these problems get harder.

So the second test is not a formality. This is the same shape of argument a condition number scaling cannot move makes for a linear system, where a row scaling leaves the componentwise number alone, and it is the same warning: the transformation that a quantity survives has to be one it could have failed, and surviving it is half of a certificate rather than the whole of one.

The condition number the problem has, and the one the solver's error analysis is written againstThe same eigenvalue of the same overdamped chain of 10 masses, in seven systems of units. Its condition number as an eigenvalue of the QUADRATIC — Tisseur's, with the three coefficient norms in the numerator and yᵀQ′(λ)x in the denominator — is 5.554 at γ = 1 and 5.554 at γ = 10⁶, a spread of 1 over six decades: it cannot move, because a change of units is not a change of problem. Its condition number as an eigenvalue of the LINEARISED MATRIX runs 55.22 to 8.685·10¹¹, a factor of 1.57·10¹⁰. The forward error follows the second one, and the first one is the honest description of the problem — so the substitution has manufactured an ill conditioning that belongs to the algorithm rather than to the question.012345610⁻¹10²10⁵10⁸10¹¹log₁₀ γ, the change of unitscondition numberthe linearisationthe quadraticone problem, two amplifiersκ(quadratic), first5.6κ(quadratic), last5.6κ(linearisation), last8.7·10¹¹how far the first moved1the problem is as well conditioned as everand the method is not
Fig. 3 Ten masses. The quadratic’s number is 5.554 at every stop and the linearisation’s runs 55.22 to 8.685·10¹¹. Between four masses and ten the flat line has risen by 54% and the shape of the figure has not changed at all.

The second test, and the rate at which it is passed

Six sizes, the same overdamped chain, the same eigenvalue at each — the median real one, which is the eigenvalue the figures on this page are drawn for. The quadratic’s condition number reads 3.596, 4.338, 4.980, 5.554, 6.076 and 6.557 at four, six, eight, ten, twelve and fourteen masses. It moves, which is the whole of what the second test asks, and it moves at a rate worth stating: divided by √n it reads 1.798, 1.771, 1.761, 1.756, 1.754 and 1.753, so about 1.75√n describes the whole range to within 2.5% while the number of masses rises by a factor of 3.5. Fitted as a power of n over that range the exponent is 0.480.

The mechanism is available in the same three norms, and it is worth reading off rather than asserting, because it explains why the rate is a square root and not something else. The condition number is a numerator |λ|²‖M‖ + |λ|‖C‖ + ‖K‖ over a denominator |λ| times |yᵀQ′(λ)x|/(‖x‖‖y‖). Between four masses and fourteen the numerator goes from 12.131 to 24.092, a factor of 1.986. The eigenvalue itself barely moves — −0.494647 to −0.531269, a factor of 1.074 — and the derivative term moves less still, 6.8197 to 6.9156, a factor of 1.014. The three combine to 1.823, which is 6.557/3.596 to four figures. The growth is entirely the numerator’s, and the numerator grows like √n because all three coefficient norms do: ‖M‖ runs 2.000 to 3.742, ‖C‖ runs 14.05 to 26.32, and ‖K‖ runs 4.690 to 9.055, each of them the Frobenius norm of a banded matrix whose entries do not change and whose number of rows does.

The linearised matrix’s number responds to the size too — 36.5 to 65.03 at γ = 1 — so the second test does not disqualify it. That matters, and it is why the first test had to come first. Neither quantity is inert; one of them is inert under a transformation it should have been inert under, and the other is not.

The condition number the problem has, and the one the solver's error analysis is written againstThe same eigenvalue of the same overdamped chain of 12 masses, in seven systems of units. Its condition number as an eigenvalue of the QUADRATIC — Tisseur's, with the three coefficient norms in the numerator and yᵀQ′(λ)x in the denominator — is 6.076 at γ = 1 and 6.076 at γ = 10⁶, a spread of 1 over six decades: it cannot move, because a change of units is not a change of problem. Its condition number as an eigenvalue of the LINEARISED MATRIX runs 60.31 to 9.461·10¹¹, a factor of 1.57·10¹⁰. The forward error follows the second one, and the first one is the honest description of the problem — so the substitution has manufactured an ill conditioning that belongs to the algorithm rather than to the question.012345610⁻¹10²10⁵10⁸10¹¹log₁₀ γ, the change of unitscondition numberthe linearisationthe quadraticone problem, two amplifiersκ(quadratic), first6.1κ(quadratic), last6.1κ(linearisation), last9.5·10¹¹how far the first moved1the problem is as well conditioned as everand the method is not
Fig. 4 Twelve masses: 6.076 flat, against 60.31 climbing to 9.461·10¹¹. The gap between the two curves at the left-hand edge of the figure is the same gap it was at four masses.

Ten times too large before anything has been done

The finding a single size cannot produce is at the left-hand edge of every one of these figures, where γ is one and nothing has been rescaled. The ratio of the two condition numbers there reads 10.15, 10.02, 9.97, 9.94, 9.93 and 9.92 at four, six, eight, ten, twelve and fourteen masses. The linearisation’s number is already an order of magnitude too large on the problem as posed, and the overstatement is stable enough across six sizes to be a property of the construction rather than an accident of one chain.

At one size that number is unreadable. Ten is exactly the sort of factor a reader forgives as the looseness of a definition, and there is no way to tell a systematic factor from a coincidence with a single measurement. Six of them agreeing to within 2.3% is a different statement.

Where the factor comes from is measurable too, and the answer is not the obvious one. The natural guess is the numerator: the linearised matrix is 2n × 2n and carries identity blocks, so its norm must be the larger. Measured, ‖A‖ divided by the quadratic’s numerator |λ|²‖M‖ + |λ|‖C‖ + ‖K‖ is 1.232 at four masses, 1.181 at eight and 1.166 at fourteen — a fifth of the factor at most, and shrinking.

The rest is the denominator, and there the two quantities are not comparable at all. The linearisation’s denominator is |vᵀu|/(‖u‖‖v‖), the cosine of the angle between the left and right eigenvectors, which measures 0.828, 0.816 and 0.813 at the three sizes and cannot exceed one by definition. That is the quantity a condition number for one eigenvalue is built from, and its being a cosine is the whole of its meaning. The quadratic’s denominator is |yᵀQ′(λ)x|/(‖x‖‖y‖), which is not a cosine: Q′(λ) is 2λM + C, so the term carries the size of the coefficients, and it measures 6.820, 6.894 and 6.916. The ratio of the two denominators is 8.236, 8.446 and 8.508, and multiplying by the numerator ratio reproduces 10.150, 9.971 and 9.917 exactly. The factor of ten is the norm of Q′ in this family, sitting in a denominator where the other definition has a number bounded by one, and two condition numbers of one matrix is the earlier case of a ratio between two definitions carrying information rather than noise.

The condition number the problem has, and the one the solver's error analysis is written againstThe same eigenvalue of the same overdamped chain of 14 masses, in seven systems of units. Its condition number as an eigenvalue of the QUADRATIC — Tisseur's, with the three coefficient norms in the numerator and yᵀQ′(λ)x in the denominator — is 6.557 at γ = 1 and 6.557 at γ = 10⁶, a spread of 1 over six decades: it cannot move, because a change of units is not a change of problem. Its condition number as an eigenvalue of the LINEARISED MATRIX runs 65.03 to 1.018·10¹², a factor of 1.57·10¹⁰. The forward error follows the second one, and the first one is the honest description of the problem — so the substitution has manufactured an ill conditioning that belongs to the algorithm rather than to the question.012345610⁻¹10²10⁵10⁸10¹¹log₁₀ γ, the change of unitscondition numberthe linearisationthe quadraticone problem, two amplifiersκ(quadratic), first6.6κ(quadratic), last6.6κ(linearisation), last10¹²how far the first moved1the problem is as well conditioned as everand the method is not
Fig. 5 Fourteen masses, the largest size at which both null vectors are computed here. The quadratic’s number is 6.557 and the linearisation’s runs 65.03 to 1.018·10¹².

Which of the two closes the accounting

Neither test appeals to what the answers are worth, which is deliberate — the two tests are about the shape of a quantity rather than about any run. But the inequality is an empirical claim, and the verdict can be checked against the errors themselves.

The check is the eight-decade sweep at eight masses, where the exact spectrum is known in closed form and the forward error is a measurement rather than an estimate:

γ forward error η, the linearisation η, the quadratic
1 7.49·10⁻¹⁴ 6.76·10⁻¹⁶ 1.40·10⁻¹⁵
10⁴ 8.57·10⁻⁹ 7.81·10⁻¹⁵ 1.44·10⁻¹⁰
10⁸ 1.26·10⁻³ 7.59·10⁻¹³ 1.23·10⁻⁴

Each condition number has a backward error that belongs with it, and the two products are the two predictions the inequality makes. The quadratic’s pair gives 6.98·10⁻¹⁵ at γ = 1 and 6.12·10⁻⁴ at γ = 10⁸, against forward errors of 7.49·10⁻¹⁴ and 1.26·10⁻³ — so the answer is 10.7 times the prediction at one end and 2.06 times it at the other, a factor of five of drift across eight decades of units. The linearisation’s pair gives 3.36·10⁻¹⁴ and 5.96·10³. The second of those is not an error; it is a prediction of five thousand for a quantity that measured 0.00126, and the drift between the two ends is a factor of 10⁷.

Both products are legitimate arithmetic and only one of them is informative. The linearisation’s condition number at γ = 10⁸ is 7.844·10¹⁵ and its backward error is 7.59·10⁻¹³, and the product is true in the sense that the forward error is indeed below it. A bound that overshoots by seven orders has told nobody anything, which is the standing three errors and one number sets for whether an accounting has closed: the test of the middle factor is that the product tracks the answer, not that it exceeds it.

One quadratic eigenvalue problem in 9 systems of units: what the solver reports and what the answer is worthAn overdamped chain of 8 masses, with λ replaced by γμ so that the coefficients become (γ²M, γC, K). That substitution is exact in both directions and divides the spectrum by γ exactly, so the closed form is still available and every error here is measured against it. The backward error of the eigenpair for the LINEARISED MATRIX — the residual a solver's own error analysis is about — is 6.76·10⁻¹⁶ at γ = 1 and 7.59·10⁻¹³ at γ = 108 — it moves by a factor of 1928 while the other two move by 1.68·10¹⁰. The backward error for the QUADRATIC, which is what the person who posed the problem is entitled to, grows by 8.76·10¹⁰ across the same sweep, and the forward error follows it: 7.49·10⁻¹⁴ to 0.00126. Nothing went wrong with the solver at any stop.0246810⁻¹⁷10⁻¹⁴10⁻¹¹10⁻⁸10⁻⁵10⁻²10¹log₁₀ γ, the change of unitsrelative errorforward errorη, the quadraticη, the linearisationagainst a closed formη(linearisation), worst7.6·10⁻¹³η(quadratic), worst1.2·10⁻⁴forward error, worst0.0013coefficient spread4.2·10¹⁵the solver is right at every stopabout a problem nobody asked
Fig. 6 The same problem across nine systems of units, with the two backward errors and the forward error the condition numbers above are being tested against. The linearisation’s backward error is 6.76·10⁻¹⁶ at γ = 1 and 7.59·10⁻¹³ at γ = 10⁸; the forward error goes from 7.49·10⁻¹⁴ to 1.26·10⁻³.

This is also the page’s published refusal, and it is stated as the arithmetic a reader would do rather than as a doctrine. The claim fed to it is that the forward error is the residual the solver prints multiplied by a condition number. At γ = 10⁸ the forward error divided by the printed residual is 1.66·10⁹, so the multiplier the claim requires is 1.66·10⁹, and the condition number of the eigenvalue is 4.98. The assertion that the two agree within four orders of magnitude is fed those numbers every time the figure is drawn, and fails by five. The residual is real, the condition number is real, and the thing blamed for the gap between them is neither.

Where the factor of ten stops being a factor of ten

Three qualifications, each measured, and the first is the one most likely to be over-read from the figures above.

The factor of ten is stable in the size and not in the coefficients. Holding the chain at eight masses and changing only the damping, the ratio of the two condition numbers reads 3.756, 5.569, 9.971, 15.58 and 30.59 as the mass-proportional damping coefficient goes 3, 4, 6, 8, 12, and 7.439, 9.971, 14.88 and 27.46 as the stiffness-proportional part goes 0.2, 0.5, 1, 2. The quadratic’s own number over that same set moves only between 4.068 and 7.100. So the overstatement is a property of the balance between the three coefficients, and “about ten” describes this family rather than linearisation in general. What survives is the direction, which is what a perturbation that moves every coefficient also finds when it restricts a backward error to one coefficient: the substitution never understates.

The square root describes one eigenvalue and not the spectrum. The eigenvalue these figures follow is the median real one, and in this family it is the best-conditioned eigenvalue there is. Taking instead the worst over the whole spectrum, the quadratic’s condition number reads 14.74, 32.47, 59.73, 97.92, 148.3 and 211.8 at the six sizes — a factor of 14.4 while the number of masses rises by 3.5, which is an exponent of 2.13 rather than 0.48. The hardest eigenvalue is the smallest in modulus, −0.00727 at fourteen masses, and it is hard for the reason a matrix that depends on its own eigenvalue gives: at a small λ the numerator is dominated by ‖K‖ and the denominator is being multiplied by |λ|. Both readings are responses to the problem and they are responses of different orders, so the rate is a property of a chosen eigenvalue rather than of the family.

And no change of units removes the inflation. The linearisation’s condition number has a minimum in γ, and it is not at one: swept finely, the least value is at γ = 1.778 at every size measured, where it reads 34.13, 46.00 and 60.06 at four, eight and fourteen masses. The ratio to the quadratic’s number at that best stop is 9.49, 9.24 and 9.16 — so the optimal system of units buys about 7% of the factor of ten and no more. The inflation is also two-sided: at eight masses, lowering γ to 10⁻⁶ raises the linearised number to 7.297·10¹², which is the same climb with ‖K‖ in the dominant position instead of ‖M‖.

One more control belongs here, because it is the trap the whole essay is about, laid for the two survivors. Multiplying M, C and K all by one constant is another transformation that changes nothing — it multiplies Q by a scalar and leaves every eigenvalue where it was. Both condition numbers pass it: at eight masses, with the constant at 1, 10³ and 10⁶, both read 4.98046 and 49.6592 unchanged to six figures. A test that both candidates pass separates nothing, which is exactly what the invariant three norms above demonstrated one level up, and it is the reason the change of variable has to be one that moves the coefficients relative to each other.

What follows for anything that reports one

Report the quadratic’s number, and report it beside the quadratic’s backward error. The pair costs two null vectors and three norms, and it is the pair whose product tracked the answer to within a factor of five across eight decades. The pair a solver hands over by default drifts by 10⁷ over the same range.

A condition number quoted without saying which object it is about is not a measurement. Two numbers on this page describe one eigenvalue of one problem, and at the honest end of the sweep they differ by a factor of ten and at the far end by about 1.6·10¹¹. Every essay in this collection that reports an amplifier now says which object it amplifies for, and the exact answer to a nearby problem is where the requirement to name the object came from.

Rescale before conditioning is discussed at all. The whole climb in these figures is undone by putting γ back where the coefficient norms want it, which is the scaling that buys ten orders, and the argument for bothering is that it is otherwise invisible: at every stop the residual is at the rounding level and nothing announces itself. A transformation chosen for what it does to the amplifier is the general version, and changing the condition number on purpose is the case where the amplifier is what the transformation is chosen for.

A comparison of linearisations is a different question from this one. Six routes to one spectrum differ in forward error by a factor of forty at badly chosen units, and six routes to one spectrum reports that the comparison anybody runs first — comparing the computed eigenvalues — cannot separate them. Choosing between linearisations and choosing between definitions of a condition number are two different choices, and the second one is worth ten orders where the first is worth forty.

And the certificate is two tests, in that order. Invariance first, because it is the one a plausible quantity fails; response second, because it is the one a plausible quantity passes vacuously. Applied to the two candidates here the verdict is unambiguous: the quadratic’s number is unmoved to 2.02·10⁻⁸ by a change of units and rises by 82% across the sizes, and the linearisation’s moves by 1.6·10¹⁰ under the transformation that changes nothing. A quantity that answers to the units is describing the units, and there is no arrangement of the inequality in which it is the middle factor.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Backward errorCompanion formCondition numberConditioningForward errorInvarianceLinearisationMatrix polynomialQuadratic eigenvalue problemScaling