The state that is removed is not a mode
Worth reading first: The bound that is known in advance · A model that is a rational function.
Balanced truncation has a price that is not estimated. Removing exactly one state from a model costs exactly twice the Hankel singular value that was removed, and the essay that established it reports the ratio of the measured H∞ error to σₙ as 2.0000 on system after system. That equality is what the whole a-priori bound is built from: the many-state bound 2Σσ is the one-state result applied once per removed state, with the terms added up.
An equality that clean invites a reading of what it is an equality about, and there is an obvious one. A stable model written in modal form is a sum of first-order terms,
H(s) = Σₖ rₖ / (s − pₖ),
and each of those terms, taken alone as a one-state system, has a Hankel singular value of exactly |rₖ| ÷ 2|pₖ|. A one-state model has one σ and removing it leaves nothing, so its own cost is twice that number — the same factor of two, arrived at independently. The reading writes itself: the state balanced truncation removes is the mode with the smallest |rₖ| ÷ 2|pₖ|, and σₙ is that number.
The reading is self-consistent, it is arithmetically cheap, it needs no Lyapunov solve, and it is wrong. On the five systems drawn below it over-estimates σₙ by factors of 1.7153, 1.6995, 7.4644, 2.9715 and 1.2071 — never under, and by an amount no arithmetic on the poles and residues reproduces.
The guess is the diagonal of a matrix whose eigenvalues are wanted
The modal number is not an analogy. It is an entry of the object the σ come from, and the coincidence is exact.
For a diagonal model the controllability Gramian solves AP + PAᵀ + BBᵀ = 0, and the routine that solves it returns a matrix whose k-th diagonal entry is −bₖ² ÷ 2pₖ. With the input vector scaled so that bₖ² = |rₖ|, that entry is |rₖ| ÷ 2|pₖ|. The observability Gramian’s k-th diagonal entry is −cₖ² ÷ 2pₖ, and cₖ² is also |rₖ|, because the sign of the residue lives in the sign of cₖ and squaring discards it. Measured on all five systems, the two diagonals agree with the modal sizes to every digit printed: 5.000000·10⁻¹, 8.333333·10⁻² and 3.333333·10⁻⁴ for the three-pole family, 1.000000, 1.750000·10⁻¹, 2.222222·10⁻² and 1.250000·10⁻³ for the four-pole one.
What the guess ignores is everything off that diagonal, and the off-diagonal entries have a closed form too: Pᵢⱼ is −bᵢbⱼ ÷ (pᵢ + pⱼ), so the coupling between two modes is set by how close their poles are. It is not a small correction. Measured as the Frobenius norm of the off-diagonal part against that of the diagonal, it is 0.0738 on the system whose poles are spread over −0.2 to −90, 0.4802 and 0.4934 on the four- and three-pole systems, and 1.0141 on the clustered pair — where the part being discarded is larger than the part being read. Those four numbers order the two extremes correctly: the least coupled system has the smallest over-estimate at 1.2071 and the most coupled has the largest at 7.4644. They do not order the middle, since 0.4802 produces 2.9715 and the slightly larger 0.4934 produces 1.7153. Coupling is the mechanism and it is not a formula for the answer.
So the modal reading is doing something specific and identifiable. The Hankel singular values are the singular values of a product of the two Gramians, computed without ever forming that product; the modal sizes are the diagonal entries of those same Gramians. The guess reads the diagonal of a matrix and calls the answer its spectrum. That is a mistake with a known direction and a known size, and both are measurable here.
Where the diagonal and the spectrum can be compared exactly
There is one family where the comparison is not merely numerical. When every residue carries the same sign, cₖ = bₖ for every k, so the two Gramians are the same matrix — measured, the relative Frobenius difference between P and Q is exactly zero on both such systems here, not small but zero — and the Hankel singular values are simply the eigenvalues of P.
Then the diagonal and the spectrum belong to one symmetric matrix, and everything that is true of a symmetric matrix’s diagonal is available. The trace is preserved: on the three-pole system with residues 1, 0.5 and 0.02, the modal sizes sum to 5.836667·10⁻¹ and the Hankel singular values sum to 5.836667·10⁻¹, the same seven digits. The partial sums do not agree, and the direction in which they fail is the whole finding. Sorted downward, the σ are 5.649435·10⁻¹, 1.852881·10⁻² and 1.943243·10⁻⁴, against modal sizes of 5.000000·10⁻¹, 8.333333·10⁻² and 3.333333·10⁻⁴. The first σ is larger than the first modal size, the first two σ sum to 0.58347234 against 0.58333333, and the totals are equal.
The same holds on the other same-sign system, whose poles at −0.2, −4 and −90 make the Gramian much closer to diagonal. Its modal sizes sum to 2.537778 and its Hankel singular values sum to 2.537778; the leading partial sums are 2.50688964 against 2.50000000 and 2.53754767 against 2.53750000, so the inequality is present but only in the fourth digit, and the over-estimate at the bottom is 1.2071 — the smallest of the five, on the model whose modes interact least.
That pattern is not a coincidence of this model. A symmetric matrix’s eigenvalues always dominate its diagonal in exactly that sense — every leading partial sum is at least as large, with equality at the end — so the smallest eigenvalue can never exceed the smallest diagonal entry. The modal guess cannot under-estimate σₙ. On this system it over-estimates it by 1.7153, and the “never under” that the five measurements show is not an empirical regularity to be checked on a sixth system. It is what the comparison is, whenever the residues share a sign.
One sign, and a fifth of the total disappears
The equality of the two Gramians is fragile in a way the modal sizes cannot see, and the cleanest demonstration is a system built to differ from the one above in one bit.
Keep the poles at −1, −3 and −30 and the residue magnitudes at 1, 0.5 and 0.02, and flip the middle residue to −0.5. Every |rₖ| is unchanged, every |pₖ| is unchanged, so all three modal sizes are identical to the previous system’s, to every digit. The Gramians are no longer equal: their relative Frobenius difference is 8.85·10⁻¹, which is not a perturbation. And the Hankel singular values move — 5.649435·10⁻¹ becomes 4.403556·10⁻¹, a fall of 22.1 per cent, while the second rises from 1.852881·10⁻² to 2.355177·10⁻², a climb of 27.1 per cent.
The system itself has moved by as much. Its H∞ norm falls from 1.16733 to 0.83400, so the same-sign model’s largest gain over all frequencies stands 39.97 per cent above the mixed one’s on the strength of a single residue’s sign. The trace identity is gone with it: the modal sizes still sum to 5.836667·10⁻¹ and the Hankel singular values now sum to 4.641035·10⁻¹, so a fifth of the total the modal reading accounts for is not there. On the four-pole system the shortfall is 19.3 per cent of the same total.
What survives is the bottom. σₙ moves from 1.943243·10⁻⁴ to 1.961339·10⁻⁴, a change of 0.92 per cent, and the over-estimate is 1.6995 where it was 1.7153. So the sign flip is nearly invisible in the number the truncation actually pays and is worth a fifth of everything above it — which is the strongest available statement that the σ are properties of the triple (A, B, C) taken together and the modal sizes are properties of A and the residues taken apart.
The case where the modes and the system have almost nothing in common
Poles at −1 and −1.05 with residues +1 and −1, and a third mode at −8 with residue 0.3. The two clustered modes have individual sizes 5.000000·10⁻¹ and 4.761905·10⁻¹ — the two largest numbers in the whole family — and the system’s largest Hankel singular value is 3.396662·10⁻².
That is 14.7 times smaller than the first modal size and 14.0 times smaller than the second. The model is not a small perturbation of two large modes; it is what is left after they nearly annihilate each other. Over a common denominator the pair is (p₁ − p₂) ÷ ((s − p₁)(s − p₂)), and the numerator is 0.05 because that is how far apart the poles were put. This is the subtraction that takes the answer rather than a digit happening in a map between function spaces rather than in a register: nothing is lost to rounding, the two terms genuinely do cancel, and the cancellation is the model.
The totals record it. The modal sizes sum to 9.949405·10⁻¹ and the Hankel singular values sum to 4.758334·10⁻², a factor of 20.9 — so the modal reading accounts for twenty-one times more model than exists. At the bottom of the spectrum the over-estimate is 7.4644, the largest of the five, and it comes from the third mode: the guess nominates the mode at −8, with size 1.875000·10⁻², and the truncation pays 2.511910·10⁻³.
And the state that is actually removed is not that mode or any other. Written in modal coordinates, the discarded direction has weights 0.6899, 0.7200 and 0.0749 on the three modes — very nearly the difference of the two clustered ones, which is exactly the combination the near-cancellation left weakly connected to the output. The reduced model’s two poles come back at −0.4340 and −8.9776, and neither is one of the three the model was built from. A mode that is hard to reach and hard to see is what a reduction should drop, and here no mode is either; a combination of two of them is both.
The worst case for the guess is the case worth least
The obvious worry is that the modal reading fails where a reader most needs it. The measurement says the opposite, and the two statements turn out to be one statement.
The clustered system’s Hankel spectrum spans a factor of 13.5 from top to bottom. The other four span 2.907·10³, 2.245·10³, 2.136·10³ and 1.089·10⁴. That ratio is what decides whether reduction is worth making at all: it is the gap that turns a rank into a decision, read on the one spectrum whose decay has a closed-form rate, and a factor of 13.5 across a whole spectrum is a model with nothing much to remove.
Priced in the currency that matters, removing one state costs 5.902·10⁻² of the clustered system’s own H∞ norm — six per cent of the model, for one state out of three. The corresponding figures on the other four are 3.329·10⁻⁴, 4.703·10⁻⁴, 4.973·10⁻⁴ and 9.067·10⁻⁵: between one part in two thousand and one in eleven thousand. The system on which the modal guess is off by 7.46 is the system on which no reduction is worth making, and the systems on which it is off by 1.21 and 1.72 are the ones where a state can be removed for a few hundredths of a per cent. The guess is worst where its answer is least needed, which is a real mitigation and not an excuse: it is precisely the shape of an error that survives, because the cases that would expose it are the cases nobody runs.
What does predict the gap, and how far that goes
Something has to be responsible for a factor that runs from 1.21 to 7.46, and the candidate visible in the poles is how close together they are. On a two-mode family it can be swept directly: poles at −1 and −a, residues 1 and ±1, with a running from 100 down to 1.05.
a same sign opposite sign
100 1.0412 1.0404
20 1.2332 1.2110
5 2.5215 2.0225
2 13.1580 5.3423
1.2 221.3779 24.3138
1.05 3281.4645 96.6176
The over-estimate is unbounded, and it is unbounded in both sign conventions for different reasons. With the same sign, two nearly coincident modes add to something very close to a single first-order system, so σ₂ collapses while both diagonal entries stay near a half. With opposite signs the modes cancel instead and σ₁ collapses. Either way the diagonal entries are unmoved, because neither mechanism is visible in |rₖ| ÷ 2|pₖ|.
The two mechanisms leave different models behind, and the difference is the one the section above was about. At a = 1.05 with the same sign, the Hankel spectrum spans a factor of 6.726·10³ — the pair has collapsed into something a single state very nearly describes, and there is a great deal to remove. At a = 1.05 with opposite signs the spectrum spans 5.831, which is a two-state model with two states in it. So the sweep’s largest over-estimate, 3,281, sits on a model that reduces beautifully, and its second largest, 96.6, sits on one that does not reduce at all. Both are failures of the same guess and only one of them would ever be noticed.
Read backwards, that sweep is a candidate predictor: take the mode the guess nominates, take its nearest neighbour in the spectrum, and use the two-mode over-estimate at that pole ratio. It works where the gap is small and fails where it is large. On the widely spread system — poles at −0.2, −4 and −90, a nearest-neighbour ratio of 22.5 — it predicts 1.2037 against a measured 1.2071, three digits. On the three-pole systems, at a ratio of 10, it predicts 1.5466 and 1.4476 against 1.7153 and 1.6995. On the four-pole system it predicts 2.1927 against 2.9715, and on the clustered system it predicts 1.7941 against 7.4644 — wrong by a factor of four, on the one system where the answer mattered.
A one-parameter model of the coupling reproduces the gap only in the regime where there is barely a gap. That is the honest end of this line of enquiry, and it is the same shape as a rank that belongs to the object rather than to how it was written down: σₙ is invariant under every change of state coordinates, the modal sizes are not, and no function of coordinate-dependent quantities is going to reconstruct a coordinate-free one.
The modal sum is a correct decomposition and still the wrong one
None of this is a complaint about modal decomposition as arithmetic. It is exact.
The sum over the modes and the factorised solve are the same function to fifteen digits, so Σ rₖ ÷ (s − pₖ) is not an approximation of H — it is H, in a basis the model was never written in. What fails is the next step. This collection keeps an assertion that a modal sum with one residue deleted still equals the transfer function, feeds it a sixteen-state model with its fourth residue set to zero, and requires it to refuse; it does. Dropping a term from the sum is exactly the operation the modal reading assumes a truncation performs, and it is refused by a computation that runs every time this collection is assembled.
Doing it anyway is a perfectly legitimate reduced model, and it can be priced. Deleting a single mode outright leaves an error of |rₖ| ÷ |pₖ| in the H∞ norm, twice that mode’s own size, so the cheapest available deletion costs 6.6667·10⁻⁴, 6.6667·10⁻⁴, 3.7500·10⁻², 2.5000·10⁻³ and 5.5556·10⁻⁴ on the five systems, against balanced truncation’s 3.8865·10⁻⁴, 3.9227·10⁻⁴, 5.0238·10⁻³, 8.4132·10⁻⁴ and 4.6022·10⁻⁴. The ratios are the same five numbers as before, because they are the same measurement seen from the other end.
The lower bound closes it. No model of order n − 1 whatever has an H∞ error below σₙ, so balanced truncation is within a factor of two of the best possible and modal deletion is between 2.41 and 14.93 times off it. That is the same kind of statement as the equality Eckart–Young makes about a truncated SVD, one factor of two weaker, and it is what makes the comparison a fact rather than a preference.
What follows for anything that reads a σ
A Hankel value is not a mode’s importance, and ranking modes by |r| ÷ 2|p| is not ranking them by what a truncation would pay. Where the spectrum is well separated the two rankings agree and the sizes are within a factor of two; where it is not, they nominate different objects entirely, as the clustered system’s discarded direction does.
Choosing an order from a σ curve is untouched by any of this. The curve is what the bound is a sum of, and it says what an order costs without reference to what a state is. The reading being refused here is about the identity of the removed direction, not about the arithmetic of the bound.
A reduced state has no modal interpretation, and this is the sharpest available demonstration. The discarded direction of the clustered system has components of 0.6899 and 0.7200 on the two clustered modes and 0.0749 on the third, and neither reduced pole is one of the model’s own. Asking which physical mode was dropped has no answer, in the same way that asking which variable a null-space basis eliminated has no answer once a basis nobody chose deliberately has been picked for numerical reasons.
Near-degeneracy defeats a modal reading in the same way it defeats a single-vector Krylov method. Two modes at −1 and −1.05 are not two things one instrument sees separately, which is the same obstruction as an eigenvalue one vector cannot see: in both cases the object that resists is a pair, and a method whose vocabulary has only individual modes in it has no name for the thing that matters.
And a σ still needs a scale before it decides anything. 2.512·10⁻³ is 3.0 per cent of a system whose H∞ norm is 8.512·10⁻² and would be 4.9·10⁻⁴ of one whose norm is 5.076, which is the question of what small is compared to arriving in a field where the answer is the H∞ norm of the model rather than the largest entry of a matrix. The relative figures above — 5.902·10⁻² against 9.067·10⁻⁵ — are the same five measurements with that scale divided out, and they are what a caller deciding an order should read.
The one place the modal reading is not merely harmless is where it is used to choose the reduction rather than to describe one. A method that ranks modes and keeps the largest is the other kind of reduction in miniature — cheap, requiring no Lyapunov solve, and bounding nothing — and its cost here is a factor of between 1.21 and 7.46 in H∞ error, paid on models where the two Gramians would have been affordable. That is the whole of the trade, measured on systems small enough that both sides could be built.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A condition number that is not the model's — both name balanced truncation, gramian, hankel singular values, lyapunov equation, mcmillan degree, singular values
- Exact at the points that were named — both name a-priori bound, balanced truncation, transfer function
- A model that cannot be run — both name balanced truncation, transfer function
- A model with no matrices behind it — both name mcmillan degree, transfer function
- Bracketing an error nobody can measure — both name gramian, lyapunov equation
- The definition asks for more of what defeats it — both name mcmillan degree, transfer function
Named objects
A flat tag is an object no other essay names yet.
A-priori boundBalanced truncationCancellationGramianHankel singular valuesLyapunov equationMcMillan degreeSingular valuesTransfer function