Where the answer stops being in the data
Worth reading first: When the answer is a choice.
The previous essay establishes that the answer to an ill-posed problem is a choice, and that both standard methods make it by weighting the terms of one sum. It does not say where the weight should fall.
There is one quantity that answers that from the data alone — no knowledge of the true signal, no knowledge of the noise level, nothing but A and b. It is the discrete Picard condition, and it is the only thing in this field that is a measurement rather than a heuristic.
It also does not give the right answer, and the amount by which it is wrong is consistent enough to be a finding rather than an error.
The condition
The filtered solution is Σₖ fₖ (uₖᵀb/σₖ) vₖ, so every term is a coefficient divided by a singular value. For that sum to mean anything, the coefficients have to fall faster than the singular values — otherwise the later terms grow without bound and the sum is dominated by whatever is smallest.
The fₖ are the filter factors, and the only thing that distinguishes one regularisation method from another is what it puts in them. Truncation sets fₖ to 1 below some index and 0 above it — a step, and a decision about exactly where to put the step, which is what everything below is about. Tikhonov sets fₖ = σₖ²/(σₖ² + λ²), which is the same decision made smoothly: near 1 where σ is large against λ, near 0 where it is small, with a transition a decade or so wide instead of a cliff. Both answer the same question — how much of each singular direction to believe — and the Picard crossing is a claim about where the answer changes from “all of it” to “none of it”.
That is the Picard condition, and for an exact right-hand side it holds. Measured on the noise-free b: the ratio |uₖᵀb|/σₖ over the range where σ is above rounding spans less than four orders of magnitude across the whole spectrum, which is a bounded sum. The problem has a solution and the sum computes it.
Add noise and the picture changes in a specific place. The noise is not smooth — it has no reason to concentrate in the leading singular directions, and it does not — so its coefficients are roughly equal in every direction, at about ‖e‖/√n each. The signal’s coefficients fall exponentially. So below some index the coefficient is no longer the signal’s; it is the noise’s, and it stops falling.
Measured with 0.1% noise on 64 points:
| k | |uₖᵀb| exact | |uₖᵀb| noisy | σₖ | ratio |
|---|---|---|---|---|
| 20 | 7.4·10⁻³ | 7.5·10⁻³ | 5.3·10⁻² | 0.14 |
| 26 | 6.8·10⁻⁴ | 1.1·10⁻⁴ | 7.6·10⁻³ | 0.01 |
| 30 | 3.0·10⁻⁵ | 1.4·10⁻⁴ | 1.6·10⁻³ | 0.09 |
| 34 | 6.8·10⁻⁷ | 2.5·10⁻⁴ | 2.7·10⁻⁴ | 0.92 |
| 38 | 1.5·10⁻⁶ | 1.2·10⁻³ | 3.8·10⁻⁵ | 31.3 |
At k = 20 the exact and noisy coefficients agree to two figures — the signal dominates there. By k = 34 the exact coefficient has fallen to 6.8·10⁻⁷ and the noisy one is 2.5·10⁻⁴, a factor of 370 apart: everything being measured at that index is noise. And the last column is the term’s own contribution to the sum, |uₖᵀb|/σₖ, which is 0.14 at the top and 31.3 at k = 38.
That last number is the whole problem in one figure. The term at k = 38 contributes thirty-one times more to the answer than the term at k = 20, and every part of it is noise.
Finding the crossing, and one way that does not work
The crossing has to be located without knowing which coefficients are the signal’s, so it is found as the minimum of the smoothed coefficient curve: below the noise floor the curve cannot fall further, so its minimum is where the floor begins.
The smoothing is over a window of four either side and it is not decoration. A single |uₖᵀb| can be small by accident — the coefficient of a component the signal happens not to contain — and a crossing detected on one term is a crossing detected on an accident.
The first version used the obvious rule instead: the first index at which the smoothed curve stops falling. It returned 6 at every noise level from 1% to 0.001%.
A detector that is perfectly stable across four orders of magnitude of the thing it is supposed to be detecting is not a detector. What it had found was a shoulder in the signal’s own coefficient curve, at k ≈ 5, where the two smooth bumps stop contributing and the step takes over — a real feature of the signal and nothing to do with noise. The argmin has no such failure mode, because the floor is genuinely the smallest the curve gets.
That is worth recording because the failure looked like success. Six is a plausible truncation for this problem, the number came out of a defensible-looking rule, and only running the rule at four noise levels showed it measuring the wrong thing.
The argmin survives the other end of the range too, which is the harder direction. At 10% noise the floor is so high that only twenty-one components sit above it, and a detector calibrated on clean data has very little curve left to work with.
And it overshoots
Now the finding. Measured at every stop of the slider, against the truncation that actually minimises the error — knowable only because the problem was constructed:
| noise | Picard crossing | best truncation | gap | best relative error |
|---|---|---|---|---|
| 10% | 21 | 14 | +7 | 0.1630 |
| 1% | 32 | 21 | +11 | 0.1445 |
| 0.1% | 32 | 28 | +4 | 0.1058 |
| 0.01% | 40 | 32 | +8 | 0.0975 |
| 0.001% | 45 | 38 | +7 | 0.0952 |
| 0.0001% | 46 | 44 | +2 | 0.0764 |
The crossing is later than the best truncation at every level. Not once, not by a random amount — by two to eleven indices, always in the same direction.
The gap is not monotone in the noise, and the reason is in the two columns beside it. Read the best truncation down the table and it is a ramp: 14, 21, 28, 32, 38, 44 — six or seven indices a decade, five times out of five. Read the crossing and it is a staircase: 21, 32, 32, 40, 45, 46, which advances by 11, then by nothing at all, then by 8, 5 and 1. A whole decade of noise between 1% and 0.1% does not move it — same index, and the same σ = 6.79·10⁻⁴ underneath it.
So the gap is a smooth thing minus a lumpy one, and its sawtooth is the lumpiness rather than anything about the problem. That matters for the practical reading below: a correction fitted to the gap is being fitted partly to where an integer happened to land.
Which is right, and is obvious once measured and not before. The crossing marks where the data stops carrying signal at all. The last few components before it carry signal and noise in comparable amounts, and a component whose signal-to-noise is near one contributes more error than information to the sum — because the contribution is divided by σ, and σ down there is small enough to make a modest error into a large one.
So there are two boundaries and they are different indices:
The boundary of the information, which is the crossing, and which the data knows. The boundary of the usefulness, which is earlier, and which the data does not know.
A rule of thumb that finds the first and stops there has found something real and has answered a different question.
The floor is the noise, checked
The crossing is only a measurement of the noise if the level it sits at is the noise level, so that is asserted rather than assumed. The noise is spread roughly equally over 64 directions, so each coefficient of it should be about ‖e‖/√n, and the smoothed coefficient at the crossing should be within an order of magnitude of that.
It is, and not only at the levels drawn here. The comparison is an assertThat inside the generator
rather than a number checked once while writing, so it re-runs on every call — every figure above,
and every stop the slider passes through on its way between them. So does the overshoot: the
generator refuses to draw at all unless the crossing it found is later than the best truncation.
That is the check which distinguishes this detector from the one it replaced. The shoulder-finder returned 6 at every noise level and the coefficient there had nothing to do with ‖e‖ — an assertion of this shape would have rejected it on the first frame.
It is worth being exact about what those two assertions do and do not establish, because one of them is contradicted later in this essay. They are statements about this problem and this seed, checked everywhere on the slider. The overshoot assertion holds across the whole range here and fails for about one seed in forty, which the sweep below measures; it is calibrated rather than universal, and if it ever fires it will be reporting that the seed changed rather than that the mathematics did.
What the exact right-hand side is for
Three separate things in this essay need a b with no noise in it, and none of them could be done with measured data.
Establishing that the condition holds at all. Without the exact coefficients there is no way to say that the flattening is the noise rather than the problem — a right-hand side that violated the Picard condition before any noise was added would be an inconsistent system, and would look identical from the data.
Locating the true crossing. The gap of two to eleven indices is between two things, and one of them requires the answer.
And refusing a right-hand side that is pure noise. assertTheRegularisationAssertionsReject feeds
the Picard machinery a random vector and requires it to report no usable crossing. A right-hand side
that satisfies nothing must not be reported as satisfying something, and a detector that finds a
crossing in noise would find one anywhere.
This is exact ground truth doing what it does elsewhere on this site — the Hilbert matrix’s rational inverse, the discrete Laplacian’s closed-form spectrum, the Toeplitz family’s tridiagonal inverse. What is unusual here is that the constructed problem is not a convenience: the entire field is about a quantity nobody has, and without one problem where somebody does, there would be nothing to say.
What this says about numerical rank
Rank is a decision establishes that a floating-point matrix does not have a rank, that a numerical rank is a decision about a gap, and that the gap is worth printing beside it.
This field is the case where there is no gap, and the Picard picture is what replaces it. A matrix with a rank has a cliff in its singular values and a threshold anywhere in the cliff gives the same answer. A matrix like this one has no cliff, and what decides the truncation is not the matrix at all — it is the right-hand side, through the coefficients uₖᵀb, and the noise in it.
Which is a genuinely different statement. The rank of a matrix, however decided, is a property of the matrix. The right truncation here changes with the noise level of the data, on the same matrix: 21 at 1% and 38 at 0.001%. Two people with the same operator and differently-measured data should truncate differently, and neither of them is looking at a rank.
The two boundaries, as a signal-to-noise statement
The gap has a clean reading and it is worth having, because “the crossing overshoots” sounds like a defect in the detector rather than a fact about the problem.
Term k contributes (uₖᵀb/σₖ)vₖ to the answer. Split the coefficient into its signal part sₖ and its noise part eₖ, and the term contributes sₖ/σₖ of signal and eₖ/σₖ of error. Keeping the term is worth doing when the first is larger than the second, which is when
|sₖ| > |eₖ|
— a comparison that has nothing to do with σₖ, since it divides both. So the right truncation is where the signal’s own coefficient falls below the noise’s.
The crossing, by contrast, is where the total coefficient stops falling, which happens when the noise part starts to dominate the total — that is, when |eₖ| exceeds |sₖ| by enough to flatten the curve. Detecting a flattening needs the noise to be several times the signal, not merely equal to it, and the difference between “equal” and “several times” is the two to eleven indices in the table.
Both boundaries are therefore about the same crossing of the same two quantities, measured with different sensitivity. One is available and blunt; the other is sharp and requires knowing sₖ, which is knowing the answer.
Where this leaves the practical advice
Three positions are defensible and the field takes the third.
Use the crossing. It is computable, it is a genuine measurement, and it costs a factor of about 1.15 in error on this problem at 0.1% noise. That is not nothing and it is not much.
Correct it by a fixed amount. Subtract five, say. The gaps are +11, +4, +8, +7 — consistent in direction and not in size — so a fixed correction is right on average and wrong at every individual noise level, and there is no reason to expect the average to transfer to another problem.
Or use the crossing as a bound and choose inside it by some other rule. Which is what any real code does, and is why choosing without knowing exists: the crossing says where the search should stop and something else has to say where inside it to land.
The essay reports the bias rather than subtracting it, for the reason this collection reports the looseness of the conjugate gradient bound and the 5.43× slack in the randomised one rather than tuning either. A correction fitted on one problem is a fact about that problem wearing the clothes of a method.
Why the crossing moves and the singular values do not
One more reading of the table, because it is the sentence a practitioner needs.
The singular values are fixed. They are a property of the blur, they do not depend on what was measured, and they are the same at every row of the table. What moves across those rows is the noise floor, and it moves because the measurement got better.
So the number of usable components is not a property of the instrument’s optics. It is a property of the optics and the exposure, and buying two more digits of measurement precision buys about six more usable singular directions here — 32 at 0.1% noise and 45 at 0.001%.
And the returns on that purchase are poor, which is the whole meaning of the word ill-posed. Five decades of measurement precision — 10% noise down to 0.0001% — move the best relative error from 0.1630 to 0.0764. A million-fold better instrument buys a factor of 2.1 in the answer. The components being bought are the ones with the smallest σ, so each is divided by a smaller number and each carries proportionately more of whatever error is left; the sum converges towards the truth about as slowly as it is possible to converge and still be converging.
That is the practical content of the Picard picture and it is why it is worth plotting even though it overshoots. A reader who sees σ falling exponentially and concludes that the problem is hopeless has read half of it; the other half is |uₖᵀb|, and where those two curves separate is a decision that was made when the data was collected rather than when the solve was written.
The gate this field does not have
Worth being explicit about a limitation, because the site’s habit is to state where a claim was checked.
Every number in this essay is measured on one problem: one blur width, one signal, one noise distribution, one seed. The seeds are fixed so the figures are byte-identical on every build, which is this collection’s rule everywhere — but a fixed seed is not the same as a claim about typical behaviour, and the randomised field is the site’s standing reminder that one draw is an anecdote.
The overshoot’s direction was argued from the mechanism above, on the reasoning that it would hold on any problem where the noise is spread evenly over the singular directions. Its size — two to eleven indices across the slider, and four to eleven over the four levels the sweep below covers — was a measurement on this one problem. That gap was recorded here rather than concealed, and it has since been closed by running the same measurement over seeds. Neither half came out the way the recording implied.
Gap between the crossing and the best truncation, 24 seeds at each noise level:
| noise | min | median | max | this essay’s seed |
|---|---|---|---|---|
| 10⁻² | 2 | 18 | 42 | 11 |
| 10⁻³ | 4 | 25 | 35 | 4 |
| 10⁻⁴ | 5 | 22 | 31 | 8 |
| 10⁻⁵ | −2 | 18 | 25 | 7 |
The size does not transfer, and the draw used here is among the smallest. Four to eleven, over these four levels, becomes 2 to 42 — a factor of twenty — and the median at every level is two to five times what this problem showed. The original measurement is reproduced exactly by its own seed, so nothing was computed wrongly; it was one draw, and an unrepresentatively mild one.
And the direction is not universal. Over 800 seed–noise combinations, twenty have a gap of zero or less: one draw crosses at index 37 against a best truncation of 39, another at 36 against 40. Two and a half per cent of the time the crossing arrives at or before the optimum, which is the opposite of what the mechanism was taken to guarantee. The mechanism argument is a statement about expectations, and the crossing is an integer index read off a single realisation — the same distance between an average and an instance that a bound that holds with probability is about.
That is also the answer to a question this essay leaves open earlier: the crossing is not a decision rank is a decision would recognise, because it is read off the data rather than chosen against a stated tolerance, and a quantity read off one draw carries the draw’s variance into the verdict.
So the honest form is stronger in one direction and weaker in the other. The overshoot is real, typical, and larger than this problem shows; and it is not an invariant. A rule that stops at the crossing is usually a few dozen indices late and is occasionally one or two early, and code that treated “never earlier than optimal” as a guarantee would be relying on something that fails about once in forty.
One more thing the sweep shows that a single draw hides: the median gap is nearly flat in the noise — 18, 25, 22, 18 across four decades — where this problem’s numbers suggest it shrinks as the noise falls. The gap at 10⁻² and the gap at 10⁻⁵ correlate at 0.44 across seeds, so a large part of it is a property of the particular noise realisation rather than of its size.
assertTheOvershootIsTypicalButNotGuaranteed measures the whole table, requires this problem’s seed
to reproduce the recorded range, requires the median to exceed it at every level, and requires the
scan to find the draws where the direction fails.
What the figure draws that the table cannot
Three curves on one pair of axes, and the arrangement is the argument.
Plotted alone, the singular values are a straight line on a logarithmic axis falling to rounding — alarming and uninformative, since every ill-posed problem’s do. Plotted alone, the noisy coefficients are a curve that falls and then flattens, which could be anything.
Plotted together with the exact coefficients between them, the picture says the thing no single curve does: the exact coefficients fall faster than σ and the noisy ones do not, and the index where the second stops tracking the first is the index where the problem stops being solvable. Two of those three curves are unavailable in practice and the figure needs all three, which is the whole reason this field’s problem is constructed rather than measured.
The two vertical lines are the essay. One is where the data says to stop; one is where stopping is actually best. Everything between them is a component that can be resolved and should not be kept.
What is left
Choosing λ rather than K, which is choosing without knowing’s subject. The Picard crossing is a truncation index; Tikhonov’s parameter is continuous, and the three published rules for choosing it use different information and get different answers.
Correcting the overshoot. The gap of two to eleven is consistent in direction and not in size, and nothing here turns it into a rule. A correction would have to know how the signal’s coefficients decay relative to the noise’s, which is most of the way to knowing the answer — so the honest position is that the crossing is a bound rather than an estimate, and the essay reports the bias rather than subtracting it.
And correlated noise. Everything above assumes the noise is spread equally over the singular directions, which is what makes the floor flat. Noise with structure — a systematic instrument error, a drift — concentrates in particular directions and does not produce a flat floor at all, so the crossing is not a crossing and the whole detector has nothing to detect. That is the realistic case and it is not measured here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An answer that changes with the seed — both name filter factors, ill-posed problem, truncated svd
- Four knobs and one floor — both name filter factors, ill-posed problem, truncated svd
- The corner reads the norm it is drawn in — both name filter factors, noise floor, regularisation
- A nearest point that is not there — both name ill-posed problem, truncated svd
- A rule that has to be told how good its answer will be — both name filter factors, regularisation
- A second penalty is not a second parameter — both name filter factors, regularisation
Named objects
A flat tag is an object no other essay names yet.
Filter factorsIll-posed problemNoise floorPicard conditionRegularisationSignal to noiseSingular value decompositionTruncated SVD