A test at year 18 reads 46% against a threshold of 37%. Observed: pass. But 46% from 50 seeds carries a Wilson interval of roughly 33–60% — the lower edge dips below the threshold. Is the interval "supported"? That's a stricter question than "did the point pass," and it deserves its own flag rather than silently replacing the observation.
The workflow asks it two independent ways. The Wilson check interrogates the single test: does even the pessimistic edge of that one binomial observation clear the bar? The GLM check interrogates the whole series: fit the decline curve, predict at that year with a 95% band (built on the link scale, then transformed), and ask whether the band's floor clears. One can pass while the other fails — a strong series with one noisy test, or one great test on a ragged series — and the disagreement is itself diagnostic.
Practical reading: both flags = act on it; one flag = act, but schedule the next test sooner; neither = it's a single plate of seeds; treat it as a lead.