Skip to lesson content

Lesson 17 · A test against independent observations

What happens when the mathematical model is confronted with real-world evidence?

We have defined and computed model-conditional prospective Health. Now I want to ask a more demanding question: how well do the forest mortality rates fitted from earlier observations predict what was later observed? One adult-tree group reveals an important discrepancy.

By Zed JamesPaper V · Sections 5.8, 6.5 and 7.1Independent mortality comparison

01 / An out-of-fitting-period evaluation

Can previously fitted mortality hazards predict later adult-tree deaths?

The survival model uses estimated mortality hazards for different forest cohorts. The question now is whether those hazards predict later outcomes that were unavailable when the model was fitted. This is a direct comparison against a separate observational record.

Paper V uses the Harvard Forest HF453 adult-tree health and mortality survey. It selects adult stems observed alive at a 2021 baseline and matches eligible observations through 2024, subject to explicit stem identity, size, and status restrictions.

Matched adult risk set21,576Adult stems with eligible later observed endpoints
Eastern hemlock7,557Stems in the hemlock comparison
Other taxa14,019All remaining selected adult stems

This is a retrospective out-of-fitting-period assessment. The later mortality outcomes were not used to fit the earlier hazards, but the predictions were evaluated retrospectively; they were not a publicly issued forecast made before 2021–2024.

The risk set is conditional on surviving and being recorded alive in 2021, then meeting the study's matching requirements. That limits which broader forest populations the findings can describe.

02 / What did the model predict?

Most of the discrepancy appears in one tree group.

Paper V, Table 8 (printed page 33) · adult deaths from the 2021–2024 selected risk set
Adult-tree groupModel-expected deathsObserved secure deaths
Hemlock280.3685
Other taxa769.9792
All adults1,050.31,477
Hemlock · expected deaths280.3Previously fitted model
Hemlock · observed deaths685Secure endpoints in the later record

There are approximately 405 more hemlock deaths than the model's expected count. For the other tree groups, approximately 770 deaths were expected and 792 observed. The deviation therefore has a pronounced taxon-specific pattern.

Expected deaths are not deterministic predictions that exactly that many trees must die. They are the average count implied by the fitted predictive model. To interpret the discrepancy scientifically we need its predicted variability, as well as the mean.

03 / Compare the proportions

Different population sizes call for a mortality-rate comparison.

The hemlock and other-taxa groups contain different numbers of adults. Converting counts to cumulative 2021–2024 mortality rates helps put them on a comparable scale.

PredictedObserved
Hemlock
Expected · 3.71%
Observed · 9.06%
Other taxa
Expected · ~5.49%
Observed · ~5.65%

For hemlock, the observed cumulative mortality rate of approximately 9.06% is about 2.4 times the model's 3.71% prediction. For other taxa the rates are much closer, at roughly 5.5–5.6%.

These are cumulative percentages over the selected observation interval, not annual mortality rates. The nonhemlock values are rounded here from the Table 8 counts and group size.

04 / How much variability did the model expect?

The observed hemlock mortality lies well outside the model's predicted range.

A statistical model predicts a distribution of death counts, not only its average count. Under the fitted hazard model the central 95% predictive intervals are:

Paper V, Table 10 (printed page 35) · predictive death counts
GroupExpected deathsCentral 95% predictive intervalObserved deaths
Hemlock280.3220–347685
Other taxa769.9654–896792
All adults1,050.3918–1,1921,477
Mortality predictive distributionPredicted central interval versus observed count
Hemlock
220–347 vs 685
Other taxa
654–896 vs 792
Central 95% rangeExpectedObserved

Both horizontal scales run from 0 to 1,000 deaths, allowing comparison of interval width and observed position. The model's interval already incorporates the specified fitted-hazard uncertainty.

The observed hemlock count is nearly twice the upper endpoint of its central 95% interval. The other-taxa count falls inside its reported interval. The model's predictive disagreement is strongest for hemlock.

05 / Quantifying the unusual observation

Under the fitted hemlock model, 685 or more deaths has an extraordinarily small probability.

Hemlock upper-tail predictive probability

What is the chance of at least 685 deaths under the specified model?

P(D≥685)=4.6308×10−21

D is the total death count from the fitted predictive distribution for the selected hemlock stems. The event D ≥ 685 includes 685 or more deaths. The probability is conditional on that predictive model and risk set.

To see the scale of the expression, remember that 10 raised to a negative exponent is a fraction:

Reading scientific notation

A small positive tail probability.

10−21=11021

The factor 4.6308 multiplies the reciprocal of 1021. The result is far smaller than conventional significance thresholds.

It is strong evidence that the tested mortality model does not adequately predict these selected hemlock outcomes. This number is not a probability that the model is true or false, and it is not a probability that the mathematical definition of health is correct.

Statistical evidence tells us that the observed outcome is exceptionally implausible under the specified predictive distribution; scientific explanation requires a separate analysis of the generating process.

06 / Accounting for uncertainty within the model

The death-count prediction includes both event variation and fitted-hazard uncertainty.

Each model cell has a fitted hazard distribution rather than a perfectly known hazard. In a predictive replicate, stems assigned to the same cell share the sampled hazard. The uncertainty in the predicted total therefore comes from two sources.

Event variabilityDeaths vary even with hazards fixed.

Individual stem outcomes still differ between draws.

Hazard uncertaintyEstimated mortality rates themselves vary.

Different sampled cell hazards shift the expected deaths together.

Law of total variance

Total predictive variance contains both sources.

Var(D)=E[Var(D∣H)]+Var(E[D∣H])

H denotes the collection of sampled cell hazards. The first term averages outcome variance conditional on fixed hazards; the second captures how the conditional mean changes when hazard values vary.

Even after accounting for these modeled uncertainties, 685 deaths remains far outside the predictive distribution. This also leaves open whether the uncertainty law or its conditioning on 2021 survivors appropriately represents the real observational population.

07 / Why many simulation draws are still insufficient

Zero simulated exceedances do not establish a zero underlying probability.

The study generates 262,144 predictive mortality replicates. None reaches 685 hemlock deaths. The direct simulation proportion of exceedances is thus 0/262,144—but that is an observation about this finite simulation experiment.

Monte Carlo sample0 / 262,144

Replicates with at least 685 hemlock deaths in this finite run

Direct tail calculation4.6308 × 10−21

A small positive upper-tail probability under the stipulated fitted distribution

Events this rare can be absent from a finite simulation even when their theoretical probability is positive. The paper therefore evaluates the extreme upper tail numerically and reports independent numerical checks rather than declaring the probability to be zero.

A simulation can characterize the middle of a distribution well while having insufficient resolution for its far tail. Exact or independently checked numerical probability methods can answer a different question about that distribution.

08 / Evidence for miscalibration and the search for mechanism

Finding where a model fails does not uniquely identify why it failed.

The adult mortality test identifies a substantial hemlock-specific underprediction. The author discusses hemlock woolly adelgid, an insect associated with eastern hemlock decline, as a plausible mechanism deserving further investigation.

That biological explanation has not been isolated by the matched mortality comparison. Testing it would require evidence on infestation pressure, crown condition, damage history, changing hazards, and independently held-out outcomes. A useful next model would include biologically supported mechanisms and determine whether they improve prospective prediction.

The current result is a strong component-level test of the previously fitted mortality law on a selected adult population. It establishes a predictive calibration problem within that scope.

09 / What exactly does the evidence establish?

The fitted hazard component has failed a demanding observational test.

Three different kinds of scientific claim
ClaimWhat the mortality result addresses
Formal definition of prospective HealthThe discrepancy does not contradict the definition or its mathematical consequences.
Correct execution of declared probability calculationsThe program may implement its probability model correctly even where that model predicts actual observations poorly.
Reliable real-forest mortality predictionThe selected hemlock comparison provides strong evidence against the tested hazard component's predictive calibration.

The complete prospective-Health claim involves present forest reconstruction, future growth, entry, mortality, viability thresholds, and a distribution over entire histories. This adult-mortality comparison tests one important part of that chain.

It cannot independently validate or reject the entire prospective-health measurement pipeline. The 2021 survivor baseline and matching rules also constrain the interpretation: deaths before that baseline and outcomes outside the eligible risk set were not directly tested here.

10 / A question to carry forward

The model's central 95% hemlock interval is 220–347 deaths, and 685 are observed. What can we conclude?

Choose the strongest conclusion supported by the comparison.

11 / Three achievements that require different evidence

A proof, a reproducible calculation, and a reliable measurement answer different questions.

Mathematical verificationDo the conclusions follow from the formal assumptions?

Exact statements, definitions, and proof obligations.

Computational reproducibilityCan we replay the stipulated calculation faithfully?

Inspectability, checksums, code, and numerical execution.

Empirical validityDoes the calculation accurately describe the observed system?

Independent observational evidence and calibrated prediction.

Paper V demonstrates why all three are needed in an empirical measurement program. The hemlock discrepancy is a substantive finding about the predictive system and a guide to what requires better scientific modeling.

What I want you to carry forward

A measured discrepancy can tell us precisely which model component requires new evidence.

In the selected adult forest population, the previously fitted hazards predict far fewer hemlock deaths than were subsequently observed. The discrepancy remains extreme after accounting for the model's specified hazard uncertainty. It tests the empirical reliability of one important mechanism without rewriting the mathematical definition of Health.

Lesson 18 returns to an exact question about information. A cohort representation can preserve stem count, RMS diameter, and basal area while losing the number of stems below a 10-centimeter juvenile threshold. We'll examine why this is a mathematical information-loss result with consequences for the very continuation questions the numerical model tries to answer.

Source: Zed James, Prospective Health under Declared Specifications, Paper V of Health, Formally Defined (2026), especially Sections 5.8, 6.5 and 7.1, Tables 8 and 10, and the predictive hemlock tail calculation. Full publication record · Zenodo DOI. The mortality findings are a component-level retrospective evaluation on a selected adult risk set, not an independently validated real-forest Health classification.