01 / An out-of-fitting-period evaluation
Can previously fitted mortality hazards predict later adult-tree deaths?
The survival model uses estimated mortality hazards for different forest cohorts. The question now is whether those hazards predict later outcomes that were unavailable when the model was fitted. This is a direct comparison against a separate observational record.
Paper V uses the Harvard Forest HF453 adult-tree health and mortality survey. It selects adult stems observed alive at a 2021 baseline and matches eligible observations through 2024, subject to explicit stem identity, size, and status restrictions.
This is a retrospective out-of-fitting-period assessment. The later mortality outcomes were not used to fit the earlier hazards, but the predictions were evaluated retrospectively; they were not a publicly issued forecast made before 2021–2024.
The risk set is conditional on surviving and being recorded alive in 2021, then meeting the study's matching requirements. That limits which broader forest populations the findings can describe.
02 / What did the model predict?
Most of the discrepancy appears in one tree group.
| Adult-tree group | Model-expected deaths | Observed secure deaths |
|---|---|---|
| Hemlock | 280.3 | 685 |
| Other taxa | 769.9 | 792 |
| All adults | 1,050.3 | 1,477 |
There are approximately 405 more hemlock deaths than the model's expected count. For the other tree groups, approximately 770 deaths were expected and 792 observed. The deviation therefore has a pronounced taxon-specific pattern.
Expected deaths are not deterministic predictions that exactly that many trees must die. They are the average count implied by the fitted predictive model. To interpret the discrepancy scientifically we need its predicted variability, as well as the mean.
03 / Compare the proportions
Different population sizes call for a mortality-rate comparison.
The hemlock and other-taxa groups contain different numbers of adults. Converting counts to cumulative 2021–2024 mortality rates helps put them on a comparable scale.
For hemlock, the observed cumulative mortality rate of approximately 9.06% is about 2.4 times the model's 3.71% prediction. For other taxa the rates are much closer, at roughly 5.5–5.6%.
These are cumulative percentages over the selected observation interval, not annual mortality rates. The nonhemlock values are rounded here from the Table 8 counts and group size.
04 / How much variability did the model expect?
The observed hemlock mortality lies well outside the model's predicted range.
A statistical model predicts a distribution of death counts, not only its average count. Under the fitted hazard model the central 95% predictive intervals are:
| Group | Expected deaths | Central 95% predictive interval | Observed deaths |
|---|---|---|---|
| Hemlock | 280.3 | 220–347 | 685 |
| Other taxa | 769.9 | 654–896 | 792 |
| All adults | 1,050.3 | 918–1,192 | 1,477 |
Both horizontal scales run from 0 to 1,000 deaths, allowing comparison of interval width and observed position. The model's interval already incorporates the specified fitted-hazard uncertainty.
The observed hemlock count is nearly twice the upper endpoint of its central 95% interval. The other-taxa count falls inside its reported interval. The model's predictive disagreement is strongest for hemlock.
05 / Quantifying the unusual observation
Under the fitted hemlock model, 685 or more deaths has an extraordinarily small probability.
Hemlock upper-tail predictive probability
What is the chance of at least 685 deaths under the specified model?
D is the total death count from the fitted predictive distribution for the selected hemlock stems. The event D ≥ 685 includes 685 or more deaths. The probability is conditional on that predictive model and risk set.
To see the scale of the expression, remember that 10 raised to a negative exponent is a fraction:
Reading scientific notation
A small positive tail probability.
The factor 4.6308 multiplies the reciprocal of 1021. The result is far smaller than conventional significance thresholds.
It is strong evidence that the tested mortality model does not adequately predict these selected hemlock outcomes. This number is not a probability that the model is true or false, and it is not a probability that the mathematical definition of health is correct.
Statistical evidence tells us that the observed outcome is exceptionally implausible under the specified predictive distribution; scientific explanation requires a separate analysis of the generating process.
06 / Accounting for uncertainty within the model
The death-count prediction includes both event variation and fitted-hazard uncertainty.
Each model cell has a fitted hazard distribution rather than a perfectly known hazard. In a predictive replicate, stems assigned to the same cell share the sampled hazard. The uncertainty in the predicted total therefore comes from two sources.
Individual stem outcomes still differ between draws.
Different sampled cell hazards shift the expected deaths together.
Law of total variance
Total predictive variance contains both sources.
H denotes the collection of sampled cell hazards. The first term averages outcome variance conditional on fixed hazards; the second captures how the conditional mean changes when hazard values vary.
In the reported hemlock calculation, about 74.6% of the predictive variance comes from fitted-hazard uncertainty. The remainder is event variation.
Even after accounting for these modeled uncertainties, 685 deaths remains far outside the predictive distribution. This also leaves open whether the uncertainty law or its conditioning on 2021 survivors appropriately represents the real observational population.
07 / Why many simulation draws are still insufficient
Zero simulated exceedances do not establish a zero underlying probability.
The study generates 262,144 predictive mortality replicates. None reaches 685 hemlock deaths. The direct simulation proportion of exceedances is thus 0/262,144—but that is an observation about this finite simulation experiment.
Replicates with at least 685 hemlock deaths in this finite run
A small positive upper-tail probability under the stipulated fitted distribution
Events this rare can be absent from a finite simulation even when their theoretical probability is positive. The paper therefore evaluates the extreme upper tail numerically and reports independent numerical checks rather than declaring the probability to be zero.
A simulation can characterize the middle of a distribution well while having insufficient resolution for its far tail. Exact or independently checked numerical probability methods can answer a different question about that distribution.