01 / The problem comes before prediction
Prospective health begins with a present state. How do we know that state?
We've returned to this definition throughout the series: a system must presently realize its organization and have sufficient capacity for viable continuation under the specified scenario and horizon.
The prospective-health definition
One equation; a newly difficult input.
Here x is the present state, d a scenario, and t a horizon. In the five-state forest from Paper IV, I could declare x to be baseline, stressed, recovered, adapted, or collapsed. An empirical forest is not supplied as a perfectly identified mathematical state.
We have to infer it from recorded stem identities, measurements, dates, apparent survival, and missing information. Errors or gaps in that inference flow into everything we later say about continuation.
02 / The observed site
A 35-hectare forest, observed over many years rather than on one day.
The empirical study uses the Harvard Forest ForestGEO plot in Massachusetts, a 700 × 500 meter research area. Researchers recorded large numbers of woody stems during two census periods, but the observations were not all made at the same instant.
Earlier census observations, with stem identities and measurements recorded on individual dates.
Later records, including changes, missing diameters, and outcomes that remain unresolved.
The common reference date for reconstructed forest states and future simulations.
Measurements taken years apart do not describe the same point in time. Before a future simulation can begin on January 3, 2020, the method must reconcile differing observation dates and reconstruct a plausible common-time state.
The research footprint and dates describe the source study, not a new field survey undertaken for this lesson.
03 / What was actually recorded?
The record is extensive. It is also incomplete in consequential ways.
| Observation category | Records |
|---|---|
| E0 stem records | 116,227 |
| E1 stem records | 123,218 |
| Living stems in E0 | 108,632 |
E1 records marked stem_gone | 37,577 |
| E1 records without a diameter measurement | 61,985 |
| Living stem IDs recorded only in E1 | 6,992 |
These counts describe different record categories and should not be added together as though they were mutually exclusive groups of trees. Each matters for a different reason.
stem_gone is not proof of biological death. A missing diameter introduces uncertainty in physical structure. A stem appearing only in E1 may be a genuine entrant, a previously missed stem, or another identity-recording outcome. The raw census is evidence about the forest, rather than a complete, error-free state description.
04 / Observed evidence and unobserved state
I need to distinguish what was seen from what I am inferring.
Let Y contain the two census records. Let X denote the modeled forest state at our common starting date. In statistical language, X is a latent state: the state of interest is not fully and directly observed.
The data and the target state
Observations are not identical to the latent forest state.
The symbol Y denotes recorded evidence. The symbol X denotes the forest state we seek to reconstruct from that evidence. We require a reconstruction law rather than simply asserting X = Y.
Imagine a purely illustrative census of 100 trees. Seventy are confirmed alive later, ten are known dead, and twenty have unresolved fates. Neither “only seventy survive” nor “all twenty unresolved trees survive” follows from those observations.
Explore a declared assumption
How much can our estimate change while the observed records stay fixed?
Suppose we assign the twenty unresolved trees the same modeled survival probability. Move the slider to see the expected number surviving, not a directly observed number of survivors.
The numbers are invented to explain conditional state reconstruction. The expected count need not be an integer; the actual outcome remains uncertain. The independence or common-probability assumptions used by a given model need separate justification.
Our data have not changed when the slider moves. The statistical assumption has. That difference is exactly why the resulting inferred state cannot be treated as a direct observation.
05 / Compressing the forest
Hundreds of thousands of records become 64 cohort cells.
Following each individual stem through every demographic mechanism would require an extremely detailed system. The paper defines a computational representation with three grouping dimensions, four groups in each dimension.
Each cohort cell c stores a stem count nc and a root-mean-square diameter dc. One spatial/protocol category is a wet-candidate proxy; the study does not independently validate a corresponding hydrological classification.
As we saw in Paper III, compression must be judged against the questions the representation will later need to answer. A summary that retains one useful quantity may erase another.
06 / What the compression preserves
Root-mean-square diameter keeps the basal-area calculation intact.
Tree basal area scales with the square of trunk diameter. For diameters d1 through dn, the root-mean-square (RMS) diameter is the square root of the average squared diameter:
Root-mean-square diameter
Square, average, then take the square root.
Unlike the ordinary average, this summary preserves the sum of squared diameters across the stems in each cohort.
With diameter measured in centimeters and basal area expressed in square meters, the cohort-state basal area is:
Paper V · Section 4.1
Total stem basal area from 64 cohorts.
nc is the number of stems in cohort c. dc is its RMS diameter in centimeters. The factor π/40,000 converts diameter squared into circular basal area in square meters.
This exact aggregation protects the basal-area calculation, given the grouped stem values. It does not recover every original diameter. Two cohorts may share nc and dc while containing different counts of stems below the study's 10-centimeter continuation threshold. That limitation becomes important later in Paper V.
07 / Different stems, different observation windows
A modeled annual mortality hazard converts elapsed time into a survival probability.
Because individual stems were observed on different dates, each stem has its own elapsed observation interval. Paper V uses a conditional constant-hazard survival model.
Paper V · Section 4.3, Eq. (30)
Survival over an observation interval.
Zi = 1 denotes survival of stem i, hc is its cohort's annual mortality hazard, and Δi is elapsed time in years. The bar means “conditional on.” The exponential is e raised to the negative hazard multiplied by the elapsed time.
For a positive constant hazard, a longer interval yields a lower modeled probability of survival. This allows the reconstruction to respect actual observation dates rather than treating all intervals as equally long.
It remains a model assumption. The real hazard is not known in advance, and a constant-hazard law is not established by writing the formula. The method must estimate its parameters and test how the resulting predictions behave.
08 / Sampling possible present states
Instead of one certain present, we generate possible present forests.
Paper V · Section 4.5, Eq. (50)
The conditional reconstruction law.
Krec is the reconstruction kernel; Y is the observed record; ϑ denotes the model parameters. Drawing X from the kernel produces one possible common-time forest state conditional on the records and assumptions.
Measurements, identities, dates, and missing entries
Hazard, growth, entry, and observation assumptions
In the primary computational design, the paper samples 256 reconstructed states. Combining those draws with parameter variations and two growth-model variants gives 1,024 outer computational units, each with 64 simulated future histories per scenario. These are conditional computational units, not 1,024 separately observed real forests.
Every possible future begins downstream of this uncertain starting point. The simulation pipeline can be executed exactly as designed while the scientific accuracy of its assumptions remains an open question.