Skip to lesson content

Lesson 14 · From formal definition to incomplete observations

What happens when we don't actually know the system's present state?

The first four papers worked with explicitly defined mathematical systems. In Paper V, I begin with a real forest whose state was never completely observed. Before we can calculate what futures remain possible, we must face a prior question: what was the forest's state at the time the future begins?

By Zed JamesPaper V · Sections 2–4 and 6.3Beginning empirical reconstruction

01 / The problem comes before prediction

Prospective health begins with a present state. How do we know that state?

We've returned to this definition throughout the series: a system must presently realize its organization and have sufficient capacity for viable continuation under the specified scenario and horizon.

The prospective-health definition

One equation; a newly difficult input.

Health(x;d,t)⇔Realizes(x)∧Adequate(Capacity(x;d,t))

Here x is the present state, d a scenario, and t a horizon. In the five-state forest from Paper IV, I could declare x to be baseline, stressed, recovered, adapted, or collapsed. An empirical forest is not supplied as a perfectly identified mathematical state.

We have to infer it from recorded stem identities, measurements, dates, apparent survival, and missing information. Errors or gaps in that inference flow into everything we later say about continuation.

02 / The observed site

A 35-hectare forest, observed over many years rather than on one day.

The empirical study uses the Harvard Forest ForestGEO plot in Massachusetts, a 700 × 500 meter research area. Researchers recorded large numbers of woody stems during two census periods, but the observations were not all made at the same instant.

E0June 2010 – March 2014

Earlier census observations, with stem identities and measurements recorded on individual dates.

E1May 2018 – January 2020

Later records, including changes, missing diameters, and outcomes that remain unresolved.

Modeled originJanuary 3, 2020

The common reference date for reconstructed forest states and future simulations.

Measurements taken years apart do not describe the same point in time. Before a future simulation can begin on January 3, 2020, the method must reconcile differing observation dates and reconstruct a plausible common-time state.

The research footprint and dates describe the source study, not a new field survey undertaken for this lesson.

03 / What was actually recorded?

The record is extensive. It is also incomplete in consequential ways.

Selected reported data, Paper V, Table 1 (printed page 3)
Observation categoryRecords
E0 stem records116,227
E1 stem records123,218
Living stems in E0108,632
E1 records marked stem_gone37,577
E1 records without a diameter measurement61,985
Living stem IDs recorded only in E16,992

These counts describe different record categories and should not be added together as though they were mutually exclusive groups of trees. Each matters for a different reason.

stem_gone is not proof of biological death. A missing diameter introduces uncertainty in physical structure. A stem appearing only in E1 may be a genuine entrant, a previously missed stem, or another identity-recording outcome. The raw census is evidence about the forest, rather than a complete, error-free state description.

04 / Observed evidence and unobserved state

I need to distinguish what was seen from what I am inferring.

Let Y contain the two census records. Let X denote the modeled forest state at our common starting date. In statistical language, X is a latent state: the state of interest is not fully and directly observed.

The data and the target state

Observations are not identical to the latent forest state.

Y=(YE0,YE1)X≠Y

The symbol Y denotes recorded evidence. The symbol X denotes the forest state we seek to reconstruct from that evidence. We require a reconstruction law rather than simply asserting X = Y.

Imagine a purely illustrative census of 100 trees. Seventy are confirmed alive later, ten are known dead, and twenty have unresolved fates. Neither “only seventy survive” nor “all twenty unresolved trees survive” follows from those observations.

Explore a declared assumption

How much can our estimate change while the observed records stay fixed?

Illustration only

Suppose we assign the twenty unresolved trees the same modeled survival probability. Move the slider to see the expected number surviving, not a directly observed number of survivors.

Confirmed alive70
Expected alive among 20 unresolved16
Expected alive out of 10086

The numbers are invented to explain conditional state reconstruction. The expected count need not be an integer; the actual outcome remains uncertain. The independence or common-probability assumptions used by a given model need separate justification.

Our data have not changed when the slider moves. The statistical assumption has. That difference is exactly why the resulting inferred state cannot be treated as a direct observation.

05 / Compressing the forest

Hundreds of thousands of records become 64 cohort cells.

Following each individual stem through every demographic mechanism would require an extremely detailed system. The paper defines a computational representation with three grouping dimensions, four groups in each dimension.

Taxon4groups
Space and protocol4strata
Stem size4diameter classes
Computed representation64cohort cells

Each cohort cell c stores a stem count nc and a root-mean-square diameter dc. One spatial/protocol category is a wet-candidate proxy; the study does not independently validate a corresponding hydrological classification.

As we saw in Paper III, compression must be judged against the questions the representation will later need to answer. A summary that retains one useful quantity may erase another.

06 / What the compression preserves

Root-mean-square diameter keeps the basal-area calculation intact.

Tree basal area scales with the square of trunk diameter. For diameters d1 through dn, the root-mean-square (RMS) diameter is the square root of the average squared diameter:

Root-mean-square diameter

Square, average, then take the square root.

dRMS=d12+d22+⋯+dn2n

Unlike the ordinary average, this summary preserves the sum of squared diameters across the stems in each cohort.

With diameter measured in centimeters and basal area expressed in square meters, the cohort-state basal area is:

Paper V · Section 4.1

Total stem basal area from 64 cohorts.

BA(X)=π40000∑c=063ncdc2

nc is the number of stems in cohort c. dc is its RMS diameter in centimeters. The factor π/40,000 converts diameter squared into circular basal area in square meters.

This exact aggregation protects the basal-area calculation, given the grouped stem values. It does not recover every original diameter. Two cohorts may share nc and dc while containing different counts of stems below the study's 10-centimeter continuation threshold. That limitation becomes important later in Paper V.

07 / Different stems, different observation windows

A modeled annual mortality hazard converts elapsed time into a survival probability.

Because individual stems were observed on different dates, each stem has its own elapsed observation interval. Paper V uses a conditional constant-hazard survival model.

Paper V · Section 4.3, Eq. (30)

Survival over an observation interval.

P(Zi=1∣hc,Δi)=e−hcΔi

Zi = 1 denotes survival of stem i, hc is its cohort's annual mortality hazard, and Δi is elapsed time in years. The bar means “conditional on.” The exponential is e raised to the negative hazard multiplied by the elapsed time.

For a positive constant hazard, a longer interval yields a lower modeled probability of survival. This allows the reconstruction to respect actual observation dates rather than treating all intervals as equally long.

It remains a model assumption. The real hazard is not known in advance, and a constant-hazard law is not established by writing the formula. The method must estimate its parameters and test how the resulting predictions behave.

08 / Sampling possible present states

Instead of one certain present, we generate possible present forests.

Paper V · Section 4.5, Eq. (50)

The conditional reconstruction law.

X∼Krec(·∣Y,ϑ)

Krec is the reconstruction kernel; Y is the observed record; ϑ denotes the model parameters. Drawing X from the kernel produces one possible common-time forest state conditional on the records and assumptions.

ObservedCensus records Y

Measurements, identities, dates, and missing entries

Inferred conditionallyKrec(· | Y, ϑ)

Hazard, growth, entry, and observation assumptions

Possible presentsX₁X₂X₃Many alternative common-time states

In the primary computational design, the paper samples 256 reconstructed states. Combining those draws with parameter variations and two growth-model variants gives 1,024 outer computational units, each with 64 simulated future histories per scenario. These are conditional computational units, not 1,024 separately observed real forests.

Every possible future begins downstream of this uncertain starting point. The simulation pipeline can be executed exactly as designed while the scientific accuracy of its assumptions remains an open question.

09 / Confronting reconstruction with held-out observations

The nominal 90% diameter intervals contained the true held-out values only about half the time.

One diagnostic temporarily hides selected recorded information, reconstructs it, and compares the inferred interval with the held-out recorded diameter. This is a masking experiment, and it is a direct test of reconstruction calibration on the observations available for checking.

Interval coverage under four declared masking mechanisms90% nominal target

Displayed values are approximate midpoints of the three-replicate coverage ranges in Paper V, Table 4 (printed page 31), not new estimates. Full reported ranges: M0 49.9–50.3%, M1 46.8–47.4%, M2 49.1–50.3%, M3 55.1–56.0%.

Individual diameters~47–56%

Observed coverage under nominal 90% intervals, across the reported masking families.

Aggregated basal area~92–97%

Reported coverage for a different, aggregated target under the corresponding masking families.

The sharp contrast matters. The procedure performs poorly at this individual-stem task while showing much better interval coverage on aggregated basal area. As Paper III taught us, adequacy and sufficiency depend on the question being asked.

These tests use stems whose information can be held out and later inspected. They do not certify the model's calibration for genuinely unresolved stem fates. That is a separate evidence boundary.

10 / A deeper boundary than numerical uncertainty

The later census cannot by itself tell us how every later-only stem entered the record.

The identity audit finds 6,992 living stem IDs appearing only in E1, and none can be certified from the available two-census identity evidence as an unequivocal biological threshold entrant. This result does not mean that no biological entry occurred.

Explanation AGenuine entry

The stem crossed the inventory's inclusion threshold between censuses.

Explanation BEarlier omission

The stem was already eligible earlier but was not recorded then.

Both histories may produce the same later-only observation. When two underlying explanations give indistinguishable evidence, no deterministic function of that evidence alone identifies which one occurred. This is non-identifiability, and it is the representation problem from Paper III encountered in an empirical study.

We can introduce an assumption-based entry model. We cannot turn an unobserved event into an observed fact by adding a parameter. Additional identifying evidence would be needed.

11 / Test the distinction

One stem is recorded in E1 but is absent from E0. What follows from that fact alone?

Choose the conclusion supported by the observed records.

12 / What kind of knowledge do we have?

There are three different levels of inference.

Separate the evidence from the assumptions used downstream
LevelWhat it means
ObservedRecorded measurements, dates, classifications, and missingness in the two censuses.
ReconstructedA conditional estimate or distribution over possible January 2020 forest states.
SimulatedFuture histories produced by assumed demographic dynamics, scenarios, and sampled starting states.

Each stage inherits the limits of the previous one and introduces additional modeling commitments. Correct execution of the declared mathematics and reliable measurement of the real forest are distinct achievements. Paper V documents both the construction and the points at which empirical evidence does not yet justify a stronger conclusion.

This study does not establish an operational health classification of Harvard Forest. It establishes an auditable path from recorded evidence toward model-conditional continuation estimates, together with substantial calibration and identification limits.

What I want you to carry forward

Before we can measure future capacity, we must identify what we actually know about the present.

The observed census is not the latent forest state. Cohort compression preserves some aggregate quantities and loses others. Reconstruction depends on estimated hazards, missingness assumptions, and imperfectly identified biological events. Holding those distinctions together is how we move from a sound formal definition toward a measurement system that can be tested.

In Lesson 15, we'll begin simulating alternative futures from these possible present states. We'll learn how to assign probability to viable continuations—and why deleting failed futures before recomputing the probability would corrupt the result.

Source: Zed James, Prospective Health under Declared Specifications, Paper V of Health, Formally Defined (2026), particularly Sections 2–4 and 6.3, Tables 1 and 4, and Equations (30) and (50). Publication record · Zenodo DOI. All estimates are conditional on the stated data and model; no operational ecological health verdict is asserted.