01 / A striking headline result
An estimator can classify a broad set of synthetic forests extremely well.
The first experiment creates 324 mathematically generated forest worlds. The experimenter knows the underlying state and demographic parameters, so reference Health classifications can be calculated. The estimator instead receives noisy or incomplete observations and must reconstruct the state, estimate the dynamics, and classify prospective Health.
State and generating parameters are controlled
The estimator cannot see the hidden truth
Reconstruct, simulate, and compare
Balanced accuracy · 324 synthetic worlds
What this result includes
69 reference-positive and 255 reference-negative cases. The primary reference viable masses are all more than 0.10 away from the 0.75 threshold.
The test is useful for internal model-based verification, particularly on cases with margins large enough to make the decisions easier.
The result is a property of this controlled synthetic population and estimator. It is not a measured diagnostic accuracy for the real Harvard Forest.
02 / Understand the metric
Balanced accuracy gives positive and negative cases equal weight.
Sensitivity asks how often reference-positive cases are correctly classified. Specificity asks how often reference-negative cases are correctly classified. Balanced accuracy averages these two rates.
Balanced classification accuracy
The mean of sensitivity and specificity.
For example, 98% sensitivity and 90% specificity produce balanced accuracy (0.98 + 0.90)/2 = 0.94, or 94%. Balanced accuracy prevents an imbalanced test set from rewarding majority-category predictions too heavily.
The metric helps us compare two kinds of correct answer. It does not make the test population representative of the hard decisions we eventually care about. In particular, a test with few examples near its decision boundary cannot establish precision at that boundary.
03 / Deliberately difficult conditions
What happens when we deliberately place worlds near 75%?
In Lesson 15, Q₁ asks whether viable mass reaches θ = 0.75. A viable mass of 0.90 passes easily and 0.40 fails easily. But 0.749 and 0.751 lie on opposite sides despite differing by just 0.002—two-tenths of one percentage point.
Paper V constructs 81 synthetic worlds: 72 designed around specified capacity targets and nine present-realization controls. Among the 72 capacity-targeted worlds, 67 end up within 0.10 of the threshold and 23 within 0.02. This is a more demanding designed population.
324 worlds · broad-range
31 certified, presently realizing cases
46 certified, presently realizing cases
Source: Paper V, Section 6.6. The Q₁ and Q₂ near-boundary subsets differ from one another and from the 324-world broad-range experiment. The figures compare performance across different designed test populations, not repeated measurements on one identical sample.
These results can all be correct simultaneously. Broad-range accuracy can be excellent while near-threshold classification is unreliable in the chosen experimental conditions. The reported approximately 56% values do not establish a universal accuracy for real-world health measurement.
04 / Where a tiny error becomes a different answer
A small probability error can flip a Boolean Health judgment.
Q₁ at the declared threshold
At 75%, the decision changes sides.
The equality belongs to the passing side. Thus 0.749 fails and 0.751 passes, even though the two probabilities are close.
Suppose the underlying probability in a specified model is 0.751 while the estimator reports 0.748. The numerical error is only −0.003, or −0.3 percentage points. Nonetheless, the estimated decision is false while the reference decision is true.