Skip to lesson content

Lesson 16 · Classification near the decision boundary

Why a model can be 99% accurate and still struggle when it matters most

We can now calculate model-conditional prospective Health. But how trustworthy is the result when the true viable mass is close to the threshold? Paper V asks this through synthetic experiments, controlled information tests, and deliberately difficult comparisons.

By Zed JamesPaper V · Sections 5.4–5.7 and 6.4–6.7Calibration and decision reliability

01 / A striking headline result

An estimator can classify a broad set of synthetic forests extremely well.

The first experiment creates 324 mathematically generated forest worlds. The experimenter knows the underlying state and demographic parameters, so reference Health classifications can be calculated. The estimator instead receives noisy or incomplete observations and must reconstruct the state, estimate the dynamics, and classify prospective Health.

GenerateKnown synthetic world

State and generating parameters are controlled

ObserveImperfect measurements

The estimator cannot see the hidden truth

EvaluateEstimated Health

Reconstruct, simulate, and compare

Correct-model broad-range experiment99.22%

Balanced accuracy · 324 synthetic worlds

What this result includes

69 reference-positive and 255 reference-negative cases. The primary reference viable masses are all more than 0.10 away from the 0.75 threshold.

The test is useful for internal model-based verification, particularly on cases with margins large enough to make the decisions easier.

The result is a property of this controlled synthetic population and estimator. It is not a measured diagnostic accuracy for the real Harvard Forest.

02 / Understand the metric

Balanced accuracy gives positive and negative cases equal weight.

Sensitivity asks how often reference-positive cases are correctly classified. Specificity asks how often reference-negative cases are correctly classified. Balanced accuracy averages these two rates.

Balanced classification accuracy

The mean of sensitivity and specificity.

BAbalanced=Sensitivity+Specificity2

For example, 98% sensitivity and 90% specificity produce balanced accuracy (0.98 + 0.90)/2 = 0.94, or 94%. Balanced accuracy prevents an imbalanced test set from rewarding majority-category predictions too heavily.

The metric helps us compare two kinds of correct answer. It does not make the test population representative of the hard decisions we eventually care about. In particular, a test with few examples near its decision boundary cannot establish precision at that boundary.

03 / Deliberately difficult conditions

What happens when we deliberately place worlds near 75%?

In Lesson 15, Q₁ asks whether viable mass reaches θ = 0.75. A viable mass of 0.90 passes easily and 0.40 fails easily. But 0.749 and 0.751 lie on opposite sides despite differing by just 0.002—two-tenths of one percentage point.

Paper V constructs 81 synthetic worlds: 72 designed around specified capacity targets and nine present-realization controls. Among the 72 capacity-targeted worlds, 67 end up within 0.10 of the threshold and 23 within 0.02. This is a more demanding designed population.

Broad-range · correct model99.22%

324 worlds · broad-range

Near boundary · Q₁56.30%

31 certified, presently realizing cases

Near boundary · Q₂55.49%

46 certified, presently realizing cases

Source: Paper V, Section 6.6. The Q₁ and Q₂ near-boundary subsets differ from one another and from the 324-world broad-range experiment. The figures compare performance across different designed test populations, not repeated measurements on one identical sample.

These results can all be correct simultaneously. Broad-range accuracy can be excellent while near-threshold classification is unreliable in the chosen experimental conditions. The reported approximately 56% values do not establish a universal accuracy for real-world health measurement.

04 / Where a tiny error becomes a different answer

A small probability error can flip a Boolean Health judgment.

Q₁ at the declared threshold

At 75%, the decision changes sides.

Q1(mV)={TruemV≥0.75FalsemV<0.75

The equality belongs to the passing side. Thus 0.749 fails and 0.751 passes, even though the two probabilities are close.

Suppose the underlying probability in a specified model is 0.751 while the estimator reports 0.748. The numerical error is only −0.003, or −0.3 percentage points. Nonetheless, the estimated decision is false while the reference decision is true.

An interactive decision-boundary experiment

Move a probability through the 75% line.

We hold the adequacy threshold fixed and vary an illustrative estimated viable mass. Watch how the classification changes at 75%, and how small the signed margin can be.

74.8%
Signed margin−0.2 percentage points
Estimated adequacyFalse · below threshold

This interactive example is a deterministic threshold illustration. It does not calculate an empirical forest estimate or a confidence interval. The threshold remains 75%.

05 / Uncertainty from sampling futures

With 64 simulated histories, one different outcome can reverse the decision.

In a conditional bank with 48 viable histories, the finite-bank estimate is exactly 48/64 = 0.75. But 47 successful futures out of 64 is below 75%, while 49 is above it.

Three neighboring finite banks

Each history represents 1/64 of the bank.

4764≈0.73444864=0.75004964≈0.7656

A single history changes the estimated viable mass by 1/64, about 1.56 percentage points. That change can reverse a threshold classification.

There are two questions here. We can evaluate the finite-bank rule exactly for the histories actually sampled. Estimating the probability of the underlying stochastic generator needs a separate uncertainty analysis. Paper V uses Wilson intervals for binomial proportions to flag cases where the generator-level threshold remains unresolved at the available numerical resolution.

More simulated futures can reduce Monte Carlo uncertainty. They do not demonstrate that the mortality, growth, or entry laws actually describe biological dynamics.

06 / Where is classification information being lost?

A four-way oracle audit separates two information sources.

Because the synthetic worlds have known generating states and parameters, I can supply the estimator with one or both instead of requiring it to infer them. “True” here means known from a synthetic generating world, not independently measured in the real forest.

Paper V, Section 5.7, Eq. (101) · four oracle conditions
ConditionStarting stateFuture-model parameters
E00EstimatedEstimated
E10TrueEstimated
E01EstimatedTrue
E11TrueTrue

Now compare the balanced accuracies measured in those conditions on the near-boundary certified cases.

Paper V, Table 11 (printed page 36) · near-boundary balanced accuracy
Information suppliedQ₁Q₂
E00 · Estimated state and parameters56.30%55.49%
E10 · True state, estimated parameters65.13%58.90%
E01 · Estimated state, true parameters51.47%52.84%
E11 · True state and parameters100.00%100.00%

The true starting state helps in this particular experiment when model parameters remain estimated. Providing true parameters without the true state performs worse here than the ordinary estimated/estimated procedure. This is an interaction inside a particular approximate inference pipeline, not a general principle that more reliable information has negative value.

With both generating inputs supplied, the correctly specified synthetic classifier reaches 100% balanced accuracy on its certified test cases. That verifies agreement within the experiment's model family. It provides no independent validation of real-world forest forecasts.

07 / A deliberately selected reconstruction extreme

Similar starting summaries can generate dramatically different continuation estimates.

Paper V compares two approaches to the unequal dates of individual stem measurements. One uses pooled observation times; another respects individual exposure times. In a selected one-year comparison, the starting-state summaries are close—but the separately sampled future banks disagree strongly.

Paper V, Table 6 (printed page 32) · selected local extreme
QuantityPooled-time reconstructionIndividual-time reconstruction
Living stems90,43190,690
Juvenile-support proxy66,60066,900
Viable futures out of 4,096433,959
Estimated viable mass1.05%96.66%
Pooled-time conditional bank1.05%
43 / 4,096 viable histories
Individual-time conditional bank96.66%
3,959 / 4,096 viable histories

This specific comparison requires juvenile support at one year to reach 80% of the census reference J₀ = 82,077:

The binding continuation requirement in the selected test

Terminal juvenile support must clear the declared minimum.

J1≥0.8J0J0=82077J1≥65661.6

Both reconstructed starting counts lie above this nominal level. Their modeled paths can nevertheless move below it before the one-year evaluation.

Path-by-path failure audit for the two selected banks
Path outcomePooled-timeIndividual-time
Model-lawfulness failure00
Juvenile-support failure only4,053137
Basal-area failure00
Viable continuation433,959

The audit localizes every failed trajectory in these banks to the terminal juvenile-support condition. But the two response banks were separately generated from different reconstructions. We cannot attribute the entire difference to 300 juvenile-proxy stems alone. Nor do these numbers measure an observed jump in the real forest's survival probability. They reveal an extreme sensitivity of this stipulated computational specification in a selected case.

08 / When an important process is missing

Simulating a misspecified model more precisely does not make it biologically right.

The next experiment keeps the comparison within synthetic worlds but deliberately changes the truth-generating process. It includes an extra annual hemlock mortality hazard of 0.12 that the estimator omits.

Correct-model broad-range test4False-positive Health judgments
Omitted hemlock hazard34False-positive Health judgments
Viable-mass mean absolute error0.01259Correct-model experiment
Viable-mass mean absolute error0.10705Omitted-hazard experiment

These are broad-range synthetic comparisons across 324 worlds (Paper V, Section 6.6). The estimator may assign excessive probability to continued organization because its future law excludes a mortality mechanism present in the generating law.

More future draws can reduce sampling noise around that estimator's answer. They cannot supply the omitted biological mechanism. This is a problem of model misspecification, which requires model evaluation and independent evidence rather than simply more Monte Carlo draws.

09 / Three problems with different remedies

To improve a measurement, I first need to know what kind of error I'm confronting.

01 · State uncertaintyThe current system is incompletely known.

Better direct observations, identification, and calibrated reconstruction may improve the state estimate.

02 · Monte Carlo uncertaintyThe generator's probability is estimated from a finite future bank.

More simulated histories and appropriate uncertainty intervals can improve numerical resolution.

03 · Model misspecificationThe assumed dynamics omit or misstate generating processes.

Model revision, independent testing, and additional biological evidence are needed.

These problems can coexist. The oracle experiment shows state and parameter errors interacting; the near-boundary experiments show classification changing under tiny mass errors; the misspecified hazard experiment shows a source of error that cannot be fixed by counting more histories.

Specification execution and empirical measurement are separate achievements. Paper V demonstrates how to execute declared Health calculations and audit their limitations. It does not assert that those conditional classifications establish the actual ecological Health of Harvard Forest.

What I want you to carry forward

Accuracy far from the threshold does not certify decisions made on the threshold.

Knowing a formal Health requirement tells us exactly which comparison to make. Empirically, we also need state reconstruction accurate enough for the question, a sufficiently resolved future probability, and a dynamics model supported by the real system. Near sharp requirements, apparently modest errors can reverse the final classification.

Lesson 17 leaves the synthetic experiment and asks how the fitted mortality model fares against independent observations. In the adult-hemlock comparison, the model expects approximately 280 deaths but later records report 685 secure deaths. We'll examine that predictive failure, its uncertainty calculation, and the precise limits of what it establishes.

Source: Zed James, Prospective Health under Declared Specifications, Paper V in Health, Formally Defined (2026), Sections 5.4–5.7 and 6.4–6.7, especially Tables 6 and 11, Equation (101), and the broad-range and near-boundary experiments. Full publication record · Zenodo DOI. Synthetic experiments, selected extreme examples, and conditional future simulations do not constitute validated ecological diagnoses.