Skip to lesson content

Lesson 08 · The information required to measure Health

How do we know whether a measurement contains enough information to determine health?

I want to begin Paper III with a distinction that matters to any science of measurement. A measurement may record something perfectly accurately and still lack the information needed to answer the question we're asking.

By Zed JamesPaper III · Sections 1–5 and 9Opening the representation series

01 / An accurate observation can be insufficient

Let's begin with two systems that give the same simple measurement.

Imagine two biological systems. Each has a possible successful continuation. If I ask only whether success is possible, both answer yes.

Now I want to enrich the example. Suppose the two systems are described by different probability distributions over success and failure. This is a stipulated mathematical example, developed later in Paper III, Section 12.3.

System A99.9%

999 successes out of 1,000, according to the stipulated distribution.

System B0.1%

1 success out of 1,000, according to the stipulated distribution.

These percentages are given by the formal example. They are not observed biological frequencies or empirically validated probabilities for real systems.

The support of the two distributions—the set of outcomes with positive probability—is the same: success and failure are both possible in each. So a measurement recording only Is successful continuation possible? returns yes for both.

But suppose my declared health question requires at least a 50% chance of successful continuation. System A passes that threshold and System B does not. The simple yes/no possibility measurement has merged two cases that this new question must distinguish.

Different declared probabilities, identical success-possible answers
What I askSystem ASystem B
Successful continuation possible?YesYes
Stipulated success probability99.9%0.1%
Meets the declared 50% requirement?YesNo

Nothing was necessarily measured incorrectly. We simply didn't record the probability information that the new question needs. That is the problem I want to solve mathematically.

02 / How information is encoded

What exactly do I mean by a representation?

In this framework, a representation is a rule for describing the objects we want to study. It may be a measurement, a classification, or the result of a mathematical procedure. It maps each state into some observable or descriptive value.

The basic mathematical object

A rule that assigns a description.

e:X→O

X is the collection of states or objects, O is the collection of possible descriptions, and e is the rule that produces a description for each state.

Imagine a forest. I might record only its number of trees. Or I might describe the distribution of species, the size of every tree, juvenile regeneration, or some representation of its future responses. All of those can be meaningful descriptions, while preserving different distinctions.

Suppose two forests each contain one thousand trees. The tree-count representation gives both the same value. The forests may nevertheless differ in regeneration and their capacity to continue under a specified challenge.

That observation leads to a practical question: Has the representation discarded something that matters to the exact health question I intend to answer?

03 / What does sufficient mean?

Sufficiency means preserving the answers the questions require.

Suppose I'm considering a collection of health-relevant questions, and I have a representation e that gives me a description of each state. To test sufficiency, I compare two states whenever the representation returns exactly the same description.

If their answers agree for every declared question, the representation is sufficient for that question family. If even one question gives opposite answers, the representation has merged two states that the question must distinguish.

Paper III · Representation sufficiency

The answer must survive the representation.

e(x)=e(y)⟹∀q∈Queries,A(q,x)↔A(q,y)

This is the condition illustrated in Paper III, Section 2. The phrase “every question” refers to the declared question family, not every question anyone could imagine.

e(x)
What the representation records about state x.
e(x) = e(y)
The observation cannot distinguish the two states.
A(q,x)
The answer to declared question q about state x.
∀
“For every.” Every question in the declared family must agree.
⟹ and ↔
The first means “implies.” The second means the two answers have the same truth value.

This condition does not require me to reconstruct the full state. It asks for exactly the distinctions that the declared questions can detect. That's a much more focused information requirement.

04 / The principal theorem

There is a canonical way to describe the information each question needs.

In Lesson 3 we learned how to group capacities that give identical adequacy answers. Paper III applies the same underlying idea to a declared family of questions over states. I call the resulting answer-class map the canonical query-visible quotient, written qA.

It groups together exactly the states whose entire pattern of answers agrees. If all our questions treat two states identically, the quotient places them in the same class. If any question separates them, the quotient keeps the classes separate.

Theorem 3.1 · Principal sufficiency

Representations must refine the question's own information.

RepresentationSufficient(A,e)⇔e⪰qA

The relation ⪰ means refinement: e retains at least every distinction retained by qA. It may retain more.

Refinement has a precise meaning. If the finer representation e gives equal values on two states, the coarser representation qA must give equal values as well. Another way of saying this is that e is not allowed to merge a pair of states that the question-answer quotient needs to keep apart.

01 / Too coarseSuccess is possible?

Both systems answer yes. The 50% requirement cannot be recovered.

02 / More detailedExact success probability

99.9% and 0.1% remain distinguishable. The threshold answer is determined.

03 / Task-visibleMeets 50%?

Yes and no remain distinct. This is exactly the distinction required by the single threshold question.

Both the exact probability and the threshold-answer representation are sufficient for the 50% question. The first retains information beyond the question; the second retains the precise answer class. The success-possible representation is insufficient because it assigns the same value to two states with different threshold answers.

The theorem proves the relationship in the declared mathematical setting: a representation is sufficient if and only if it refines the quotient determined by the questions.

The same example, three descriptions

What did our representation actually retain?

All three descriptions below are accurate under the stipulated example. Select one to focus on what information it keeps for the declared 50% success question. All three remain readable without JavaScript.

Focus on a representation
01 / Success possible?
System A YesSystem B Yes

Both systems look identical to this representation.

Insufficient for 50%
02 / Success probability
System A 99.9%System B 0.1%

The exact values preserve the difference needed for the threshold.

Sufficient, more detailed
03 / Meets at least 50%?
System A YesSystem B No

Exactly the answer distinction required by this one declared question.

Sufficient, task-visible

Possibility alone is insufficient: it returns the same value for systems with opposite 50% threshold answers.

The probabilities are stipulated mathematical values from Paper III, Section 12.3. They do not indicate observed risks or scientific health assessments for real organisms.

05 / What postprocessing can and cannot do

We cannot calculate our way out of missing information.

I want to explore one especially consequential implication of the sufficiency condition. Suppose a representation e gives the same value for two states x and y, but our declared health question gives different answers.

Now imagine processing e through an arbitrarily sophisticated deterministic procedure f. The procedure still receives the same input in both cases. So its output must be the same too:

The deterministic postprocessing limit

Equal inputs remain equal after the same rule.

e(x)=e(y)⟹f(e(x))=f(e(y))

This is the essential postprocessing argument from Paper III's information-loss results. A deterministic transformation cannot split identical recorded inputs into distinct correct answers without additional distinguishing information.

In the probability example, if all I've retained is “success is possible,” both systems receive yes. No deterministic calculation using only that value can recover which system passes the 50% threshold.

That does not make inference or prediction impossible. A statistical model may bring in new assumptions, training data, or other measured variables to estimate a missing distinction. Such a prediction is conditional on those added ingredients. It is not an exact deduction from the original yes/no representation alone.

This becomes important later when we move from formal histories to empirical and model-based assessments in Paper V.

06 / What I mean by the least information

Minimal here has a precise mathematical meaning.

The canonical quotient groups together exactly those states that give identical answers to the declared question family. Every sufficient representation must preserve at least that grouping's distinctions.

For the single 50% question, the task-visible quotient only needs two answer classes: passes and does not pass. Knowing the exact probability is sufficient but gives us a finer description than that question alone demands.

If I expand the question family to include every probability threshold between zero and one, then two different probabilities can be separated by at least one threshold. In that expanded language, the exact success probability becomes the minimal task-visible quantity. Paper III studies this distinction separately in Section 10.1.

What “least” means here

The smallest distinction structure sufficient for the declared questions. It is not necessarily the smallest file, fewest bytes, easiest computation, lowest entropy, or least expensive instrument.

The measurement-design implication is subtle: we can identify what distinctions an instrument must preserve without yet knowing how to build the cheapest instrument that does so.

07 / Connecting the mathematics to measurement science

Logical sufficiency is one part of scientific validity.

Let me connect this result to the larger program we've been developing. Paper I established an explicit prospective Health predicate. Paper II made context, licensing, and transport conditions visible. Paper III now provides a mathematical test for whether a chosen representation contains enough information to determine the question's answer.

If a representation merges states that a declared health question distinguishes, it is insufficient for exact determination of that answer. If the representation refines the canonical query-visible quotient, it is logically sufficient in the declared formal setting.

But that logical property alone doesn't establish that we can measure the representation accurately, model the future correctly, or justify the biological meaning of the requirement. Those are additional scientific obligations.

It's precisely why I distinguish an exact mathematical information guarantee from empirical calibration, uncertainty, and validation. We need both kinds of work when we want a formal judgment to become a trustworthy measurement of a living system.

What I want you to take forward

The right measurement is the one that preserves the distinctions our question needs.

Let's finish with a forest again. Two forests might contain exactly the same number of trees. If my health question depends on juvenile regeneration, their shared total tree count may be perfectly accurate while leaving me unable to decide whether the requirement is met. The number alone has merged cases the health question needs apart.

This is the central lesson of Paper III's principal theorem. Representation sufficiency is a relationship between the information a description retains and the distinctions demanded by a declared family of questions. A sufficient representation refines the canonical query-visible quotient. Deterministic postprocessing of an insufficient representation cannot restore distinctions it has already erased.

Next we'll combine that theorem with transport from Paper II. Sometimes the complete future representation cannot be determined, yet a coarser representation sufficient for our health questions still can. I'll develop that relationship in Lesson 9, then turn to probability, cost, control, and robustness before we reach the empirical forest studies.

Source: Zed James, Representation Sufficiency in Prospective Health: Query-Visible Quotients, Transport, and Enriched Response Semantics, Paper III in Health, Formally Defined (2026), particularly Sections 1–5 and 9. The stipulated probability example is from Section 12.3; the full threshold-family observation is from Section 10.1. Paper III publication record · Zenodo DOI. Formal examples establish information properties under declared assumptions and are not measured biological probabilities or clinical Health judgments.