01 / An accurate observation can be insufficient
Let's begin with two systems that give the same simple measurement.
Imagine two biological systems. Each has a possible successful continuation. If I ask only whether success is possible, both answer yes.
Now I want to enrich the example. Suppose the two systems are described by different probability distributions over success and failure. This is a stipulated mathematical example, developed later in Paper III, Section 12.3.
999 successes out of 1,000, according to the stipulated distribution.
1 success out of 1,000, according to the stipulated distribution.
These percentages are given by the formal example. They are not observed biological frequencies or empirically validated probabilities for real systems.
The support of the two distributions—the set of outcomes with positive probability—is the same: success and failure are both possible in each. So a measurement recording only Is successful continuation possible? returns yes for both.
But suppose my declared health question requires at least a 50% chance of successful continuation. System A passes that threshold and System B does not. The simple yes/no possibility measurement has merged two cases that this new question must distinguish.
| What I ask | System A | System B |
|---|---|---|
| Successful continuation possible? | Yes | Yes |
| Stipulated success probability | 99.9% | 0.1% |
| Meets the declared 50% requirement? | Yes | No |
Nothing was necessarily measured incorrectly. We simply didn't record the probability information that the new question needs. That is the problem I want to solve mathematically.
02 / How information is encoded
What exactly do I mean by a representation?
In this framework, a representation is a rule for describing the objects we want to study. It may be a measurement, a classification, or the result of a mathematical procedure. It maps each state into some observable or descriptive value.
The basic mathematical object
A rule that assigns a description.
X is the collection of states or objects, O is the collection of possible descriptions, and e is the rule that produces a description for each state.
Imagine a forest. I might record only its number of trees. Or I might describe the distribution of species, the size of every tree, juvenile regeneration, or some representation of its future responses. All of those can be meaningful descriptions, while preserving different distinctions.
Suppose two forests each contain one thousand trees. The tree-count representation gives both the same value. The forests may nevertheless differ in regeneration and their capacity to continue under a specified challenge.
That observation leads to a practical question: Has the representation discarded something that matters to the exact health question I intend to answer?
03 / What does sufficient mean?
Sufficiency means preserving the answers the questions require.
Suppose I'm considering a collection of health-relevant questions, and I have a representation e that gives me a description of each state. To test sufficiency, I compare two states whenever the representation returns exactly the same description.
If their answers agree for every declared question, the representation is sufficient for that question family. If even one question gives opposite answers, the representation has merged two states that the question must distinguish.
Paper III · Representation sufficiency
The answer must survive the representation.
This is the condition illustrated in Paper III, Section 2. The phrase “every question” refers to the declared question family, not every question anyone could imagine.
- e(x)
- What the representation records about state x.
- e(x) = e(y)
- The observation cannot distinguish the two states.
- A(q,x)
- The answer to declared question q about state x.
- ∀
- “For every.” Every question in the declared family must agree.
- ⟹ and ↔
- The first means “implies.” The second means the two answers have the same truth value.
This condition does not require me to reconstruct the full state. It asks for exactly the distinctions that the declared questions can detect. That's a much more focused information requirement.
04 / The principal theorem
There is a canonical way to describe the information each question needs.
In Lesson 3 we learned how to group capacities that give identical adequacy answers. Paper III applies the same underlying idea to a declared family of questions over states. I call the resulting answer-class map the canonical query-visible quotient, written qA.
It groups together exactly the states whose entire pattern of answers agrees. If all our questions treat two states identically, the quotient places them in the same class. If any question separates them, the quotient keeps the classes separate.
Theorem 3.1 · Principal sufficiency
Representations must refine the question's own information.
The relation ⪰ means refinement: e retains at least every distinction retained by qA. It may retain more.
Refinement has a precise meaning. If the finer representation e gives equal values on two states, the coarser representation qA must give equal values as well. Another way of saying this is that e is not allowed to merge a pair of states that the question-answer quotient needs to keep apart.
Both systems answer yes. The 50% requirement cannot be recovered.
99.9% and 0.1% remain distinguishable. The threshold answer is determined.
Yes and no remain distinct. This is exactly the distinction required by the single threshold question.
Both the exact probability and the threshold-answer representation are sufficient for the 50% question. The first retains information beyond the question; the second retains the precise answer class. The success-possible representation is insufficient because it assigns the same value to two states with different threshold answers.
The theorem proves the relationship in the declared mathematical setting: a representation is sufficient if and only if it refines the quotient determined by the questions.