Mutual information is a measure from information theory that quantifies how much knowing the value of one random variable reduces uncertainty about another. Its parameters are: (1) two random variables X and Y, (2) their joint probability distribution p(x,y), (3) their marginal distributions p(x) and p(y). It is computed as the KL-divergence between the joint distribution and the product of the marginals. It persists through mathematical formalism in probability theory, machine learning feature selection, and communication theory. [formal: mutualis informatio | substrate: mind | horizon: hours | explicit: yes | epoch: 0.16]
Accepted ontology entry
mutual information
Mutual information is a measure from information theory that quantifies how much knowing the value of one random variable reduces uncertainty about another. Its parameters are: (1) two random variables X and Y, (2) their joint probability…
Definition
Why it is in scope
Mutual information is a human-made measure of how much knowing one random variable reduces uncertainty about another. It quantifies the shared information between variables — built to persist through probability theory, statistics, and information theory as a tool for feature selection, dependency detection, and model evaluation.
Names and aliases
- mutual informationen · CANONICAL
Relations from this entry
- cmsf6w1ek06ms3vv3snhiwxwlDERIVED_FROM →
Probability theory existed centuries before mutual information: the former traces to the 17th century (Pascal, Fermat, Jakob Bernoulli), while mutual information was introduced by Shannon in 1948 as a concept within information theory, which itself grew out of probability theory. The which-came-first test is decisive: probability theory is the older, broader framework that fed mutual information.
- cmsftujnv00f0qszgpolboiajDEPENDS_ON →
Mutual information was introduced by Shannon as part of information theory. It measures dependence between random variables using IT's entropy formalism. Remove information theory's framework and mutual information loses its operational machinery (Law 8). Present-tense dependency, not just historical association.
- cmsfxlnqd00qqqszg8hkj1cdgDEPENDS_ON →
Mutual information I(X;Y) = H(X)+H(Y)-H(X,Y) — defined entirely in terms of Shannon entropy. Remove entropy and mutual information has no definition.
- kullback-leibler-divergenceDEPENDS_ON →
Mutual information I(X;Y) = D_KL(p(x,y) || p(x)p(y)) — its definition is literally the KL-divergence between the joint distribution and the product of marginals. Remove KL-divergence and mutual-information has no definition. Constitutive present-tense dependency: mutual-information operates through KL-divergence as its computational mechanism (Law 8b/8c).
- conditional-entropyDERIVED_FROM →
Mutual information I(X;Y) = H(Y) - H(Y|X) is derived from the conditional entropy formula. Conditional entropy extends Shannon entropy to dependent variables; mutual information measures the reduction in conditional entropy compared to marginal entropy. Conditional entropy is the more fundamental construct that mutual information derives from.
Relations to this entry
No accepted relations in this direction.
Record identity
- Created
- Aug 5, 2026, 5:22 AM UTC
- Content hash
- 5323a8b65e21f5a06fb842ff49d00558a8722a5119f08a069e53356a90ffe5a4