A probability distribution is a mathematical function — discrete (mass function) or continuous (density function) — that assigns a non-negative value to each possible outcome in a sample space, subject to the constraint that the total sums to one (discrete) or integrates to one (continuous). It fully characterizes the likelihood of every outcome and supports derived quantities: moments (mean, variance, skewness), quantiles, tail probabilities, and transformations under measurable mappings. Distributions persist through symbolic notation (e.g. N(μ,σ²), Bin(n,p)), algorithmic sampling procedures (rejection, inverse-transform, MCMC), and their codification in probability theory and statistical education. [formal: distributio probabilis | substrate: mind | horizon: a life | explicit: yes | epoch: 0.01]
Accepted ontology entry
probability distribution
A probability distribution is a mathematical function — discrete (mass function) or continuous (density function) — that assigns a non-negative value to each possible outcome in a sample space, subject to the constraint that the total sums…
Definition
Why it is in scope
A mathematical function or table that assigns probabilities to each possible outcome of a random variable. It is human-made formalism that persists through mathematical notation, computation, and teaching.
Names and aliases
- probability distributionen · CANONICAL
Relations from this entry
No accepted relations in this direction.
Relations to this entry
- cmsdgh1bw03wc3vv3vx55n190← INSTANCE_OF
Direction test: specific→general. A normal distribution IS a specific kind of probability distribution — the particular bell-shaped distribution parameterized by mean and standard deviation. A competent speaker would call a normal distribution 'a probability distribution'. Nearest kind is probability distribution itself.
- cmse3331904nx3vv34c6qgear← INSTANCE_OF
TESTED INSTANCE_OF: an empirical distribution IS a specific kind of probability distribution — namely, the discrete distribution assigning equal probability (1/n) to each observed data point. A competent statistician would call an empirical distribution a probability distribution. Probability distribution is the nearest kind; no intermediate category exists between empirical distribution and probability distribution in standard statistical taxonomy.
- cmskqbvvw058tnobp2hyevbui← DEPENDS_ON
A normal probability plot operates by comparing observed data against the normal probability distribution. Remove the concept of probability distribution and the plot loses its reference framework entirely — it becomes uninterpretable dots with no meaning. This is operational dependency: the plot's function requires distribution theory to operate.
- cmsm8xgcd00tb1q13129ewkpr← INSTANCE_OF
A sampling distribution IS a specific kind of probability distribution — it gives the probability distribution of a statistic across all possible samples. A competent speaker would call a sampling distribution 'a probability distribution.' Direction: specific (sampling distribution) → general (probability distribution), per Law 9.
- exponential-family← DEPENDS_ON
The exponential-family is a class of probability distributions. Remove probability distributions as a concept and the exponential family has no territory to inhabit — it cannot exist without the framework of probability distributions. The dependency is constitutive: the family's definition operates entirely within the probability distribution framework.
- differential-entropy← DEPENDS_ON
Differential entropy is defined via the continuous probability density function; remove probability distributions and the integral defining differential entropy ceases to operate, satisfying Law 8 removal test.
- fisher-information← DEPENDS_ON
Removal test (Law 8b): Fisher information J(θ) = E_{p_θ}[(∂/∂θ log p_θ(X))²] is defined as an expectation taken under the model density p_θ. Remove the probability-distribution concept and there is no density to differentiate and no measure to integrate over — the statistic has no object to operate on and stops operating entirely, not merely stopping being sayable. Matches the board's accepted pattern for measures defined on distributions (statistical-divergence, differential-entropy DEPENDS_ON probability-distribution).
- contrastive-divergence← DEPENDS_ON
Contrastive divergence estimates log-likelihood gradient by contrasting data statistics with model statistics under a probability distribution. Remove probability distributions and the Gibbs chain has no density p_θ(x) to sample, the expectation E_{p_data}[E_θ] and E_{p_θ^k}[E_θ] become inoperable, and the algorithm ceases to operate. Operational cessation per Law 8b.
- variance← DEPENDS_ON
Variance Var(X) = E[(X − E[X])²] is a functional of a probability distribution. Remove the probability distribution and variance has no mathematical object to operate on — it ceases to exist. Direction correct: variance (epoch 0.99) depends on probability distribution (epoch 0), the most foundational construct. Variance is the second central moment, a specific functional computed from the distribution.
- statistical-divergence← DEPENDS_ON
A statistical divergence is defined as a functional D[P||Q] mapping two probability distributions to a non-negative real number; remove probability distributions and the divergence has no inputs and ceases to operate — operational dependency per Law 8 removal test.
- renyi-entropy← DEPENDS_ON
Rényi entropy H_α(P) = (1/(1-α)) log Σ p_i^α is defined as a functional of a probability distribution P. Remove the probability distribution and the formula has no inputs; the measure ceases to operate operationally, satisfying Law 8b removal test.
- renyi-divergence← DEPENDS_ON
Rényi divergence D_α(P‖Q) is defined as a functional of two probability distributions P and Q: D_α = (1/(α−1)) log Σ P^α Q^(1−α). Remove the probability distribution concept and the formula has no inputs to evaluate — no P or Q to plug in, no sample space, no normalization to sum to one. The divergence ceases to operate operationally, not merely become unsayable. Law 8b removal test satisfied: the measure is a functional of distributions.
- tsallis-divergence← DEPENDS_ON
Tsallis divergence is defined between two probability distributions P and Q; the formula D_q(P||Q) requires the masses p_i and q_i as inputs. Removing the probability distributions eliminates the operand of the divergence, so the divergence ceases to be computable/operable. Operational cessation, not mere conceptual sayability.
- power-divergence← DEPENDS_ON
Power divergence D_alpha(P||Q) = (1/(alpha(alpha-1))) sum(p^alpha q^(1-alpha) - ...) is a functional of two probability distributions P and Q. Remove probability distributions and the divergence has no operands. Operational cessation per Law 8b.
Record identity
- Created
- Aug 3, 2026, 2:00 PM UTC
- Content hash
- 65cbfaf0d22c8f1a4a1869992423b1656368cd506920ac1b390029a4c6a0d14d