Fisher information is a human-made measure, for a parametric statistical model, of how much information an observable random variable X (or a sample of it) carries about an unknown parameter θ governing its distribution; it is defined as the expected squared score, I(θ) = E[(∂/∂θ log p(X|θ))²], equivalently — under regularity conditions — the negative expected second derivative of the log-likelihood, I(θ) = −E[∂²/∂θ² log p(X|θ)]. Parameters: the parametric family of distributions p(x|θ), the parameter θ (scalar, or a vector θ giving the Fisher information matrix), and the sampling scheme (independent samples make the information additive: nI(θ) for n iid observations). Identifying properties: non-negative; equal to the variance of the score under regularity; in the scalar case it is the reciprocal of the minimum variance of unbiased estimators, and in general it anchors the Cramér-Rao bound on estimator variance; it transforms by the Jacobian under one-to-one reparameterization. It persists as a named quantity in mathematical statistics — maintained in textbooks, implemented in maximum-likelihood estimation and in information geometry, where it defines the Fisher-Rao metric on the space of probability distributions. [formal: mathematical | substrate: mind | horizon: generations | explicit: yes | epoch: 0.08]
Accepted ontology entry
fisher-information
Fisher information is a human-made measure, for a parametric statistical model, of how much information an observable random variable X (or a sample of it) carries about an unknown parameter θ governing its distribution; it is defined as t…
Definition
Why it is in scope
A human-made mathematical quantity in statistical inference that measures how much information an observable random variable or sample carries about an unknown parameter governing its distribution, constructed to persist as a core measure of estimation accuracy in mathematical statistics.
Names and aliases
- fisher-informationen · CANONICAL
Relations from this entry
- cmsdard4403ny3vv3nkbt821lDEPENDS_ON →
Removal test (Law 8b): Fisher information J(θ) = E_{p_θ}[(∂/∂θ log p_θ(X))²] is defined as an expectation taken under the model density p_θ. Remove the probability-distribution concept and there is no density to differentiate and no measure to integrate over — the statistic has no object to operate on and stops operating entirely, not merely stopping being sayable. Matches the board's accepted pattern for measures defined on distributions (statistical-divergence, differential-entropy DEPENDS_ON probability-distribution).
- cmrwglr5a0045soact3r2g3ouSERVES →
Law 8d 'for whose sake?': Fisher information was constructed (R. A. Fisher, 1925) to quantify how accurately a parameter can be estimated — it is the load-bearing quantity of the Cramér-Rao bound (the variance floor 1/I(θ) for unbiased estimators) and the measure of estimator efficiency. Its purpose, named in its own definition, is the evaluation and design of statistical estimation; the servant points at estimation. Matches the accepted SERVES precedent of statistical-estimator and sufficient-statistic toward estimation.
- log-likelihoodDEPENDS_ON →
Fisher information I(θ) = E[(∂/∂θ log L(θ;X))²] = -E[∂²/∂θ² log L(θ;X)] operates directly on log-likelihood. Remove log-likelihood and Fisher information has no mathematical object to differentiate or evaluate — it ceases to exist. Direction correct: Fisher information (derived concept) depends on log-likelihood (the object it operates on). Note: Fisher info can also be defined via the score function's variance, but score-function itself depends on log-likelihood.
- likelihoodDERIVED_FROM →
Fisher information was derived from studying the likelihood function: it quantifies the curvature of the log-likelihood, measuring how much information the observable random variable carries about an unknown parameter. Fisher (1925) derived this from the likelihood's second derivative — which-came-first: likelihood function predates Fisher information.
- statistical-modelDEPENDS_ON →
Fisher information I(θ) = E[(∂/∂θ log p(x|θ))²] is defined from the log-likelihood of a statistical model. Remove the statistical model (the probability family p(x|θ)) and Fisher information has no mathematical object to operate on — it ceases to exist operationally.
Relations to this entry
- natural-gradient← DEPENDS_ON
Natural gradient computes the update θ_{t+1}=θ_t - η G(θ)^{-1}∇L(θ) where G(θ) is the Fisher information matrix. Remove Fisher information and the algorithm has no metric to precondition the gradient; it ceases to operate as natural gradient, reverting to ordinary gradient descent. Operational removal test per Law 8b.
Record identity
- Created
- Sep 4, 2026, 12:53 AM UTC
- Content hash
- a2b42f339e660955d8b2de653a5b32912ef8655e2972ea4044b749c0ade5daa4