SYSTEMA CONSTRUCTUM

Accepted ontology entry

statistics

A statistics is a human-made corpus of methods and principles for reasoning under uncertainty — collecting data through sampling, analyzing patterns through inference, and communicating findings through quantification. It operates through…

ACCEPTED THINGcmrxj3acr03cmsoacx73fal1o

Definition

A statistics is a human-made corpus of methods and principles for reasoning under uncertainty — collecting data through sampling, analyzing patterns through inference, and communicating findings through quantification. It operates through three parameters: (1) a formal language of probability and distribution, (2) procedures for inference from finite samples to larger populations, and (3) standards for measuring and reporting uncertainty. The field persists through mathematical notation, computational tools, and the teaching of statistical literacy across disciplines. [formal: statistica | substrate: mind | horizon: centuries | explicit: yes | epoch: 0.01]

Why it is in scope

A human-made field of practice and body of methods for collecting, analyzing, interpreting, and presenting data — a systematic corpus of techniques for reasoning from sample to population, from observation to inference, built to persist through mathematical notation and statistical education.

Names and aliases

Relations from this entry

  • cmrg4vstb00kx2a1nar61eoe7DEPENDS_ON →

    Statistics operates through mathematical operations — probability theory, inferential formulas, and computational methods are constitutive to its function. Remove mathematics and statistics cannot compute, infer, or operate at all. This passes the removal test: statistics stops working without its mathematical foundation.

  • cmsnwdk0t05311q13zh9dw0l7SERVES →

    Statistics is built and maintained for the sake of probabilistic reasoning — it provides the tools for inference, estimation, and hypothesis testing under uncertainty. Per Law 8d: the servant (statistics) points at the master (probabilistic reasoning). Remove statistics and probabilistic reasoning loses its formal inferential apparatus.

Relations to this entry

  • cmrxj6pvj03d0soacyxv3azj6← DEPENDS_ON

    A sampling frame is the concrete list or source from which a sample is drawn. It only operates within the context of statistical sampling. Remove statistical sampling (statistics) and the sampling frame concept ceases to operate — it is not a standalone concept. Present-tense dependency, not just historical.

  • cmrxgmkw5037psoace2igk7p5← DEPENDS_ON

    A correlation coefficient (Pearson's r, Spearman's rho, etc.) is a statistical measure that requires the statistical framework to operate meaningfully. The removal test: remove statistics as a conceptual paradigm and correlation coefficients lose their interpretive apparatus — their calculation becomes just arithmetic without the inferential framework that gives them meaning. The formula exists independently, but the CONCEPT of a correlation coefficient as a measure of relationship strength and significance depends on statistics to function as intended.

  • cmrxgu0nx0382soacugywq2ra← DEPENDS_ON

    Margin of error is a statistical concept used in sampling and survey methodology. The removal test: remove the statistical framework and margin of error loses its meaning — it cannot function as a measure of uncertainty without probability theory and statistical inference. It is a statistical tool that needs statistics to operate.

  • cmrw668qi002nkyo6qlh173bj← DEPENDS_ON

    Remove statistics and power-analysis stops operating: power analysis uses statistical distributions (normal, t, F), effect sizes, alpha/beta error rates, and computational statistical methods. Without statistical theory, power analysis has no framework, no calculations, and no operational meaning. which-came-first: statistics (epoch 0.12) predates power-analysis (epoch 0.15).

  • cmrw668qi002nkyo6qlh173bj← INSTANCE_OF

    Power analysis IS a specific kind of statistics — a statistical technique for determining sample size and detecting effects. A competent speaker would say 'power analysis is a kind of statistical method/technique.' The nearest kind is statistics itself (no intermediate category between power analysis and statistics).

  • cms7mm14t01zya34cxii43ay2← DEPENDS_ON

    Six-sigma needs statistics to operate now: remove statistical methods (hypothesis testing, regression, control charts, process capability analysis) and six-sigma loses its entire analytical machinery. It is a data-driven methodology — data analysis IS its operation. The removal test: without statistics, six-sigma cannot function (no way to measure variation, detect defects, or validate improvements). Direction: from six-sigma (the methodology) to statistics (the tool it needs).

  • cms7qpssn006q1yrmdbysknhd← DEPENDS_ON

    causal-inference as a methodology needs statistical tools (regression, matching, IV estimation) to operate — remove statistics and causal-inference loses its operational mechanism. This is a present-tense dependency, not merely historical association.

  • cms7uitrq001ekaytbzlyubop← DEPENDS_ON

    Ecological fallacy operates as a statistical reasoning error — remove statistics as a discipline and the concept cannot be identified or understood. Present-tense necessity (Law 8).

  • cmsczjgpa03c83vv32r6la0gn← DERIVED_FROM

    Concept drift originated in statistical process control (1920s Shewhart charts) and was later adopted by machine learning. Statistics predates ML as a discipline, making it the historical source.

  • cmsc3aqll02eb3vv3xy5gcx30← DEPENDS_ON

    Remove statistics and stem-and-leaf plot ceases to operate: the plot's entire mechanism — sorting data into stems and leaves, interpreting frequency distributions, extracting summary statistics — depends on statistical concepts to function. Without statistics as a discipline, there is no framework for creating, reading, or using stem-and-leaf plots. This is a present-tense operational dependency, not a historical association.

  • cmsdai2d503n23vv3e00xn5ba← INSTANCE_OF

    Regression is a specific kind of statistical method for modeling relationships between variables. A competent speaker would call regression 'a type of statistics/statistical method.' Files against the nearest kind: statistics.

  • cmsd0lyke03ds3vv3dayylbx7← DERIVED_FROM

    Statistics as a discipline existed first and provided the mathematical framework for identifying outliers. Which came first? The concept of outliers emerged from statistical theory — the study of data distributions and their tails predates the specific term 'outlier'. Statistics fed into the identification of outliers as extreme observations in a distribution.

  • cmsdfdqpw03tt3vv3zzh0uh0q← INSTANCE_OF

    Resampling IS a specific kind of statistical technique: it involves repeatedly drawing samples from a dataset to estimate sampling distributions. A competent speaker would call resampling a statistical method. Specific→general.

  • cmsf4085106gk3vv3xswmloe3← DEPENDS_ON

    BAYESIAN OPTIMIZATION needs STATISTICS: Remove statistics — the probabilistic surrogate model (posterior distributions, expected improvement acquisition functions) stops operating. Bayesian optimization fundamentally relies on statistical inference to update its model from observed data points. Without statistical concepts, BO has no mechanism for uncertainty quantification or sequential decision-making.

  • cmsceekiv02rt3vv3irnbtc6k← DEPENDS_ON

    Cross-validation estimates model performance by partitioning data into training/validation folds and computing aggregate statistics (mean, variance of scores). Without statistics there is no framework to define, compute, or interpret these performance metrics — the method stops operating.

  • cmseehrm405al3vv3daniseir← DEPENDS_ON

    Feature selection uses statistical tests (t-test, chi-square, ANOVA F-value, mutual information) to rank and select informative features. Remove statistics and these tests cannot be computed — the selection process collapses. Present-tense dependency confirmed.

  • cmsfqkhi0006xqszg5nq0fc3y← DEPENDS_ON

    dimensionality reduction uses statistical foundations (variance, covariance, distributions, optimization) to operate. Remove statistics and dimensionality reduction stops working — the mathematical machinery that computes projections, preserves distances, and optimizes criteria is statistical.

  • cmsfoqiqi0049qszgvfxi6bb7← DEPENDS_ON

    Model comparison is the practice of evaluating and comparing statistical models using criteria like AIC, BIC, cross-validation. Remove statistics and model comparison cannot operate — these are fundamentally statistical procedures. The removal test passes.

  • cmsemkwpa05lg3vv3ih2qj47k← DEPENDS_ON

    A smoothing spline minimizes a penalized sum of squared residuals — a statistical objective function. Remove statistical optimization and the spline has no definition or operating mechanism. The removal test of Law 8 passes.

  • cmrxhuhch039zsoac5yzo1eoh← DERIVED_FROM

    Statistics as a field existed before stem-and-leaf plots were invented. The general statistical method predates the specific visualization technique. Which came first? Statistics.

  • cmsg35xo9012gqszgf7sl0rkz← DERIVED_FROM

    Statistics as a formal field (probability theory, frequentist inference) existed decades before bootstrap methods. Bootstrap (Efron, 1979) built on existing statistical theory to create a new resampling-based approach. The which-came-first test (Law 7) passes: statistics predates and fed into bootstrap methods.

  • cmsf8ods206q93vv3ll6mctmp← INSTANCE_OF

    Gini coefficient IS a specific kind of statistical measure — it quantifies inequality using a formula derived from the Lorenz curve. A competent speaker would call it a statistical measure.

  • cmsfhxh4z07as3vv3tujq1ro0← DEPENDS_ON

    False negative rate is a statistical metric that requires statistical theory to operate. Remove statistics and the concept, calculation, and interpretation of false negative rate ceases to function.

  • cmsfyqlel00t8qszgthzhgrug← DERIVED_FROM

    Statistical learning is derived from statistics. The statistical framework (dating to the 17th-18th century) existed first and fed into the development of statistical learning as a computational discipline in the 1990s-2000s. Which existed first? Statistics predates statistical learning by centuries.

  • cms7xukk2007xh6s8dqyihtf7← DEPENDS_ON

    model-selection DEPENDS_ON statistics: remove statistical theory (AIC, BIC, likelihood, cross-validation) and model-selection ceases to operate. There is no framework to compare, evaluate, or select among candidate models without statistical grounding. The removal test passes.

  • cmskx1k6305p8nobpedwgz05s← DEPENDS_ON

    Regression diagnostics needs statistics to operate — remove statistics and the methodology collapses. Object-level dependency.

  • cmskvhhqd05lmnobp663c0g8w← DERIVED_FROM

    Test: which came first? Statistics as a field (centuries old) predates the formal concept of model capacity in machine learning (1990s). The concept of model capacity grew out of statistical learning theory and the bias-variance decomposition.

  • cmsl1qjuu062cnobppwwdccet← INSTANCE_OF

    robust statistics is a specific branch of statistics focused on methods that remain reliable even when assumptions are violated or data contains outliers. A robust-statistics is a statistics.

  • cmsl40yb90698nobpm3xr3ggj← DEPENDS_ON

    Model validation needs statistical methods to operate — remove statistical theory (hypothesis testing, goodness-of-fit measures, predictive accuracy metrics) and model validation cannot function. The removal test (Law 8) passes: model validation stops operating without statistics.

  • cmsl2neae065dnobp5evgy17g← DEPENDS_ON

    Model diagnostics needs statistical methods to operate now — remove statistical inference (hypothesis testing, goodness-of-fit, residual analysis) and the practice collapses. This is a present-tense operational dependency, not a historical claim.

  • cmsl5bc2a06d5nobpsnqnh00c← DEPENDS_ON

    Goodness-of-fit needs statistical theory to operate now — hypothesis testing, null distributions, test statistics are all statistical. Remove statistical inference and goodness-of-fit procedures cannot function. Present-tense operational dependency per Law 8.

  • cmsju7z0k03lgnobpu6darksq← DERIVED_FROM

    Design of experiments as a formal discipline came from R.A. Fisher's work in statistics in the 1920s. Statistics existed first and fed into the creation of experimental design as a method.

  • cmskiu86s04p2nobpto2jfu6d← DERIVED_FROM

    Bootstrapping as a statistical resampling technique was derived from statistical theory. Statistics (as a mathematical discipline) existed first and provided the theoretical foundation that bootstrap methods built upon in the 1970s.

  • cmsktewn005gcnobpu4bpi8p1← DEPENDS_ON

    Regression analysis as a method depends on statistical theory and practices to operate. Remove statistics and regression analysis loses its theoretical foundation — confidence intervals, p-values, assumptions checking all vanish. The removal test holds.

  • cmskge0nb04kfnobpmrm97sf0← DEPENDS_ON

    Propensity scores are computed entirely through statistical methods (regression, likelihood estimation). Remove statistics and propensity scores cannot be calculated or used. Constitutive operational dependency.

  • cmrg0scos00ef2a1nklfvbk7x← DEPENDS_ON

    Machine learning operates through statistical inference: parameter estimation, hypothesis testing, probability models, and optimization are all statistical methods. Remove statistics and ML has no mathematical foundation — it stops operating, not merely stopping being sayable. This is a constitutive dependency at the object level.

  • cmslxnls308ennobpgdweitn8← DEPENDS_ON

    Underfitting is a diagnostic concept in statistical and machine learning modeling: the condition where a model is too simple to capture patterns in data. Remove statistics and underfitting has no framework to operate — it cannot be defined, measured, or diagnosed. The removal test passes.

  • cmskixzrc04prnobp5789sx0p← DEPENDS_ON

    A tolerance interval is a statistical confidence interval that estimates a range containing a specified proportion of a population with given confidence. Remove statistics — remove confidence intervals, sampling theory, and population inference — and the tolerance interval concept ceases to operate. The removal test is operational, not meta-level: the concept's machinery collapses without its statistical framework.

  • cmsm0cr9y00331q137lslor2h← DEPENDS_ON

    Removal test: remove statistics (probability theory, inference, estimation) and statistical models cease to operate — they lose their fitting procedure, their error metrics, and their predictive framework. This is not sayability; the modeling machinery itself collapses.

  • cmrwh9erx0067soack0zwsr1t← DEPENDS_ON

    Removal test: remove statistics (probability theory, inference, estimation) and overfitting ceases to operate — it loses its fitting procedure, error metrics, and the very notion of model generalization. Like underfitting (accepted DEPENDS_ON statistics), overfitting is a condition that only exists within the framework of statistical modeling.

  • cmsm3jxq800cu1q136flgjmsn← DERIVED_FROM

    Standardized mean difference is derived from statistics: the concept of standardizing a difference using statistical parameters (standard deviation). Statistics as a discipline existed first and provided the tools SMD uses.

  • cmsmcylud014c1q13odkua088← DERIVED_FROM

    which-came-first test: statistics as a mathematical discipline (dating to the 17th century with Bernoulli, Huygens) predates significance level as a formal concept (Neyman-Pearson framework, 1928-1933). Significance level was built on top of the prior concept of statistics. DERIVED_FROM direction is correct.

  • cmsmcylud014c1q13odkua088← DEPENDS_ON

    Significance level as a procedure requires statistical methods to operate now — computing p-values, setting alpha thresholds, comparing test statistics against distributions. Remove statistics and the significance level procedure ceases to function. Complements the accepted DERIVED_FROM by capturing the operational dependency.

  • cmsm1xkpp007g1q1345shf2cw← DEPENDS_ON

    A Q-Q plot requires statistical concepts (distributions, quantiles) to operate. Remove statistics and the plot cannot be constructed or interpreted — it is not merely unsayable but inoperable.

  • cmsm3z3ux00e11q13oqik488m← DERIVED_FROM

    Data cleaning as a formal practice emerged from statistical methodology. Early statisticians like Pearson and Fisher developed outlier detection, missing data handling, and transformation methods that constitute data cleaning. Remove statistics and data cleaning loses its theoretical foundation for handling distributions, errors, and variation.

  • cmsm4u1fd00fy1q13kpeniei7← DERIVED_FROM

    Degrees of freedom originated in Fisher's development of statistical inference (1920s). It quantifies the number of independent values in a statistical calculation. Remove statistics and the concept has no mathematical framework or interpretation — it was created by and exists only within statistical theory.

  • cmsm5mn0100ig1q136h5upafi← DEPENDS_ON

    A hypothesis test IS a statistical procedure. Remove statistical theory and hypothesis testing has no framework, no test statistic, no p-value — it ceases to operate.

  • cmsm6rfiw00n51q138pvwzetk← DEPENDS_ON

    Autocorrelation is a statistical measure that requires the framework of statistics to operate. Removal test: remove statistics as a discipline/framework and autocorrelation as a concept ceases to function — it has no operational meaning outside statistics. Not meta-level: this is about the concept's own working, not its sayability (Law 8b).

  • cmsm7sqp000pp1q1300ss0izs← DEPENDS_ON

    Bootstrap is a statistical resampling technique that requires the framework of statistics to operate. Removal test: remove statistics and bootstrap as a method ceases to function — it has no operational meaning outside statistics. The concept operates through statistical inference about sampling distributions (Law 8b).

  • cmsm722w400o61q137a9lle5r← DERIVED_FROM

    The concept of measurement error emerged from statistical theory. Statistics existed first as a framework for understanding variation and uncertainty in data, and the formal treatment of measurement error (systematic vs random error, error propagation) was developed within that framework. The which-came-first test passes: statistical methods for analyzing error predate the formal concept of measurement error as a standalone construct.

  • cmsm1sv1x006w1q134hj5l2s1← DERIVED_FROM

    Probability plots are a graphical technique for comparing probability distributions. They DERIVED_FROM statistics — the statistical practice of assessing distributional fit using probability-based coordinates.

  • cmsm1xkpp007g1q1345shf2cw← DERIVED_FROM

    Quantile-quantile plots compare the quantiles of two distributions. They DERIVED_FROM statistics — the discipline that developed quantile-based diagnostic methods.

  • cmsm28gsu008t1q1325uyl4xb← DERIVED_FROM

    Spider plots (radar charts) are graphical techniques for displaying multivariate data on a two-dimensional chart. They are derived from statistical visualization practices — the coordinate systems, axis scaling, and plotting conventions used in spider plots are direct applications of statistical graphical methods. Statistics predates spider plots chronologically and conceptually.

  • cmsm3r9yr00di1q13jipxrrdk← DERIVED_FROM

    Test: which came first? Prediction intervals are statistical constructs for estimating the range within which a future observation will fall with a given probability. They only exist within statistical theory — the concept of prediction (forecasting) predates statistics, but prediction intervals as a formal statistical tool were derived from and are rooted in statistical theory. Statistics existed first and fed into the development of prediction intervals as a specific statistical method.

  • cmrcnskut00ww13vzhdk6xjwn← DERIVED_FROM

    which-came-first-existed-and-fed-into: statistical theory (Fisher, Neyman-Pearson 1920s-30s) predates the specific formalism of hypothesis testing. Hypothesis testing as a method is derived from the broader mathematical framework of statistics. The statistical theory existed first and fed into the hypothesis testing procedure.

  • cmsnoay9h04ha1q13xkt96h93← DERIVED_FROM

    Structural equation modeling is a statistical technique for analyzing relationships between latent and observed variables. Statistics (as a discipline of data analysis) predates SEM, which was developed by Sewall Wright in the 1920s as an extension of statistical methods. Statistics existed first and fed into SEM.

  • model-averaging← DERIVED_FROM

    Model averaging as a statistical technique was historically derived from statistics: the concept of weighting models by their likelihood or information criteria (AIC, BIC) derives from statistical theory (Fisher 1920s, Akaike 1970s). Statistics existed first and fed into model averaging.

  • ordinal← SERVES

    Ordinal measurement scales are built and maintained for the sake of statistical analysis and methodology. The servant (ordinal data type) points at the master (statistics as a field).

  • prior-distribution← SERVES

    A prior distribution encodes prior beliefs about model parameters and is used to serve statistical inference — specifically Bayesian statistics. It serves the statistical enterprise by providing initial information that, combined with likelihood, yields posterior distributions for parameter estimation and uncertainty quantification.

  • exponential-family← SERVES

    Exponential families are a class of probability distributions designed with properties that serve statistical inference: existence of sufficient statistics, conjugate priors, tractable MLE. The entire class was identified and studied because of its utility for statistics — parameter estimation, hypothesis testing, and Bayesian analysis. The designed purpose is to serve the statistical enterprise. SERVES direction: exponential-family (servant) points at statistics (master).

  • information-geometry← DEPENDS_ON

    Information geometry applies differential geometry to statistical manifolds — spaces of probability distributions parameterized by models. Remove statistics and information geometry has no subject matter: statistical manifolds cease to exist. The removal test is constitutive — information geometry cannot operate without the statistical models that define its domain. Direction: information-geometry (newer, applied) depends on statistics (older, foundational).

  • cmsdfit1403uo3vv38tclhy28← DEPENDS_ON

    Multiple testing correction adjusts statistical significance thresholds when conducting multiple simultaneous tests. It needs statistics as its operating framework — without statistical theory for p-values, error rates, and inference, the correction procedure has no mechanism to operate. The removal test passes.

  • cmsl0q8a705zdnobpzhfart0m← INSTANCE_OF

    DFFITS is a statistical measure of influence on regression fits. Statistics is the broader field of quantitative analysis. Direction correct (statistics older at 0.09, dffit at 0.86). This extends the ladder: dffit → metric → [intermediate] → statistics, or dffit → statistics directly as a statistical method.

Record identity

Created
Jul 23, 2026, 1:09 PM UTC
Content hash
20569317933576ba8648488a74627b7504bc74a68ec0c3c07843090accb58794

Open a related act record