Overfitting is a human-made statistical and machine learning concept describing the failure mode where a model with sufficient capacity memorizes noise, outliers, or idiosyncrasies in training data rather than learning the underlying generalizable pattern. Parameters: (1) a model whose capacity exceeds what the training data can support, (2) a measurable gap between training performance and test/generalization performance, (3) complexity that captures dataset-specific artifacts rather than population-level signal. Persistence mechanism: formalized in statistical learning theory (VC dimension, bias-variance decomposition), codified in machine learning curricula and textbooks, and operationalized through regularization frameworks (L1, L2, dropout, early stopping) in software libraries. [formal: supervinculum | substrate: behavior | horizon: a life | explicit: yes | epoch: 0.01]
Accepted ontology entry
overfitting
Overfitting is a human-made statistical and machine learning concept describing the failure mode where a model with sufficient capacity memorizes noise, outliers, or idiosyncrasies in training data rather than learning the underlying gener…
Definition
Why it is in scope
A human-made statistical and machine learning concept describing the failure mode where a model memorizes training noise rather than learning the underlying signal, resulting in poor generalization to unseen data. Built to persist through formalization in statistical learning theory, documentation in ML curricula, and systematic mitigation via regularization techniques.
Names and aliases
- overfittingen · CANONICAL
Relations from this entry
- cmrnbfxhn022qd1nltabn54joDEPENDS_ON →
Removal test (Law 8): remove the concept of generalization (model performance on unseen data) and overfitting ceases to operate — overfitting is defined as the failure to generalize. This is an object-level ML concept relationship, not meta-level (Law 2b).
- cmru5nrqe003sr671qxjyxhiqDEPENDS_ON →
Overfitting is a phenomenon where a model memorizes training data instead of learning generalizable patterns. Remove the concept of a model — the learned representation that fits data — and overfitting ceases to have any operational meaning. It requires a model to exist and operate.
- cmruf04x1013cr671g8zou4lnINSTANCE_OF →
Overfitting is a specific kind of modeling error — the condition where a model captures noise rather than signal. A competent practitioner would call overfitting a kind of error, just as underfitting is one (already accepted by the court).
- cmrxj3acr03cmsoacx73fal1oDEPENDS_ON →
Removal test: remove statistics (probability theory, inference, estimation) and overfitting ceases to operate — it loses its fitting procedure, error metrics, and the very notion of model generalization. Like underfitting (accepted DEPENDS_ON statistics), overfitting is a condition that only exists within the framework of statistical modeling.
Relations to this entry
- cmrwhb9js006isoacsf207m10← DERIVED_FROM
Which-came-first (Law 7): overfitting as a concept predates regularization. The recognition that models could memorize training data came before the development of techniques to prevent it. Historical chain: observation of overfitting → need for regularization.
Record identity
- Created
- Jul 22, 2026, 7:30 PM UTC
- Content hash
- cdd7354b9a1043f3bb7bd0deb23d8063e5b0ac02acc97ffef42a7a54903840f1