Regularization is a human-made machine learning and statistical technique that constrains model complexity by adding a penalty term to the loss function, thereby reducing the risk of overfitting and improving generalization to unseen data. Parameters: (1) a model with a loss function to minimize, (2) a penalty term proportional to model complexity (e.g., L1 norm for sparsity, L2 norm for weight decay), (3) a hyperparameter controlling the trade-off between fit and simplicity. Persistence mechanism: formalized in convex optimization and Bayesian statistics (regularization = MAP estimation with priors), implemented in all major ML libraries (scikit-learn, TensorFlow, PyTorch), and taught as a core technique in machine learning curricula. [formal: regularisatio | substrate: behavior | horizon: a life | explicit: yes | epoch: 0.01]
Accepted ontology entry
regularization
Regularization is a human-made machine learning and statistical technique that constrains model complexity by adding a penalty term to the loss function, thereby reducing the risk of overfitting and improving generalization to unseen data.…
Definition
Why it is in scope
A human-made machine learning and statistical technique that constrains model complexity to prevent overfitting and improve generalization to unseen data. Built to persist through mathematical formalization (penalty terms, priors), algorithmic implementation in optimization frameworks, and widespread use in production ML systems.
Names and aliases
- regularizationen · CANONICAL
Relations from this entry
- cmrwh9erx0067soack0zwsr1tDERIVED_FROM →
Which-came-first (Law 7): overfitting as a concept predates regularization. The recognition that models could memorize training data came before the development of techniques to prevent it. Historical chain: observation of overfitting → need for regularization.
- cmrnbfxhn022qd1nltabn54joSERVES →
Regulation is built and maintained for the sake of improving generalization (Law 8d): its designed purpose is to further generalization by constraining complexity. For whose sake? The master is generalization, regularization is the servant.
Relations to this entry
- cmsli0x7x076bnobp5cbkmn1x← INSTANCE_OF
Early stopping IS a specific kind of regularization technique: it constrains model capacity by halting training before overfitting occurs. A competent speaker would call it 'a regularization method.' The nearest kind is regularization, not machine learning.
- cmslj13wx0797nobpl727n24x← INSTANCE_OF
dropout is a specific kind of regularization technique. A competent speaker would call dropout 'a regularization method.' It randomizes neuron deactivation to prevent co-adaptation, fitting squarely under the regularization umbrella.
- cmslk15iz07d0nobpj7trjre0← INSTANCE_OF
Weight decay IS a specific kind of regularization technique: it adds a penalty proportional to the squared magnitude of weights to the loss function. A competent speaker would call it 'a regularization method.' Files against the nearest kind (regularization) per Law 9.
- cmslkr22307f5nobpcytgqx50← INSTANCE_OF
Gradient clipping IS a specific kind of regularization technique: it bounds gradient magnitudes to prevent exploding gradients during training. A competent speaker would call it 'a regularization method.' Files against the nearest kind (regularization) per Law 9.
- cmsneovqi03pl1q13bt25mogd← SERVES
Weight sharing is maintained for the sake of regularization: in neural networks it reduces the effective number of parameters, constraining model capacity and preventing overfitting. Its purpose is to act as a regularization mechanism.
- cmslj13wx0797nobpl727n24x← SERVES
Dropout is built and maintained for the sake of regularization: randomly dropping neurons during training reduces co-adaptation and prevents overfitting. Its designed purpose is to serve as a regularization mechanism.
Record identity
- Created
- Jul 22, 2026, 7:32 PM UTC
- Content hash
- 36568a2a096e02f821b1dae66904c2d4be3465a324b849504e6a8da5cec5aa4a