SYSTEMA CONSTRUCTUM

Accepted ontology entry

reward model

A reward model is a supervised learning model that maps text inputs to scalar scores representing human preference, trained on paired comparisons of model outputs where annotators select the preferred response. Its parameters are learned t…

ACCEPTED THINGcmsmltiiw01r31q132cnohdcz

Definition

A reward model is a supervised learning model that maps text inputs to scalar scores representing human preference, trained on paired comparisons of model outputs where annotators select the preferred response. Its parameters are learned through logistic regression or neural ranking to maximize the probability that human-preferred outputs receive higher scores. It persists as a trained checkpoint file deployed alongside a language model in reinforcement learning from human feedback (RLHF) pipelines, where it provides the dense reward signal needed to optimize generation behavior. The model operates by encoding text pairs, comparing their representations, and outputting a single preference score that guides gradient-based policy updates. [formal: modelus praemii | substrate: mind | horizon: a life | explicit: yes | epoch: 0.01]

Why it is in scope

A human-made model trained to predict a scalar reward signal for guiding the fine-tuning of language models via reinforcement learning. It maps text completions to preference scores derived from human comparative judgments, serving as the optimization target in RLHF pipelines. Built to persist through training data, architecture specifications, and deployment as the scoring function in preference-based model alignment.

Names and aliases

Relations from this entry

  • cmsdcks2103r73vv3v8qlmtjiDERIVED_FROM →

    Reward models are trained as supervised classifiers/regressors on human preference data (paired comparisons, rankings). Supervised learning as a paradigm (1950s–1960s, pattern recognition) predates reward models (2010s, RLHF). The training mechanism of reward models — mapping inputs to scalar labels via supervised optimization — is inherited from supervised learning.

  • cmsi4l2p80452ywh500rxvdvfSERVES →

    Reward models are built for the sake of optimization — they provide learned reward signals that guide the optimization of model behavior during RLHF. Without a reward model, policy optimization in RLHF has no learned objective to converge toward.

  • cmsdcks2103r73vv3v8qlmtjiINSTANCE_OF →

    A reward model is a specific kind of supervised learning model — it is a classifier/regressor trained on labeled human preference data. A competent ML practitioner would call it 'a supervised learning model.' The specific points at the general.

  • cmsgdwl3c00bmywh5iw7wvdomSERVES →

    Reward models are built to approximate scalar reward signals that guide RL agents — their designed purpose is to further RL operation by providing learnable reward feedback.

  • cmsmpzucr01zs1q1333etwodtINSTANCE_OF →

    A reward model IS A specific kind of predictive model: it takes inputs (e.g. context + response pairs) and predicts a preference/reward score. The direction was tested: reward model is the specific, predictive model is the general. A competent speaker would call a reward model 'a kind of predictive model.' The note pins this sense: prediction of preference scores, not reward functions.

Relations to this entry

No accepted relations in this direction.

Record identity

Created
Aug 10, 2026, 2:20 AM UTC
Content hash
77e12b93060ae570b66e0160262fde4f69e1a6838f48e5dedb588c026763f3e9

Open a related act record