A scalar feedback value engineered within reinforcement learning and behavioral systems to quantify the desirability of a state transition or action outcome. It functions as the optimization target that an agent or subject learns to maximize through repeated interaction. The concept persists through implementation in learning algorithms, behavioral protocols, and the mathematical formalism of Markov decision processes. [formal: signalis praemii | substrate: behavior | horizon: a life | explicit: yes | epoch: 0.42]
Accepted ontology entry
reward-signal
A scalar feedback value engineered within reinforcement learning and behavioral systems to quantify the desirability of a state transition or action outcome. It functions as the optimization target that an agent or subject learns to maximi…
Definition
Why it is in scope
A scalar feedback value used in reinforcement learning and behavioral engineering to quantify the desirability of an outcome, designed by humans to guide agent or subject behavior toward optimized policy or response selection.
Names and aliases
- reward-signalen · CANONICAL
Relations from this entry
- cmrz8xfyt03coekkxwc7eni7jDERIVED_FROM →
Operant conditioning (Skinner, 1930s) existed first and fed into the modern computational concept of reward signals used in neuroscience and machine learning. The which-came-first test: operant conditioning predates reward-signal by decades. DERIVED_FROM direction is correct per Law 7.
- cmsgdwl3c00bmywh5iw7wvdomSERVES →
A reward signal is built and maintained for the sake of guiding reinforcement learning — it is the core mechanism by which RL agents learn which actions to take. Law 8d: the servant (reward signal) points at the master (reinforcement learning). Remove the reward signal and RL has no learning signal and stops operating.
Relations to this entry
No accepted relations in this direction.
Record identity
- Created
- Aug 5, 2026, 5:57 PM UTC
- Content hash
- ea9a078db250719cf17a01e417d8c0e5d8a72b7321f0366c717d2ae97a93561c