A dataset is a structured collection of data records assembled by humans for use in analysis, modeling, or experimentation. It consists of observations (rows) described by variables (columns), curated and formatted so that computational tools can ingest and operate on it. Its persistence mechanism is file storage and version control — datasets are saved as structured files (CSV, JSON, Parquet, SQL tables) and tracked across iterations. [formal: dataset | substrate: matter | horizon: a life | explicit: yes | epoch: 0.01]
Accepted ontology entry
dataset
A dataset is a structured collection of data records assembled by humans for use in analysis, modeling, or experimentation. It consists of observations (rows) described by variables (columns), curated and formatted so that computational to…
Definition
Why it is in scope
A human-made collection of data instances organized for storage, analysis, or training. Created, curated, and maintained by humans, it persists as structured data (tables, arrays, records) and serves as the foundational input for machine learning, statistics, and data science.
Names and aliases
- dataseten · CANONICAL
Relations from this entry
No accepted relations in this direction.
Relations to this entry
- cmsernq4y05tp3vv3usooao6w← INSTANCE_OF
A validation set IS a specific kind of dataset per Law 9: it is a human-curated subset of data held out for model validation during training. A competent speaker would call it a dataset. Nearest kind confirmed — dataset.
- cmskcrbx404blnobpspqah08q← INSTANCE_OF
A training set is a specific kind of dataset — a collection of data instances used specifically for training machine learning or statistical models. The test: a competent data scientist would call a training set 'a dataset' (specifically, a dataset used for training).
Record identity
- Created
- Aug 8, 2026, 12:47 PM UTC
- Content hash
- aa5238f67aa21b6e3ab694eefc45f7f67b2ecc829adb0fe1cee914bafe93aceb