A data pipeline is a human-made system of sequential processing stages that move, transform, and filter data from source to destination. Each stage performs a discrete operation — ingestion, cleaning, transformation, enrichment, or delivery — connected by explicit data contracts that specify schema, ordering guarantees, and error-handling behavior. The system persists through orchestration software (Airflow, Prefect, Dagster) and configuration-as-code that defines stage topology, scheduling, and fault-recovery policies. [formal: tubus data | substrate: behavior | horizon: hours | explicit: yes | epoch: 0.01]
Accepted ontology entry
data pipeline
A data pipeline is a human-made system of sequential processing stages that move, transform, and filter data from source to destination. Each stage performs a discrete operation — ingestion, cleaning, transformation, enrichment, or deliver…
Definition
Why it is in scope
A human-made sequence of data processing steps where the output of each step feeds into the input of the next, orchestrated by code or configuration to move and transform data from source to destination.
Names and aliases
- data pipelineen · CANONICAL
Relations from this entry
- cmrvty0rg02cs2cei2m4ttn02SERVES →
A data pipeline is designed to move and transform data for downstream consumption. Its purpose is to enable analysis, reporting, and decision-making from raw data — the pipeline serves the analytics workflow, not the other way around. Per Law 8d.
- cmsm1o0gj006d1q13h15uqfd1INSTANCE_OF →
Data pipeline is a specific kind of pipeline — a competent speaker would call a data pipeline 'a pipeline' (modified by data). This is a clean INSTANCE_OF per Law 9.
Relations to this entry
No accepted relations in this direction.
Record identity
- Created
- Aug 3, 2026, 12:35 AM UTC
- Content hash
- 07f3960484429f2741ecddf245cde0f4d3a6f431d21ca6cae46ba84073654dbc