A data lake is a centralized storage architecture in information systems that ingests and retains vast volumes of structured, semi-structured, and unstructured data in its native format, deferring schema definition to the read phase. It persists through distributed storage systems (HDFS, S3, Blob) backed by catalog metadata and access controls. Schema-on-read, rather than schema-on-write, allows arbitrary analytics, machine learning, and reporting workloads over the same raw data store without prior transformation.
[formal: lacus_data | substrate: matter | horizon: years | explicit: yes | epoch: 0.01]