Observability is a software engineering practice that defines how to determine the internal state of a running system from its external outputs — specifically logs (structured event records), metrics (quantified measurements over time), and traces (request journeys across services). It differs from monitoring in scope: monitoring answers known questions about a system's health, while observability enables investigation of unknown failure modes by exposing the system's behavior through queryable data. The practice persists through tooling frameworks (Prometheus, Grafana, Jaeger), operational procedures, and the shared mental models of engineering teams. [formal: observabilitas | substrate: mind | horizon: a life | explicit: yes | epoch: 0.01]
Accepted ontology entry
observability
Observability is a software engineering practice that defines how to determine the internal state of a running system from its external outputs — specifically logs (structured event records), metrics (quantified measurements over time), an…
Definition
Why it is in scope
Observability is a human-made practice in software engineering and operations that defines how to infer a system's internal state from its external outputs — logs, metrics, and traces. It is a conceptual framework built by practitioners to reason about complex systems, not the systems themselves. The concept was coined by Leslie Brown (1976) in control theory and later adapted for software by Michael Nygard (2012).
Names and aliases
- observabilityen · CANONICAL
Relations from this entry
- cmr9uz3vv00elhcxfruyltnd4DEPENDS_ON →
observability operates through measurement — metrics, logs, traces are all measured signals. Remove measurement and observability stops OPERATING (not just sayable); it needs measured data to extend system visibility. The removal test passes.
- cms88c18w000h73fkh1a5pwxvSERVES →
Observability is built for the sake of reliability engineering: by making internal system states inferable from external outputs, it enables reliability engineers to diagnose failures, understand system behavior, and improve resilience. Purpose by design.
Relations to this entry
- cmsinr1g700konobpyevl1ud7← DEPENDS_ON
SRE operates by observing system behavior to ensure reliability. Remove observability and SRE has no data to engineer from — it stops OPERATING. The removal test passes: SRE needs observability as its primary data source.
Record identity
- Created
- Aug 7, 2026, 7:57 AM UTC
- Content hash
- cf0c07aafccef458d3d4ba199f27be0582ef8414cffc480a6b4fa8ab337a7278