Service reliability engineering (SRE) is a discipline within software engineering that applies reliability targets, error budgets, and observability practices to maintain system uptime and performance at scale. Parameters: reliability targets (SLOs) set acceptable failure thresholds; error budgets quantify tolerated downtime derived from (1-SLO)*operations; observability provides logs, metrics, and traces to infer system state. Persistence mechanism: sustained through organizational processes, automated tooling pipelines, and engineering culture that treats reliability as a first-class design requirement rather than an afterthought. [formal: servitium-reliabilitas-engenia | substrate: behavior | horizon: hours | explicit: yes | epoch: 0.01]
Accepted ontology entry
service reliability engineering
Service reliability engineering (SRE) is a discipline within software engineering that applies reliability targets, error budgets, and observability practices to maintain system uptime and performance at scale. Parameters: reliability targ…
Definition
Why it is in scope
A human-made discipline in software engineering that applies reliability targets (SLOs), error budgets, and observability practices to ensure system uptime and performance. Built to persist through organizational processes, tooling, and cultural practices.
Names and aliases
- service reliability engineeringen · CANONICAL
Relations from this entry
- cmsinj44w00k3nobpqlkg56feDEPENDS_ON →
SRE operates by observing system behavior to ensure reliability. Remove observability and SRE has no data to engineer from — it stops OPERATING. The removal test passes: SRE needs observability as its primary data source.
- cmrg5g8tq00mz2a1nkmqfmaj4INSTANCE_OF →
Service reliability engineering IS a specific kind of practice — a professional engineering practice focused on ensuring system reliability through observability, testing, and incident response. Competent-speaker test passes.
- cmsip7bsg00pinobpjm0rpp62DEPENDS_ON →
Service reliability engineering needs incident-response to operate: incident response is a core pillar of SRE practice. Removal test — remove incident response from the SRE toolkit and SRE loses a fundamental operational component. This is a genuine dependency, not merely historical association.
Relations to this entry
No accepted relations in this direction.
Record identity
- Created
- Aug 7, 2026, 8:03 AM UTC
- Content hash
- 79cc600d3d5408a96fed2b06f4963c15ed4aa9ee72b7258b058e3cd13b150721