FieldsSite Reliability Engineer
Site Reliability Engineer
Keep systems fast, stable, and online when it matters.
Career · Capability
What you learn
- Reliability defined as a number a business has agreed to
- Observability engineered for the failures nobody anticipated
- Toil measured in hours and engineered away
- Incident command that stays calm and honest when everything is loud
What you will understand
- Service level indicators, objectives, and the error budget that governs the trade
- Why one hundred percent is the wrong target and what each additional nine costs
- Percentiles over averages, because the tail is the real user experience
- Blameless review, because blame suppresses the reporting reliability depends on
How you will practice
- Instrument a system, then diagnose a failure you never specifically instrumented for
- Set an objective with an error budget policy and govern an exhaustion scenario under pressure
- Command a live incident, prioritising mitigation over diagnosis, then run the review
What you will build
- An objective set with burn-rate alerting that fires on trajectory rather than on noise
- A resilience engagement where induced failures are contained by patterns you implemented
The proof you build
Evidence that you can define reliability in numbers, engineer a system to meet them, and lead the response when it fails anyway.
Your work is observed with your consent, scored for independence and assistance, and turned into proof that carries a confidence level. The career path can reach a high-assurance credential, anchored by a scored capstone.
Related capability paths
What is not live yet
The desktop app, consent-based observation, scoring, and credentials are in development. Nothing here implies they are live yet.
Get early access