콘텐츠로 건너뛰기

FieldsCloud and DevOpsSite Reliability Engineer

Site Reliability Engineer

Keep systems fast, stable, and online when it matters.

Career

Structure
8 modules, each ending in a milestone
Proof
Verified against real tasks from module 2
Ends in
A defended capstone and a high-assurance credential

Get early access

The journey

  1. Practised under observation

    1. The reliability seat and systems thinkingWhat this discipline actually is: reliability as an engineering problem with a number attached, and the systems reasoning that makes failure predictable.Practised under observation
      MilestoneReasons about failure

      The diagram says this part is redundant and you can show that it is not, a failure is traced through the system to the thing it takes down next, and this work is told apart from operations with a new name on it.

  2. From here, milestones are verified rather than practised.

    Verified against a real task

    1. Observability engineeredYou cannot make reliable what you cannot see, so instrumentation is designed for the questions failure will ask rather than the ones you thought of.Verified against a real task
      MilestoneInstrumented for the questions you haven't asked

      A failure nobody instrumented for is still findable in what you already collect, and a green dashboard does not outrank a user saying the product is broken.

    2. Service level objectives and error budgetsThe core intellectual contribution of this discipline: reliability defined as a number the business agreed to, and the budget that turns the gap into a decision.Verified against a real task
      MilestoneReliability, defined and governed

      Reliable is a number somebody agreed to rather than a word, the next nine carries a price you can state, and when the budget is gone the policy is what happens instead of the argument.

    3. Toil elimination and reliability automationThe mandate that keeps this seat an engineering role, with operational work measured, engineered away, and replaced by systems that respond without a human.Verified against a real task
      MilestoneToil measured, engineered away

      Every hour of repetitive operational work is counted and then engineered away for good, and the automation that acts knows when acting would make things worse.

    4. Designing and debugging for reliabilityReliability built into systems rather than bolted on, with resilience patterns proven under pressure and the debugging craft that finds causes instead of symptoms.Verified against a real task
      MilestoneDesigned to survive, debugged to cause

      A system you have never seen is taken to the cause rather than to a symptom, and a well-meant retry is shown making things worse before it is fixed.

    5. Incident response and the learning organizationThe craft that defines the seat in public: incidents commanded calmly, communicated honestly, and converted into systemic change rather than blame.Verified against a real task
      MilestoneCommands the incident, learns from it

      The person running the incident coordinates instead of disappearing into the debugging, and a finding that blames a person is rewritten as one that blames a system.

    6. Reliability at organizational scaleReliability as a practice rather than a person, with an operating model, a launch gate, sustainable on-call, and pattern-finding that fixes classes of failure.Verified against a real task
      MilestoneThe practice, not the person

      Saying a service is not ready comes with a path to ready, the pages an on-call rotation receives are counted and defended as survivable, and a failure that keeps happening is fixed once for its whole class.

    7. The Site Reliability Engineer in the organizationThe hardest work in this seat: holding the reliability line with evidence, negotiating with product rather than obstructing it, and staying honest when the pressure is to say yes.Verified against a real task
      MilestoneHolds the line, honestly

      You either approve with the risk named or decline with a route to yes, and when you are overruled the record still says what you said.

    1. Capstone

      A reliability practice a business trusts

      A complete reliability engagement on a realistic production system, worked under observation, from failure-mode and dependency analysis through engineered observability, objectives and error budgets, toil elimination and resilience work to incident command and the organisational layer of the seat.

      Defence

      You walk an instructor through your own numbers: what reliable means for this system and who agreed to it, why not one more nine and what that would cost, how you know who is right when your dashboards are green and a user says it is broken, and which failure you could not prevent at an acceptable price. The credential is not awarded if you cannot account for your own work.

    2. Credential

      High-assurance credential

      Evidence that you can turn reliability into a number a business agreed to, engineer toward it, and say plainly when it was missed. It does not claim seniority, and it does not oblige any employer to accept it.

      See how proof works

Related capability paths

What is not live yet

The desktop app, consent-based observation, scoring, and credentials are in development. Nothing here implies they are live yet.

Get early access