A structured, live cohort taught by working Site Reliability Engineers from Google and Meta. Built specifically for mid-level backend engineers who have already failed an SRE loop and know they need more than a YouTube playlist.
Most engineers who fail an SRE loop are strong coders. They fail because the SRE interview tests a fundamentally different skill set: operating systems at scale, designing for failure, debugging live incidents under time pressure, and communicating risk to non-technical stakeholders in real time.
Self-study covers the theory. It does not give you a live incident to debug, a hostile interviewer who keeps narrowing your blast radius, or a staff SRE who watches you draw a reliability architecture and tells you where it breaks. That gap is exactly what this programme is built to close.
We have trained engineers from 3 years of experience to 12. The sweet spot is 3 to 8 years of backend or infrastructure work who know distributed systems conceptually but have never had to defend a reliability decision under pressure. If that is you, this cohort was designed for your situation.
Every session is taught by a practising SRE, not a coach who last ran production systems five years ago.
Capacity planning, load shedding, graceful degradation and the architectural decisions that separate a 99.9% from a 99.99% system.
SLIs, SLOs, error budgets, metrics pipelines, alerting philosophy and the difference between a symptom alert and a cause alert.
Live incident simulations, pager triage, incident command, stakeholder communication and postmortem writing with blameless root cause analysis.
Systematic debugging methodology, USE and RED methods, profiling, tracing and the structured approach interviewers want to see.
Explaining trade-offs to non-technical stakeholders, writing runbooks and postmortems, and the written component many SRE loops include.
Full mock on-call interviews run by practising SREs from FAANG companies, with live feedback on your diagnosis and decision-making process.
Two live sessions per week, each 90 minutes. Recordings available within 24 hours. All assignments are graded by the instructor who taught the session.
Understand exactly what each interview format tests, how Google, Meta, Stripe and Netflix run their loops differently, and where engineers with your background typically lose points.
The system design interview, SRE-style. You are not designing for features; you are designing for failure.
Metrics, logs, traces, SLIs, SLOs and error budgets. The concepts and the practical interview answers.
Structured debugging under time pressure, with the USE and RED methods and a reliable approach to narrowing a blast radius in front of an interviewer.
The mock on-call interview is the format most engineers are least prepared for. This week runs two full simulations.
Many SRE loops include a written component. This week covers blameless postmortem writing, runbook design and the written interview format.
Full 45-minute mock interviews with a practising SRE, covering the format your target company uses. Detailed written feedback within 48 hours.
Sharpen the edges. Cover the behavioural component, negotiate the offer, and leave with a plan for the live loop.
Every instructor is an active practitioner. None of them left industry to become a coach. They teach because they remember what the loop actually asked.
Staff Site Reliability Engineer
11 years at large-scale infrastructure teams
Senior SRE, Production Infrastructure
9 years, specialising in observability and incident command
Principal Reliability Engineer
13 years in distributed systems and on-call simulation design
SRE Tech Lead
8 years, formerly led SRE interview panel design
The decision you are actually making
Most engineers try self-study first. There is no deadline, so it slips. Here is what that choice actually costs.
I had failed two SRE loops at top companies before this cohort. The on-call simulations in week 5 were harder than my actual Google interview. When the real thing came, I had already done it twice under worse conditions.
I thought I understood SLOs. I did not. The week on observability rewired how I think about reliability, and the interviewer at Meta specifically said my error budget framing was the best they had seen that quarter.
The postmortem assignment in week 6 felt like too much work at the time. Then I got a written exercise in my Netflix loop that was nearly identical. I submitted in 40 minutes and it came back as the strongest written response of the loop.
I had been self-studying for a year. Eight weeks in this cohort did more than twelve months on my own. The difference is the structured feedback on my process, not just whether I got the right answer.
The offer negotiation session in week 8 alone was worth it. I negotiated an additional $45K in RSU on top of the base offer. Nobody tells you how to have that conversation.
Arjun has run SRE interviews at Google. When he tells you what the panel is actually scoring, you believe him. That context made every session hit differently.
No upsells. No add-on packages. One price covers the full programme.
Two 90-minute sessions per week for 8 weeks, run by your assigned instructor. Not pre-recorded.
Run by a practising SRE from Google, Meta or Netflix. 45 minutes each, with a 30-minute debrief.
A private 60-minute session with your instructor reviewing your reliability architecture approach.
Every postmortem, runbook and design exercise is graded and returned with written comments within 48 hours.
All 16 sessions recorded and available indefinitely. Re-watch before your loop or for the next one.
Ask questions, book office hours and access the alumni network for 6 months after cohort end.
The cohort is designed for working engineers. Two 90-minute live sessions per week is the core commitment. Assignments take 2 to 3 hours per week on top of that. All sessions are recorded within 24 hours, so missing one does not mean falling behind. Most cohort members are working 40 to 50 hour weeks at their current job.
Yes, if you have 3 or more years of backend or infrastructure engineering experience. The programme assumes you know how distributed systems work conceptually. It teaches you how to operate, defend and improve them at scale, which is the gap most backend engineers have. You do not need prior on-call experience; that is part of what the simulations build.
The SRE book is theory. This is practice under pressure. Reading about incident command does not prepare you for a live simulation where an interviewer is probing your decision-making in real time. The programme uses the book as background reading; the sessions use that background to run the practical exercises the book cannot simulate.
You get 6 months of alumni support after the cohort ends. Book office hours, ask your instructor what went wrong in the debrief and build a targeted plan. Alumni who complete the full cohort and still do not receive an offer within 6 months can re-attend the next cohort at no additional cost.
The curriculum is built around the SRE interview formats used at Google, Meta, Netflix, Stripe, Uber, Atlassian and similar large-scale engineering organisations. If your target company runs an SRE loop with a system design component, an on-call simulation, and a reliability or postmortem exercise, this cohort prepares you for all three.
Cohorts run quarterly. The next cohort has 30 seats. Click the apply button to see the upcoming start date and check availability. Seats are allocated in order of completed applications, not payment, so applying early is worth doing even if you are not yet certain.
Pricing is shown during the application process and depends on your location and payment plan preference. Monthly payment options are available. The median salary increase for alumni who receive an SRE offer at a Tier-1 company is $95K annually; the programme fee is typically recovered within the first two months of the new role.
Applications reviewed within 2 business days. You will hear back before the seat is released.