A six-week cohort for people who run real software
Production is
a practice.
Build durable software you can change, recover, and maintain.
Shipping is a moment. Operating is everything that follows: protecting data, diagnosing failures, controlling costs, and changing the system without losing sleep.
Get cohort updates- Duration
- 6 weeks
- Format
- Cohort-based
- You bring
- One deployed app
- You leave with
- An operating manual
01 / The shift
Working software is not yet durable software.
A product can be online and still be impossible to explain, dangerous to change, expensive to run, and one bad migration away from disaster.
Operator Lab teaches the work after launch. You will apply every lesson to your own application and turn unknowns into decisions, runbooks, limits, and tested recovery paths.
02 / The program
Six weeks.
One system you trust.
Each week produces evidence from your application—not generic homework.
- W01
Take responsibility
Define the production promise, map the system, separate environments, and make access explicit.
Production baseline + system map - W02
Protect users and data
Test authorization boundaries, validate untrusted input, and decide how data is retained and removed.
Data-protection review - W03
Change software safely
Make deployments repeatable, plan migrations, preserve compatibility, and prove a rollback path.
Migration + rollback runbook - W04
See production clearly
Create useful logs, metrics, and alerts, then diagnose a failure from evidence instead of guesses.
Monitoring + alerting plan - W05
Recover from failure
Restore real data, make background jobs reliable, run an incident, and improve the system afterward.
Recovery evidence + incident review - W06
Keep the product viable
Control abuse and cost, reduce unnecessary systems, prioritize debt, and establish a maintenance rhythm.
90-day operating plan
03 / Your evidence
The final deliverable
Your application's
Operating Manual.
Not a certificate for watching. A working body of evidence that shows you can operate the software you are responsible for.
- 01System + dependency map
- 02Production risk register
- 03Environment + secrets inventory
- 04Deployment + rollback runbook
- 05Database migration procedure
- 06Monitoring + alerting plan
- 07Backup + recovery evidence
- 08Incident response checklist
- 09Cost limits
- 10Maintenance schedule
- 11Technical debt decisions
- 1290-day operating plan
04 / Final exercise
Operator Day
Something breaks.
You take the lead.
Diagnose a controlled production failure, assess the impact, contain it, recover, communicate, and change the system so it is less likely to happen again.
- Detect
- Assess
- Contain
- Recover
- Verify
- Communicate
- Review
- Improve
05 / Readiness check
Bring a product
with consequences.
This is for you if
- You have a deployed application.
- You have real users or realistic production data.
- You know basic Git and deployment.
- You are responsible for keeping the app working.
This is not
- An introduction to programming.
- A build-your-first-app course.
- A tour of every Cloudflare product.
- A promise that incidents never happen.
06 / First cohort
Enrollment is not open yet
Be there when
the lab opens.
Join the waiting list and I’ll email you when dates, enrollment, and the first cohort are ready.