Cheatsheets
A condensed, exam-oriented companion to the course notes — one card per concept, comparison tables, and a full tool directory. Pair with the exam simulator to test yourself after reading.
Practice with the exam simulatorBy Ernest Mueller & James Wickett · The values, principles, and practices that make DevOps work — culture first, then everything else.
1.1The core model
What is DevOps?
Developers and operations engineers working together across the entire service lifecycle — design, development, and production support. It's a cultural and professional movement, not a tool, a framework you buy, or a job title. Outcomes: faster deployments, fewer failures, quicker recovery, less burnout.
CAMS — the four core values
- Culture — change people's behavior first; collaboration between Dev and Ops. Everything else is built on this.
- Automation — remove inefficiency and improve quality, but only on a solid cultural foundation.
- Measurement — track key metrics to understand and improve technical and business outcomes.
- Sharing — information and collaboration across teams drive continuous improvement.
The Three Ways
- First Way — Systems thinking & flow: optimize the whole system, not individual parts; fast flow from Dev → Ops → customer.
- Second Way — Amplify feedback loops: create, shorten, and amplify feedback so issues surface fast.
- Third Way — Experimentation & learning: a culture of continual experimentation, risk-taking, and learning from failure.
The five practice areas (your playbook)
Culture · Process · Infrastructure as Code · Continuous Delivery · Site Reliability Engineering. They're interdependent — advance them together, iteratively. Unbalanced adoption (all tooling, no culture) leads to frustration.
1.2Culture
The wall of confusion
Teams throw work "over the wall" to the next group. The real cause is institutional incentives: Dev is rewarded for change, Ops for stability. Fix the incentives (shared goals), embed ops engineers in dev teams, and build cross-functional ownership.
Communication & trust
Establish clear, agreed channels for information flow. Good communication builds trust, and trust is what makes information actually flow — the engine of organizational performance.
Kaizen — continuous improvement
Small, ongoing changes made by everyone. Five principles: know your customer · enable smooth workflow · go to the real place (gemba) · empower people · be transparent.
1.3Process building blocks
Agile
DevOps grew out of Agile-infrastructure discussions (2008–09) and extends Agile past "working software" into deployment and operations. Iterative work, active collaboration, faster time to market.
Lean
Eliminate waste so value reaches the customer. Three wastes: muda (unnecessary activity), mura (irregular flow), muri (overburdening). Techniques: kaizen, value stream mapping (visualize flow, expose wait time), visual management.
Lightweight change control
- Change control should be fast, scalable, and repeatable — not a biweekly board meeting.
- The most effective reviews are peer reviews by a technologist close to the team, performed quickly and in a distributed fashion.
- Keep changes small — easier to review, easier to fix.
- CI with automated testing validates changes early.
1.4Metrics that matter
| Metric | What it measures |
|---|---|
| Lead time for changes | Commit → running in production. Core speed measure. |
| Deployment frequency | How often you ship to production. Elite teams: on demand. |
| MTTR (time to restore) | How fast service recovers after failure. Optimize recovery, not just prevention. |
| Change failure rate | % of deployments causing a production failure needing remediation. |
1.5Site Reliability Engineering
What SRE is
Applying software engineering to IT operations (born at Google). Build for reliability from the start, use operational feedback to improve, and spend ≥50% of time building tools rather than firefighting.
Observability — five areas to measure
- Synthetic checks — "is it working?" simulated user probes.
- System & app metrics — CPU, memory, custom app metrics.
- End-user performance — RUM + APM, the user's real experience.
- Logs — what happened, when, where; troubleshooting and forensics.
- Security monitoring — threats detected from logs/metrics.
Incident response & postmortems
- Three activities: troubleshooting, automation (runbooks), communication. Process inspired by the Incident Command System.
- Postmortems: no single root cause (incidents come from multiple deficiencies), blameless (fix the system, not the person), transparent.
Building for reliability
Design decisions determine production behavior. Integration points are the #1 failure source — use the circuit breaker pattern (stop calling a failing service, allow recovery) plus timeouts to prevent cascading failures. Resilience = redundancy, load balancing, auto-scaling, failover. Systems are sociotechnical: people are part of resilience. Key reading: Release It! (Nygard), the Twelve-Factor App, Martin Fowler.
1.6Advanced topics
Platform engineering — the paved road
Self-service golden paths (common CI/CD, observability) that make the right way the easy way. Treat the platform as a product: discover user needs, iterate. Don't overbuild — blaze a trail, then pave the road, then build infrastructure.
DevSecOps
CAMS with a security lens: security works alongside dev (culture), tools shift left into IDE + CI (automation), joint security metrics (measurement), and shared responsibility (sharing). Use security champions in each team; inspire, don't rely on FUD.
Cloud native & Kubernetes
K8s automates deployment, scaling, and management of containers — powerful but complex, costly, and it can create silos if adopted without DevOps thinking. Assess whether serverless or lighter orchestration is enough before committing.
Chaos engineering
Deliberate experiments (fault injection) to build confidence in resilience — Netflix's Chaos Monkey. Game days exercise the human side of incident response. It's a feedback loop for organizational learning.
MLOps
DevOps for ML: versioning big datasets and models, heavy compute, continuous learning from user input, richer feedback loops, and close collaboration between data scientists, dev, and ops.
AIOps — three waves
Wave 1: AI-generated code, tests, docs in the IDE. Wave 2: automating systems — alerting, cluster health, runbooks. Wave 3: self-service platforms and cross-team collaboration. Prompt with context to get good results.
1.7Resources
Books: The DevOps Handbook · Accelerate · The Phoenix Project. Research: dora.dev. Community: DevOpsDays, DevOps Enterprise Summit. Authors to follow: Martin Fowler, Julia Evans.