Back to DevOps Learning
$cat cheatsheets.html

Cheatsheets

A condensed, exam-oriented companion to the course notes — one card per concept, comparison tables, and a full tool directory. Pair with the exam simulator to test yourself after reading.

Practice with the exam simulator

By Ernest Mueller & James Wickett · The values, principles, and practices that make DevOps work — culture first, then everything else.

1.1The core model

What is DevOps?

Developers and operations engineers working together across the entire service lifecycle — design, development, and production support. It's a cultural and professional movement, not a tool, a framework you buy, or a job title. Outcomes: faster deployments, fewer failures, quicker recovery, less burnout.

CAMS — the four core values

  • Culture — change people's behavior first; collaboration between Dev and Ops. Everything else is built on this.
  • Automation — remove inefficiency and improve quality, but only on a solid cultural foundation.
  • Measurement — track key metrics to understand and improve technical and business outcomes.
  • Sharing — information and collaboration across teams drive continuous improvement.

The Three Ways

  • First Way — Systems thinking & flow: optimize the whole system, not individual parts; fast flow from Dev → Ops → customer.
  • Second Way — Amplify feedback loops: create, shorten, and amplify feedback so issues surface fast.
  • Third Way — Experimentation & learning: a culture of continual experimentation, risk-taking, and learning from failure.
Three levels of DevOps understanding: values (what we believe) → principles (how we formalize beliefs into a plan) → practices (how we put them into action).

The five practice areas (your playbook)

Culture · Process · Infrastructure as Code · Continuous Delivery · Site Reliability Engineering. They're interdependent — advance them together, iteratively. Unbalanced adoption (all tooling, no culture) leads to frustration.

People over process over tools. Identify the right people and processes first, then pick tools that fit. Keep the toolchain simple and well-integrated.

1.2Culture

The wall of confusion

Teams throw work "over the wall" to the next group. The real cause is institutional incentives: Dev is rewarded for change, Ops for stability. Fix the incentives (shared goals), embed ops engineers in dev teams, and build cross-functional ownership.

Communication & trust

Establish clear, agreed channels for information flow. Good communication builds trust, and trust is what makes information actually flow — the engine of organizational performance.

Kaizen — continuous improvement

Small, ongoing changes made by everyone. Five principles: know your customer · enable smooth workflow · go to the real place (gemba) · empower people · be transparent.

1.3Process building blocks

Agile

DevOps grew out of Agile-infrastructure discussions (2008–09) and extends Agile past "working software" into deployment and operations. Iterative work, active collaboration, faster time to market.

Lean

Eliminate waste so value reaches the customer. Three wastes: muda (unnecessary activity), mura (irregular flow), muri (overburdening). Techniques: kaizen, value stream mapping (visualize flow, expose wait time), visual management.

Lightweight change control

  • Change control should be fast, scalable, and repeatable — not a biweekly board meeting.
  • The most effective reviews are peer reviews by a technologist close to the team, performed quickly and in a distributed fashion.
  • Keep changes small — easier to review, easier to fix.
  • CI with automated testing validates changes early.
ITIL can coexist with DevOps — adapt its change/service management ideas into lightweight versions rather than importing its full 34-process weight.

1.4Metrics that matter

MetricWhat it measures
Lead time for changesCommit → running in production. Core speed measure.
Deployment frequencyHow often you ship to production. Elite teams: on demand.
MTTR (time to restore)How fast service recovers after failure. Optimize recovery, not just prevention.
Change failure rate% of deployments causing a production failure needing remediation.
These four are the DORA metrics (dora.dev). Measure the system, never individuals — individual metrics get gamed and punish collaboration. An SLO is your internal reliability target (e.g. 99.9%); an SLA is the contractual version.

1.5Site Reliability Engineering

What SRE is

Applying software engineering to IT operations (born at Google). Build for reliability from the start, use operational feedback to improve, and spend ≥50% of time building tools rather than firefighting.

Observability — five areas to measure

  • Synthetic checks — "is it working?" simulated user probes.
  • System & app metrics — CPU, memory, custom app metrics.
  • End-user performance — RUM + APM, the user's real experience.
  • Logs — what happened, when, where; troubleshooting and forensics.
  • Security monitoring — threats detected from logs/metrics.

Incident response & postmortems

  • Three activities: troubleshooting, automation (runbooks), communication. Process inspired by the Incident Command System.
  • Postmortems: no single root cause (incidents come from multiple deficiencies), blameless (fix the system, not the person), transparent.

Building for reliability

Design decisions determine production behavior. Integration points are the #1 failure source — use the circuit breaker pattern (stop calling a failing service, allow recovery) plus timeouts to prevent cascading failures. Resilience = redundancy, load balancing, auto-scaling, failover. Systems are sociotechnical: people are part of resilience. Key reading: Release It! (Nygard), the Twelve-Factor App, Martin Fowler.

1.6Advanced topics

Platform engineering — the paved road

Self-service golden paths (common CI/CD, observability) that make the right way the easy way. Treat the platform as a product: discover user needs, iterate. Don't overbuild — blaze a trail, then pave the road, then build infrastructure.

DevSecOps

CAMS with a security lens: security works alongside dev (culture), tools shift left into IDE + CI (automation), joint security metrics (measurement), and shared responsibility (sharing). Use security champions in each team; inspire, don't rely on FUD.

Cloud native & Kubernetes

K8s automates deployment, scaling, and management of containers — powerful but complex, costly, and it can create silos if adopted without DevOps thinking. Assess whether serverless or lighter orchestration is enough before committing.

Chaos engineering

Deliberate experiments (fault injection) to build confidence in resilience — Netflix's Chaos Monkey. Game days exercise the human side of incident response. It's a feedback loop for organizational learning.

MLOps

DevOps for ML: versioning big datasets and models, heavy compute, continuous learning from user input, richer feedback loops, and close collaboration between data scientists, dev, and ops.

AIOps — three waves

Wave 1: AI-generated code, tests, docs in the IDE. Wave 2: automating systems — alerting, cluster health, runbooks. Wave 3: self-service platforms and cross-team collaboration. Prompt with context to get good results.

1.7Resources

Books: The DevOps Handbook · Accelerate · The Phoenix Project. Research: dora.dev. Community: DevOpsDays, DevOps Enterprise Summit. Authors to follow: Martin Fowler, Julia Evans.