Back to DevOps Learning
$cat notes-*.html

Course Notes

Complete study notes for all three courses, refactored for clarity — every concept restructured into a logical flow, with diagrams in place of the original screenshots.

Complete study notes, refactored for clarity: every concept from the course, restructured into a logical flow, with diagrams in place of screenshots. Read top to bottom or jump by section.

01The core model

What DevOps is

DevOps combines the roles of developers and operations engineers working together across the entire service lifecycle — from design through development to production support. The goal is to pull specialized roles (front-end devs, test engineers, DBAs, sysadmins) out of isolation and into collaboration. The measurable payoff: faster deployments, fewer failures, quicker recovery, and less burnout.

Three levels of understanding

DevOps is understood at three levels that build on each other — this framing organizes the whole course:

VALUES what we believe CAMS PRINCIPLES beliefs formalized into a plan The Three Ways PRACTICES how we act on them 5 practice areas
Values → principles → practices: each level makes the previous one actionable

Values — CAMS

  • Culture — DevOps is fundamentally about changing people's behavior and building collaboration between Dev and Ops. Everything else stands on this.
  • Automation — remove inefficiency and manual error to improve quality — but only on top of a solid cultural foundation, or you automate the dysfunction.
  • Measurement — track key metrics to understand and improve both technical and business outcomes.
  • Sharing — open information flow and collaboration across teams drives continuous improvement.

Principles — The Three Ways

Dev Ops Customer 1 · FLOW — optimize the whole system, left to right 2 · FEEDBACK — amplified loops, right to left 3 · LEARNING — continual experimentation and risk-taking, everywhere
Systems thinking & flow · shortened, amplified feedback · a culture of experimentation and learning from failure

Practices — the five-area playbook

  • Culture · Process (Agile/Lean) · Infrastructure as Code · Continuous Delivery · Site Reliability Engineering.
  • These areas are interdependent — advance them together, iteratively. Unbalanced adoption (all tooling, no culture) breeds frustration and stalls.

Choosing tools

People over process over tools. Identify the right people and processes first, then pick tools that fit. Keep the toolchain small, well-integrated, and adaptable — too many tools creates its own complexity.

02Culture — the hardest, highest-leverage practice

Why culture change is needed: the wall of confusion

IT departments suffer misalignment and conflict — within IT and with the business. The classic picture is the wall of confusion: each team finishes its part and throws the work "over the wall" to the next group.

Developmentrewarded for CHANGE THE WALL Operationsrewarded for STABILITY "works on my machine" "your code broke prod"
The wall isn't a personality problem — it's built by opposing institutional incentives
The real cause of the division is institutional incentives that reward opposing behaviors: Dev is measured on shipping change, Ops on preventing it. Fix the incentives (shared goals and metrics), and the wall comes down. Tactically: embed operations engineers inside development teams to build cross-functional ownership and mutual respect.

Communication and trust

  • Establish clear, agreed channels for how information flows — ambiguity about "who needed to know" is where trust dies.
  • Good communication builds trust, and trust is what makes information actually flow — the engine of organizational performance.

Continuous learning: kaizen

Kaizen = continuous improvement: small, ongoing changes made by everyone rather than big-bang transformations. Its five principles:

  • Know your customer — improvement is defined by the value it delivers.
  • Enable smooth workflow — remove friction and irregularity.
  • Go to the real place (gemba) — observe where the work actually happens, don't theorize from afar.
  • Empower people — those doing the work improve the work.
  • Be transparent — visible work, visible metrics, visible problems.

Learning also means mastering core skills and taking deliberate risks to learn practically — experimentation is a feature, not indiscipline.

03Process — the building blocks

Agile

  • DevOps grew out of Agile-infrastructure discussions at conferences in 2008–2009 — it extends Agile past "working software" into deployment and operations.
  • Agile vs Waterfall: iterative loops with active collaboration instead of a linear handoff sequence.
  • Reported benefits: higher productivity, faster time to market, more predictable delivery, better quality.

Lean

Lean (adapted from manufacturing) exists so that the value stream reaches the customer while waste is eliminated. Three wastes to hunt:

MUDAunnecessary activitywork that adds no value MURAirregular flowbursts, queues, waiting MURIoverburdeningpeople or systems past capacity
The three Lean wastes — in IT, the biggest is usually mura: finished work waiting in queues
  • Kaizen — small continuous improvements (see Culture).
  • Value stream mapping — visualize every step from idea to customer; wait time between steps is where cycle time hides.
  • Visual management — kanban boards and dashboards make work and problems visible.

Lightweight change control

  • Change control should be lightweight, fast, scalable, and repeatable — not a biweekly approval board.
  • The most effective reviews are peer reviews: performed quickly, in a distributed fashion, by a technologist close to the system being changed. They're optimal for the vast majority of changes that aren't explicitly risky or cross-technology.
  • Keep changes small — easier to review, easier to fix when something goes wrong.
  • CI with automated testing validates every change early, so review effort goes where judgment is needed.

Where ITIL fits

ITIL is a comprehensive IT service management framework (service strategy, design, transition, operation, continual improvement — 34 process areas including incident, problem, and change management). It can coexist with DevOps:

  • Adapt ITIL change management into lightweight change control rather than importing full ceremony.
  • Its service-management focus aligns with CD's goal of sustained service quality.
  • Its continual-improvement process complements DevOps feedback loops.
The balance to strike: ITIL brings stability discipline, DevOps brings agility — take the ideas, not the 34-process weight.

04Infrastructure as code (practice area)

Provision and manage infrastructure through automation code instead of manual processes. Modern systems are programmable — configured entirely through software — which lets infrastructure follow a development lifecycle: versioned, reviewed, tested, deployed. This is the automation value applied to operations, and it demands a mindset shift from manual setup to software engineering. (The full deep-dive is course 2 — see the IaC notes.)

Configuration management — the three parts

PROVISIONINGmake infrastructure exist:hardware · OS · network DEPLOYMENTinstall & upgrade theapplication software ORCHESTRATIONcoordinate across systems:rolling updates · failover
Configuration management = creating and maintaining systems in a desired state, as code

Key terms

TermMeaning
Imperative (procedural)Define the exact commands/steps to reach the state — e.g. stop service, update, restart.
Declarative (functional)Define the desired end state; the tool figures out how to get there.
IdempotentRun the procedure N times, end in the same state every time — consistency and safety.
Self-serviceEnd users trigger processes without waiting on another team — velocity and satisfaction.
DriftThe running system deviating from desired configuration (manual changes, script errors); tools include drift detection.

Tool evolution & toolchain choices

  • Evolution: golden-image tools (Ghost) → infrastructure-as-code CM (CFEngine, Puppet, Chef) → orchestrated deployment (Ansible, SaltStack) → self-service orchestration and runbooks (Rundeck) → cloud provisioning (CloudFormation → Terraform, Pulumi) → containers and Kubernetes → immutable infrastructure (replace servers, never modify them).
  • Choosing: design the operational environment first, then pick a cohesive toolchain. Decide: manage your own infra or use a service; template-driven provisioning (CloudFormation, ARM) vs programmable (Boto, CDK); config-management deployment (Chef, Puppet) vs container-based (Docker, Kubernetes). Start simple, add complexity only as needed, and match tools to the team's skills.

Example: what IaC looks like

# Terraform — declare an EC2 instance
provider "aws" { region = "us-west-2" }

resource "aws_instance" "example" {
  ami           = "ami-0c55b159cbfafe1f0"
  instance_type = "t2.micro"
  tags = { Name = "example-instance" }
}
# Ansible — converge web servers to a state
- name: Configure web server
  hosts: webservers
  tasks:
    - name: Install Nginx
      apt: { name: nginx, state: present }
    - name: Start Nginx service
      service: { name: nginx, state: started }

The same idea covers CloudFormation templates, Kubernetes YAML (declare a Pod, the platform converges to it), Chef recipes for databases, and network device playbooks — if it can be described, it can be code.

Testing infrastructure code

  • Unit-ish validation — policy/compliance checks on the code itself (e.g. terraform-compliance: are instances tagged? security groups correct?).
  • Integration tests — Terratest deploys real resources (a VPC) and verifies configuration (subnets, route tables).
  • End-to-end — deploy a full environment and test component interaction.
  • Continuous — run these in the CI pipeline on every infra change; mock cloud services (LocalStack) for fast, free early tests.
  • Observe — monitoring (Prometheus/Grafana) closes the loop on whether infra behaves in reality.

05Continuous delivery (practice area)

The three definitions

CONTINUOUS INTEGRATIONbuild + unit test on every check-in CONTINUOUS DELIVERYevery build → prod-like env + tests CONTINUOUS DEPLOYMENTauto-release to production always in a working state always READY to deploy Amazon · Google · Etsy · Facebook each practice contains and extends the previous one
CI validates code · CD validates the system · Deployment delivers to users

Benefits: dramatically shorter deployment time, rapid experimentation, and better quality because testing happens earlier in the process.

Six practices for continuous integration

1 · Builds run fastunder 5 minutes 2 · Small commitseasier to manage & debug 3 · Never leave it brokenfixing red is priority #1 4 · Trunk-based devno long-running branches 5 · Fix flaky testsunreliable tests kill confidence 6 · Build artifactsstatus + log + installable output
Six CI practices — together they create the fast feedback loop that changes delivery culture

Five practices for continuous delivery

Immutableartifactsbuild once, useeverywhere unchanged Prod-likepre-prodstaging mirrorsproduction closely Automatetestingevery stepyou possibly can Stop onfailurepipeline halts,fix immediately Idempotentdeployssame deploy,same result, always
Five CD practices — the artifact you tested is exactly what ships, every time

The role of QA: automated testing

  • Automated testing is what makes CI/CD possible — minimize manual testing.
  • Types: unit testing, code hygiene (lint/format), integration testing, acceptance testing — each with its own role.
  • TDD: write a failing test first, then the code — tests guide development. BDD extends it by describing behavior from the end user's perspective in natural language, improving communication among stakeholders. TDD focuses on components; BDD on collaboration and user requirements.
  • Manage slow tests: run them in parallel, schedule them, or run them continuously against test environments.
  • Specialty testing rounds it out: infrastructure, performance, and security tests.

Continuous deployment & production release patterns

Continuous deployment auto-releases every change that passes the full pipeline. Organizations not ready for that use manual approvals (a product manager signs off) or batch changes. Either way, releases reach production through safety patterns:

ROLLING one instance at a time; traffic shifts seamlessly between versions BLUE-GREEN BLUE (live) GREEN (new) switch all traffic once verified CANARY current · 95% 5% small subset first → monitor → widen to the whole system A/B + FEATURE FLAGS feature OFF feature ON (subset) release features to user segments at runtime — deploy and release become separate decisions
Four production release patterns — pick per service, and design them collaboratively with packaging + IaC + architecture owners
Real-world example: Signal Sciences (security startup) combined rolling deploys with A/B feature flags. Their internal tool "Deployer" let any engineer push the latest build to staging and — if automated tests passed — to production with one button.

The CI toolchain — layers of an onion

  • Version control (innermost): Git-based — GitHub, GitLab, Bitbucket.
  • Build system: Jenkins, GitHub Actions — builds and initiates pipeline stages.
  • Testing tools: unit, hygiene, integration, acceptance — Pytest, Selenium, JMeter.
  • Artifact repository: Artifactory, Nexus.
  • Deployment (outermost): rolling/A-B tooling — Kubernetes, Ansible.

06Site reliability engineering

SRE = software engineering applied to IT operations (born at Google): build for reliability from the start, then use operational feedback to keep improving. Target outcomes: better change failure rate, faster time-to-restore, and met uptime/performance goals — via reliability testing and operational automation. SREs should spend at least half their time building tools (runbooks, monitoring) rather than manually firefighting.

Building for reliability — theory

  • Design decisions determine production behavior — architecture and reliability planning at design time beat heroics later.
  • Integration points are the #1 source of failure, and the classic disaster is the cascading failure: one component's problem spreads system-wide.
Service A Service Bfailing CLOSED — calls flowfailures counted against threshold OPEN — calls fail fastB gets time to recover HALF-OPEN — probe callssuccess → closed · failure → open
The circuit breaker pattern: after a failure threshold trips, calls fail immediately instead of piling onto a failing service — combined with timeouts, it stops cascades
  • Key resources: Release It! by Michael Nygard (production failure patterns), the Twelve-Factor App (e.g. config separated from code for portability), and Martin Fowler's writing on architecture.

Building for reliability — practice

  • All systems fail — design assuming it. Failure and slowdown are normal in complex systems.
  • Resilience = maintaining or regaining stability after disruption: redundancy, load balancers, auto-scaling, failover.
  • Systems are sociotechnical — technology + people. Humans cause and resolve issues; they're part of the resilience design, not outside it.

Operational feedback — observability

Observability is understanding internal system states from external outputs. Five areas to measure:

1 · Synthetic checkssimulated users: "is it working?" 2 · System & app metricsCPU · memory · custom app metrics 3 · End-user performanceRUM + APM: the real experience 4 · Logswhat · when · where — forensics 5 · Security monitoringthreats from logs & metrics
The five observability areas: detection, troubleshooting, and feedback that improves development

Operational feedback — incident response & postmortems

Incident response = 3 activities

  • Troubleshooting — knowing the system well enough to diagnose and fix.
  • Automation — pre-built tools/runbooks for fast, safe information-gathering and remediation.
  • Communication — coordinating specialists; keeping stakeholders and users informed.

The process is modeled on the real-world Incident Command System: defined detection, reporting, and management so response doesn't add chaos.

Postmortem principles

  • No single root cause — incidents result from multiple deficiencies together.
  • Blameless — understand the system and the actions that made sense at the time; fix the system, not the person.
  • Transparent — open communication during and after incidents builds trust and improves process.

The SRE toolchain

  • Building for reliability: language-level resilience libraries (e.g. Resilience4j for Java); Dev + Ops collaborating at design time on the approach.
  • Observability tools: SaaS (Datadog, Honeycomb, SumoLogic), open source (Prometheus, Grafana, Nagios), commercial (Splunk, Solarwinds). Start with a minimum viable monitoring stack — synthetics, system monitoring, latency from logs — then build-measure-learn your way to what your service actually needs.
  • Incident tooling: PagerDuty, OpsGenie, VictorOps for response; Rundeck, Ansible Tower, StackStorm for runbook automation.
  • Transparency: status pages (Atlassian Statuspage, Status.io) to communicate outages honestly.

07Advanced topics

Platform engineering — the paved road

1 · Blaze a trailone team proves the path works 2 · Pave the roadstandardize: golden CI/CD, observability 3 · Build infrastructureadvanced platform, from real needs
The lean platform path — never overbuild upfront; each stage is justified by demand from the last
  • Netflix, Meta, Google, and Spotify scaled DevOps by investing in self-service automation — the golden path makes the right way the easy way.
  • Platform engineering = designing toolchains and workflows that enable self-service, optimizing the global system rather than a central team's convenience.
  • Treat the platform as a product: uncover user requirements, ensure quality, and market it internally to its users.

DevSecOps

  • Security integrated at every stage — the CAMS values with a security lens: security works alongside developers (Culture); security tools in the IDE and CI, shifting left (Automation); joint security metrics (Measurement); shared responsibility with InfoSec (Sharing).
  • Security champions: under-resourced security teams deputize a champion inside each squad to spread practice and culture.
  • No FUD: inspire collaboration rather than pushing compliance through fear, uncertainty, and doubt.

Cloud native & Kubernetes

  • Kubernetes automates container deployment, scaling, and management, abstracting compute/network/storage — with service discovery, health monitoring, and observability built into the platform so app teams don't rebuild them.
  • Costs: high configurability → complexity and a steep learning curve; platforms are many stitched-together tools (hard upgrades); installations are expensive and need a dedicated team.
  • Assess honestly: serverless or lighter orchestration may deliver enough benefit with far less cost. And adopted without DevOps thinking, Kubernetes can create new silos — keep CAMS and systems thinking in view. The CNCF ecosystem is huge; popularity isn't a requirements analysis.

Chaos engineering

  • Experiment on the system deliberately to build confidence it withstands real-world conditions — Netflix's Chaos Monkey randomly breaks production components so weaknesses surface before real failures do.
  • Fault injection validates system behavior; game days emulate significant faults to exercise the human side of incident response — because systems are sociotechnical.
  • It's a feedback loop for organizational learning, and applies to complex platforms like Kubernetes.

MLOps

Data + modelsversioned large datasets TrainingHPC clusters · batch jobs Servingpredictions in production Monitoringquality · governance continuous learning from user input feeds back into data
The MLOps loop — a richer feedback cycle than traditional DevOps
  • MLOps = DevOps for machine learning: automated versioning of large datasets and models, high-performance compute for intensive training, and ongoing monitoring and governance because AI systems keep learning from user input.
  • New collaborators: data scientists generate demanding workloads and need close partnership with dev and ops; vector databases manage model data.

AIOps — three waves of AI in DevOps

WAVE 1 · Generating codeCopilot & co in the IDE: code,reviews, tests, docs WAVE 2 · Automating systemsalerting · cluster health · runbookse.g. k8sgpt explaining cluster state WAVE 3 · Collaborationself-service platforms; teams'work explained across silos
Each wave moves AI deeper into the operational core
  • LLMs automate repetitive work and reduce cognitive load: writing/refactoring code, converting between languages, documentation, API interaction, data transformation, better alert detection, remediation suggestions, and risk-flagging in code review.
  • Prompt engineering matters: give the model context about what you're trying to achieve to get useful results.

08Career & resources

  • DevOps is a mindset and a suite of practices, not a job title — and it improves every role: developers (reliable apps, builds, observability), sysadmins (IaC, reliability design, runbook automation), platform engineers (automation around outcomes), release managers (pipelines and rollout), SREs (runtime support), security engineers (DevSecOps) — and even sales, marketing, product, and executives benefit from understanding it.
  • Learning path: pick a target role (SRE, automation developer, release engineer…), build the basics — operating systems, a programming language, cloud, containers — then go deep with courses on Lean/Agile, IaC, CI/CD, SRE, DevSecOps, observability, and incident management. Experience is key — hands-on practice over theory.
  • Books: The DevOps Handbook · Accelerate · The Phoenix Project. Research: dora.dev. Events: DevOpsDays, DevOps Enterprise Summit. Follow: Martin Fowler, Julia Evans. Certs: cloud-specific (AWS, HashiCorp, Kubernetes) and DevOps Institute.