Course Notes
Complete study notes for all three courses, refactored for clarity — every concept restructured into a logical flow, with diagrams in place of the original screenshots.
Complete study notes, refactored for clarity: every concept from the course, restructured into a logical flow, with diagrams in place of screenshots. Read top to bottom or jump by section.
01The core model
What DevOps is
DevOps combines the roles of developers and operations engineers working together across the entire service lifecycle — from design through development to production support. The goal is to pull specialized roles (front-end devs, test engineers, DBAs, sysadmins) out of isolation and into collaboration. The measurable payoff: faster deployments, fewer failures, quicker recovery, and less burnout.
Three levels of understanding
DevOps is understood at three levels that build on each other — this framing organizes the whole course:
Values — CAMS
- Culture — DevOps is fundamentally about changing people's behavior and building collaboration between Dev and Ops. Everything else stands on this.
- Automation — remove inefficiency and manual error to improve quality — but only on top of a solid cultural foundation, or you automate the dysfunction.
- Measurement — track key metrics to understand and improve both technical and business outcomes.
- Sharing — open information flow and collaboration across teams drives continuous improvement.
Principles — The Three Ways
Practices — the five-area playbook
- Culture · Process (Agile/Lean) · Infrastructure as Code · Continuous Delivery · Site Reliability Engineering.
- These areas are interdependent — advance them together, iteratively. Unbalanced adoption (all tooling, no culture) breeds frustration and stalls.
Choosing tools
02Culture — the hardest, highest-leverage practice
Why culture change is needed: the wall of confusion
IT departments suffer misalignment and conflict — within IT and with the business. The classic picture is the wall of confusion: each team finishes its part and throws the work "over the wall" to the next group.
Communication and trust
- Establish clear, agreed channels for how information flows — ambiguity about "who needed to know" is where trust dies.
- Good communication builds trust, and trust is what makes information actually flow — the engine of organizational performance.
Continuous learning: kaizen
Kaizen = continuous improvement: small, ongoing changes made by everyone rather than big-bang transformations. Its five principles:
- Know your customer — improvement is defined by the value it delivers.
- Enable smooth workflow — remove friction and irregularity.
- Go to the real place (gemba) — observe where the work actually happens, don't theorize from afar.
- Empower people — those doing the work improve the work.
- Be transparent — visible work, visible metrics, visible problems.
Learning also means mastering core skills and taking deliberate risks to learn practically — experimentation is a feature, not indiscipline.
03Process — the building blocks
Agile
- DevOps grew out of Agile-infrastructure discussions at conferences in 2008–2009 — it extends Agile past "working software" into deployment and operations.
- Agile vs Waterfall: iterative loops with active collaboration instead of a linear handoff sequence.
- Reported benefits: higher productivity, faster time to market, more predictable delivery, better quality.
Lean
Lean (adapted from manufacturing) exists so that the value stream reaches the customer while waste is eliminated. Three wastes to hunt:
- Kaizen — small continuous improvements (see Culture).
- Value stream mapping — visualize every step from idea to customer; wait time between steps is where cycle time hides.
- Visual management — kanban boards and dashboards make work and problems visible.
Lightweight change control
- Change control should be lightweight, fast, scalable, and repeatable — not a biweekly approval board.
- The most effective reviews are peer reviews: performed quickly, in a distributed fashion, by a technologist close to the system being changed. They're optimal for the vast majority of changes that aren't explicitly risky or cross-technology.
- Keep changes small — easier to review, easier to fix when something goes wrong.
- CI with automated testing validates every change early, so review effort goes where judgment is needed.
Where ITIL fits
ITIL is a comprehensive IT service management framework (service strategy, design, transition, operation, continual improvement — 34 process areas including incident, problem, and change management). It can coexist with DevOps:
- Adapt ITIL change management into lightweight change control rather than importing full ceremony.
- Its service-management focus aligns with CD's goal of sustained service quality.
- Its continual-improvement process complements DevOps feedback loops.
04Infrastructure as code (practice area)
Provision and manage infrastructure through automation code instead of manual processes. Modern systems are programmable — configured entirely through software — which lets infrastructure follow a development lifecycle: versioned, reviewed, tested, deployed. This is the automation value applied to operations, and it demands a mindset shift from manual setup to software engineering. (The full deep-dive is course 2 — see the IaC notes.)
Configuration management — the three parts
Key terms
| Term | Meaning |
|---|---|
| Imperative (procedural) | Define the exact commands/steps to reach the state — e.g. stop service, update, restart. |
| Declarative (functional) | Define the desired end state; the tool figures out how to get there. |
| Idempotent | Run the procedure N times, end in the same state every time — consistency and safety. |
| Self-service | End users trigger processes without waiting on another team — velocity and satisfaction. |
| Drift | The running system deviating from desired configuration (manual changes, script errors); tools include drift detection. |
Tool evolution & toolchain choices
- Evolution: golden-image tools (Ghost) → infrastructure-as-code CM (CFEngine, Puppet, Chef) → orchestrated deployment (Ansible, SaltStack) → self-service orchestration and runbooks (Rundeck) → cloud provisioning (CloudFormation → Terraform, Pulumi) → containers and Kubernetes → immutable infrastructure (replace servers, never modify them).
- Choosing: design the operational environment first, then pick a cohesive toolchain. Decide: manage your own infra or use a service; template-driven provisioning (CloudFormation, ARM) vs programmable (Boto, CDK); config-management deployment (Chef, Puppet) vs container-based (Docker, Kubernetes). Start simple, add complexity only as needed, and match tools to the team's skills.
Example: what IaC looks like
# Terraform — declare an EC2 instance
provider "aws" { region = "us-west-2" }
resource "aws_instance" "example" {
ami = "ami-0c55b159cbfafe1f0"
instance_type = "t2.micro"
tags = { Name = "example-instance" }
}
# Ansible — converge web servers to a state
- name: Configure web server
hosts: webservers
tasks:
- name: Install Nginx
apt: { name: nginx, state: present }
- name: Start Nginx service
service: { name: nginx, state: started }
The same idea covers CloudFormation templates, Kubernetes YAML (declare a Pod, the platform converges to it), Chef recipes for databases, and network device playbooks — if it can be described, it can be code.
Testing infrastructure code
- Unit-ish validation — policy/compliance checks on the code itself (e.g. terraform-compliance: are instances tagged? security groups correct?).
- Integration tests — Terratest deploys real resources (a VPC) and verifies configuration (subnets, route tables).
- End-to-end — deploy a full environment and test component interaction.
- Continuous — run these in the CI pipeline on every infra change; mock cloud services (LocalStack) for fast, free early tests.
- Observe — monitoring (Prometheus/Grafana) closes the loop on whether infra behaves in reality.
05Continuous delivery (practice area)
The three definitions
Benefits: dramatically shorter deployment time, rapid experimentation, and better quality because testing happens earlier in the process.
Six practices for continuous integration
Five practices for continuous delivery
The role of QA: automated testing
- Automated testing is what makes CI/CD possible — minimize manual testing.
- Types: unit testing, code hygiene (lint/format), integration testing, acceptance testing — each with its own role.
- TDD: write a failing test first, then the code — tests guide development. BDD extends it by describing behavior from the end user's perspective in natural language, improving communication among stakeholders. TDD focuses on components; BDD on collaboration and user requirements.
- Manage slow tests: run them in parallel, schedule them, or run them continuously against test environments.
- Specialty testing rounds it out: infrastructure, performance, and security tests.
Continuous deployment & production release patterns
Continuous deployment auto-releases every change that passes the full pipeline. Organizations not ready for that use manual approvals (a product manager signs off) or batch changes. Either way, releases reach production through safety patterns:
The CI toolchain — layers of an onion
- Version control (innermost): Git-based — GitHub, GitLab, Bitbucket.
- Build system: Jenkins, GitHub Actions — builds and initiates pipeline stages.
- Testing tools: unit, hygiene, integration, acceptance — Pytest, Selenium, JMeter.
- Artifact repository: Artifactory, Nexus.
- Deployment (outermost): rolling/A-B tooling — Kubernetes, Ansible.
06Site reliability engineering
SRE = software engineering applied to IT operations (born at Google): build for reliability from the start, then use operational feedback to keep improving. Target outcomes: better change failure rate, faster time-to-restore, and met uptime/performance goals — via reliability testing and operational automation. SREs should spend at least half their time building tools (runbooks, monitoring) rather than manually firefighting.
Building for reliability — theory
- Design decisions determine production behavior — architecture and reliability planning at design time beat heroics later.
- Integration points are the #1 source of failure, and the classic disaster is the cascading failure: one component's problem spreads system-wide.
- Key resources: Release It! by Michael Nygard (production failure patterns), the Twelve-Factor App (e.g. config separated from code for portability), and Martin Fowler's writing on architecture.
Building for reliability — practice
- All systems fail — design assuming it. Failure and slowdown are normal in complex systems.
- Resilience = maintaining or regaining stability after disruption: redundancy, load balancers, auto-scaling, failover.
- Systems are sociotechnical — technology + people. Humans cause and resolve issues; they're part of the resilience design, not outside it.
Operational feedback — observability
Observability is understanding internal system states from external outputs. Five areas to measure:
Operational feedback — incident response & postmortems
Incident response = 3 activities
- Troubleshooting — knowing the system well enough to diagnose and fix.
- Automation — pre-built tools/runbooks for fast, safe information-gathering and remediation.
- Communication — coordinating specialists; keeping stakeholders and users informed.
The process is modeled on the real-world Incident Command System: defined detection, reporting, and management so response doesn't add chaos.
Postmortem principles
- No single root cause — incidents result from multiple deficiencies together.
- Blameless — understand the system and the actions that made sense at the time; fix the system, not the person.
- Transparent — open communication during and after incidents builds trust and improves process.
The SRE toolchain
- Building for reliability: language-level resilience libraries (e.g. Resilience4j for Java); Dev + Ops collaborating at design time on the approach.
- Observability tools: SaaS (Datadog, Honeycomb, SumoLogic), open source (Prometheus, Grafana, Nagios), commercial (Splunk, Solarwinds). Start with a minimum viable monitoring stack — synthetics, system monitoring, latency from logs — then build-measure-learn your way to what your service actually needs.
- Incident tooling: PagerDuty, OpsGenie, VictorOps for response; Rundeck, Ansible Tower, StackStorm for runbook automation.
- Transparency: status pages (Atlassian Statuspage, Status.io) to communicate outages honestly.
07Advanced topics
Platform engineering — the paved road
- Netflix, Meta, Google, and Spotify scaled DevOps by investing in self-service automation — the golden path makes the right way the easy way.
- Platform engineering = designing toolchains and workflows that enable self-service, optimizing the global system rather than a central team's convenience.
- Treat the platform as a product: uncover user requirements, ensure quality, and market it internally to its users.
DevSecOps
- Security integrated at every stage — the CAMS values with a security lens: security works alongside developers (Culture); security tools in the IDE and CI, shifting left (Automation); joint security metrics (Measurement); shared responsibility with InfoSec (Sharing).
- Security champions: under-resourced security teams deputize a champion inside each squad to spread practice and culture.
- No FUD: inspire collaboration rather than pushing compliance through fear, uncertainty, and doubt.
Cloud native & Kubernetes
- Kubernetes automates container deployment, scaling, and management, abstracting compute/network/storage — with service discovery, health monitoring, and observability built into the platform so app teams don't rebuild them.
- Costs: high configurability → complexity and a steep learning curve; platforms are many stitched-together tools (hard upgrades); installations are expensive and need a dedicated team.
- Assess honestly: serverless or lighter orchestration may deliver enough benefit with far less cost. And adopted without DevOps thinking, Kubernetes can create new silos — keep CAMS and systems thinking in view. The CNCF ecosystem is huge; popularity isn't a requirements analysis.
Chaos engineering
- Experiment on the system deliberately to build confidence it withstands real-world conditions — Netflix's Chaos Monkey randomly breaks production components so weaknesses surface before real failures do.
- Fault injection validates system behavior; game days emulate significant faults to exercise the human side of incident response — because systems are sociotechnical.
- It's a feedback loop for organizational learning, and applies to complex platforms like Kubernetes.
MLOps
- MLOps = DevOps for machine learning: automated versioning of large datasets and models, high-performance compute for intensive training, and ongoing monitoring and governance because AI systems keep learning from user input.
- New collaborators: data scientists generate demanding workloads and need close partnership with dev and ops; vector databases manage model data.
AIOps — three waves of AI in DevOps
- LLMs automate repetitive work and reduce cognitive load: writing/refactoring code, converting between languages, documentation, API interaction, data transformation, better alert detection, remediation suggestions, and risk-flagging in code review.
- Prompt engineering matters: give the model context about what you're trying to achieve to get useful results.
08Career & resources
- DevOps is a mindset and a suite of practices, not a job title — and it improves every role: developers (reliable apps, builds, observability), sysadmins (IaC, reliability design, runbook automation), platform engineers (automation around outcomes), release managers (pipelines and rollout), SREs (runtime support), security engineers (DevSecOps) — and even sales, marketing, product, and executives benefit from understanding it.
- Learning path: pick a target role (SRE, automation developer, release engineer…), build the basics — operating systems, a programming language, cloud, containers — then go deep with courses on Lean/Agile, IaC, CI/CD, SRE, DevSecOps, observability, and incident management. Experience is key — hands-on practice over theory.
- Books: The DevOps Handbook · Accelerate · The Phoenix Project. Research: dora.dev. Events: DevOpsDays, DevOps Enterprise Summit. Follow: Martin Fowler, Julia Evans. Certs: cloud-specific (AWS, HashiCorp, Kubernetes) and DevOps Institute.