Back to DevOps Learning
$cat interview-questions.md

Interview Guide

40 real DevOps interview questions across Kubernetes, CI/CD, Terraform, deployments, security, and observability — model answers, what interviewers listen for, and diagrams for the flows you'll be asked to whiteboard.

Practice with the simulator

Kubernetes & Containers

13 questions — The deepest section, because it's where interviews go deepest.

Q1How does a request flow from a browser to a Kubernetes pod?+

Walk the layers in order: the browser resolves DNS to the address of an external load balancer (cloud LB provisioned by a Service of type LoadBalancer or fronting the ingress controller). The LB forwards to the Ingress controller (e.g. NGINX, ALB) running in the cluster, which matches host/path rules from the Ingress resource and routes to the right Service. The Service is a stable virtual IP (ClusterIP); kube-proxy has programmed iptables/IPVS rules so traffic to that VIP is DNAT'd to one of the ready pod endpoints (from EndpointSlices — only pods passing their readiness probe). Finally the CNI network delivers the packet to the pod's network namespace and the container's port.

BrowserDNSCloud LBIngresshost/path rulesServicekube-proxy DNATPod APod Bready only
DNS → LB → Ingress → Service (ClusterIP) → ready pod endpoints
Interviewers listen for
  • ✓DNS + external load balancer entry point
  • ✓Ingress controller matching host/path rules
  • ✓Service ClusterIP + kube-proxy iptables/IPVS DNAT
  • ✓EndpointSlices contain only pods passing readiness
  • ✓CNI delivers to the pod's network namespace
Q2What happens internally when you run a Docker container?+

The docker CLI sends the request to the Docker daemon (dockerd) over its API. dockerd delegates to containerd, which pulls image layers if missing, prepares a snapshot filesystem, and asks runc (an OCI runtime) to create the container. runc creates the isolation primitives: namespaces (PID, network, mount, UTS, IPC, user) so the process sees its own world, and cgroups to enforce CPU/memory limits. The image's read-only layers are stacked with a copy-on-write writable layer via a union filesystem (overlayfs); a network namespace gets a veth pair into the bridge; then the entrypoint process is exec'd as PID 1 inside the container.

docker CLIdockerdcontainerdlayers · snapshotruncnamespacespid net mnt uts ipccgroupscpu · memory limitscontainer procoverlayfs + CoW layer
CLI → dockerd → containerd → runc → namespaces + cgroups + overlayfs = running process
Interviewers listen for
  • ✓dockerd → containerd → runc chain (OCI)
  • ✓Namespaces for isolation, cgroups for resource limits
  • ✓Image layers + copy-on-write writable layer (overlayfs)
  • ✓It's a normal Linux process — no hypervisor
Q3How does Kubernetes decide which node should run a pod?+

The scheduler watches for pods with no node assigned and runs a two-phase cycle. Filtering eliminates infeasible nodes: not enough allocatable CPU/memory for the pod's requests, node selectors/affinity not satisfied, taints without matching tolerations, volume/zone constraints, port conflicts. Scoring then ranks the survivors — spreading across zones, balancing resource utilization, honoring preferred affinity and topology spread constraints — and the highest-scoring node wins. The scheduler writes a binding; the kubelet on that node sees it and starts the containers. If no node passes filtering, the pod stays Pending (and may trigger cluster autoscaler).

Pending podrequests + rules1 · Filteringresources · taints · affinity2 · Scoringspread · balance · preferBindkubelet runschosen node
Scheduler: filter infeasible nodes → score the rest → bind → kubelet starts pod
Interviewers listen for
  • ✓Filtering vs scoring phases
  • ✓Requests (not limits) drive scheduling decisions
  • ✓Taints/tolerations, node & pod affinity, topology spread
  • ✓Pending when nothing fits → cluster autoscaler
Q4Difference between Readiness, Liveness, and Startup probes.+

Readiness answers "can this pod receive traffic right now?" — failing it removes the pod from Service endpoints but does not restart it (right for temporary states: warming caches, lost DB connection). Liveness answers "is this process irrecoverably stuck?" — failing it makes the kubelet restart the container (right for deadlocks). Startup protects slow-starting apps: while it runs, liveness/readiness are suspended, so a slow boot isn't killed as "dead". Classic mistakes: pointing liveness at a dependency (DB down → restart storm across the fleet) or making liveness and readiness identical.

Startup probe windowliveness/readiness pausedLiveness → restart if stuckReadiness → traffic on/offkubelet restarts containerremoved from endpoints
Startup gates the others; liveness restarts, readiness only gates traffic
Interviewers listen for
  • ✓Readiness → endpoints only; Liveness → restart; Startup → protects slow boot
  • ✓Never point liveness at external dependencies
  • ✓Failure consequences differ — that's the whole distinction
Q5Why would a pod be Running but the application still be unavailable?+

"Running" only means containers started — it says nothing about serving. The classic causes: readiness probe failing (pod excluded from Service endpoints), Service selector mismatch (labels don't match, endpoints list is empty), wrong targetPort/containerPort, app listening on 127.0.0.1 instead of 0.0.0.0, NetworkPolicy blocking traffic, Ingress pointing at the wrong Service, DNS issues, or the app started but crashed internally while PID 1 stays alive. Debug path: kubectl get endpoints (is the pod listed?), kubectl describe pod (probe events), port-forward straight to the pod to bypass Service/Ingress and bisect the layer that's broken.

Interviewers listen for
  • ✓Running ≠ Ready — readiness gates endpoints
  • ✓Selector/label mismatch → empty endpoints
  • ✓Port and bind-address mistakes, NetworkPolicy
  • ✓Bisect with port-forward to isolate the layer
Q6A pod is stuck in CrashLoopBackOff. How do you debug it?+

CrashLoopBackOff means the container starts, exits, and the kubelet restarts it with growing backoff. First, get the reason: kubectl logs --previous for the crashed instance's output, and kubectl describe pod for events and the exit code — 137 is OOMKilled or SIGKILL (check memory limits), 1/2 are app errors, 126/127 mean bad command or missing binary. Common causes: failing config (missing env/Secret/ConfigMap), unreachable dependency at startup, a misconfigured liveness probe killing a healthy-but-slow app, or wrong entrypoint. If the container dies too fast to inspect, override the command with sleep or use kubectl debug with an ephemeral container to poke around.

Interviewers listen for
  • ✓logs --previous + describe events
  • ✓Exit codes: 137 = OOM/SIGKILL, 126/127 = command problems
  • ✓Liveness probe misconfiguration as a cause
  • ✓kubectl debug / command override to inspect
Q7What happens when you run kubectl apply -f deployment.yaml?+

kubectl sends the manifest to the API server, which authenticates, authorizes (RBAC), runs admission controllers (validating/mutating webhooks, defaults), and persists the object in etcd. Nothing "runs" it — Kubernetes is a set of control loops. The Deployment controller notices the new/changed Deployment and creates/updates a ReplicaSet; the ReplicaSet controller creates Pod objects to match replicas; the scheduler assigns each pod a node; the kubelet on that node sees its assignment and starts containers via the container runtime, reporting status back. Every layer continuously reconciles declared desired state against observed state — that's also what heals it later.

kubectlAPI serverauthn · RBAC · admissionetcdDeploy ctrl→ ReplicaSet → PodsSchedulerkubeletPod
Declarative store + independent control loops reconciling toward desired state
Interviewers listen for
  • ✓API server: authn → RBAC → admission → etcd
  • ✓Controller chain: Deployment → ReplicaSet → Pods
  • ✓Scheduler binds; kubelet runs; status flows back
  • ✓Reconciliation loops, not imperative execution
Q8Deployment vs StatefulSet vs DaemonSet — when do you use each?+

Deployment is for stateless, interchangeable replicas — pods have random names, share storage patterns, scale freely, and roll with ReplicaSets. StatefulSet is for workloads needing stable identity: ordered names (db-0, db-1), a stable DNS entry per pod via a headless Service, and a dedicated PersistentVolumeClaim per replica that survives rescheduling — databases, Kafka, anything with per-node state and ordered startup. DaemonSet runs exactly one pod per (matching) node — log shippers, metric agents, CNI plugins — and adds pods automatically as nodes join. Rule of thumb: identity/storage per replica → StatefulSet; node-level agents → DaemonSet; everything else → Deployment.

Interviewers listen for
  • ✓Deployment = stateless interchangeable replicas
  • ✓StatefulSet = stable identity, per-pod PVC, ordered operations
  • ✓DaemonSet = one per node, agents/infra
Q9How does a Kubernetes Service (ClusterIP) actually route traffic?+

A Service allocates a virtual IP that exists only in rules, not on any interface. The endpoints controller maintains EndpointSlices listing the ready pod IPs behind the selector. On every node, kube-proxy watches Services/EndpointSlices and programs iptables (or IPVS) rules: traffic to the ClusterIP:port is DNAT'd to a randomly selected pod IP:targetPort, with connection tracking keeping a flow on the same backend. Service discovery is via cluster DNS (CoreDNS): svc.namespace.svc.cluster.local resolves to the ClusterIP. NodePort adds a listener on every node forwarding into the same rules; LoadBalancer adds a cloud LB in front of NodePorts.

Interviewers listen for
  • ✓ClusterIP is virtual — implemented in iptables/IPVS
  • ✓kube-proxy programs DNAT to ready endpoints
  • ✓EndpointSlices + readiness; CoreDNS names
  • ✓ClusterIP ⊂ NodePort ⊂ LoadBalancer layering
Q10Requests vs limits — and what actually happens on OOM or CPU pressure?+

Requests are the scheduler's currency — guaranteed reservation used for placement. Limits are runtime ceilings enforced by cgroups. The behavior differs by resource: CPU is compressible — hitting the limit means throttling (latency, not death). Memory is not — exceeding the limit gets the container OOMKilled (exit 137) and restarted. Requests/limits also set QoS class: Guaranteed (requests = limits), Burstable, BestEffort — which determines eviction order under node memory pressure (BestEffort dies first). Practical guidance: always set memory requests = limits for predictability; set CPU requests honestly and be cautious with CPU limits (throttling surprises).

Interviewers listen for
  • ✓Requests → scheduling; limits → cgroup enforcement
  • ✓CPU throttles, memory OOMKills (exit 137)
  • ✓QoS classes and eviction order
Q11How does the Horizontal Pod Autoscaler work?+

The HPA controller periodically (default 15s) queries the metrics API — resource metrics from metrics-server, or custom/external metrics via adapters (Prometheus) — and computes: desiredReplicas = currentReplicas × (currentMetric / targetMetric), clamped between min and max. Crucially, resource-based HPA compares usage against requests, so wrong requests make scaling wrong. Stabilization windows and scale-down policies stop flapping. Know its limits: it can't exceed cluster capacity (that's cluster autoscaler's job adding nodes), it needs metrics that actually correlate with load, and for queue-driven work an event-based scaler like KEDA (scale on queue depth, even to zero) is often the better fit.

Interviewers listen for
  • ✓metrics-server / custom metrics; the ratio formula
  • ✓Targets are relative to requests
  • ✓Stabilization to prevent flapping; min/max bounds
  • ✓HPA scales pods; cluster autoscaler scales nodes; KEDA for queues
Q12Docker image layers and build cache — how do you optimize a Dockerfile?+

Each Dockerfile instruction creates an immutable layer; builds reuse cached layers until the first changed instruction, after which everything below rebuilds. So: order by change frequency — copy dependency manifests and install dependencies before copying source, so code changes don't bust the dependency cache. Use multi-stage builds to compile in a heavy stage and copy only artifacts into a minimal runtime image (distroless/alpine) — smaller attack surface and faster pulls. Add .dockerignore, pin base image versions (never bare latest in prod), combine related RUN steps to avoid junk layers, and run as a non-root USER.

Interviewers listen for
  • ✓Layer caching invalidates downward from first change
  • ✓Dependencies before source code
  • ✓Multi-stage builds; minimal final image
  • ✓Pinned tags, .dockerignore, non-root
Q13A node goes NotReady — what happens to its pods, and how do you investigate?+

The kubelet stops heartbeating; after the node controller's grace period, the node is marked NotReady and gets NoExecute taints. After the eviction timeout (~5 min by default), its pods are marked for deletion and controllers reschedule replacements on healthy nodes — but StatefulSet pods wait for confirmation the old one is gone (to protect single-writer storage). Investigation: kubectl describe node for conditions (MemoryPressure, DiskPressure, PIDPressure, network), then on the node itself: is the kubelet running, can it reach the API server, is the container runtime healthy, disk full? Common causes: resource exhaustion from workloads without limits, network partition, runtime crash, or cloud instance issues.

Interviewers listen for
  • ✓Heartbeats → NotReady → taint → eviction after timeout
  • ✓Stateless pods reschedule; StatefulSets are cautious
  • ✓Node conditions; kubelet/runtime/network/disk checks

CI/CD & Git

6 questions

Q14Explain the complete lifecycle of a CI/CD pipeline from commit to production.+

Structure the answer as stages, each a quality gate: commit to a short-lived branch triggers the pipeline → build + fast checks (compile, lint, unit tests, SAST/dependency scan — under ~5 min) → on merge to trunk, produce a versioned immutable artifact (container image, tagged with commit SHA) pushed to a registry → deploy to staging (an environment built from the same IaC as prod) and run integration/acceptance/smoke tests → promote the same artifact to production via a progressive strategy (rolling/canary) → post-deploy verification: deployment-correlated metrics, automated rollback triggers. Emphasize: build once, promote everywhere; every gate stops the pipeline; humans approve at most, they never rebuild.

commitbuild + testlint · unit · scanartifactimage :shastagingintegration · smokeprod canarysame artifactverifymetrics · rollback
Every stage is a gate; one artifact travels the whole path
Interviewers listen for
  • ✓Fast feedback first; pipeline stops on failure
  • ✓Immutable versioned artifact built once
  • ✓Prod-like staging from the same IaC
  • ✓Progressive rollout + post-deploy verification
Q15If a deployment fails midway in production, how do you recover safely?+

First stop the bleeding, then diagnose. Halt the rollout so it doesn't progress further (kubectl rollout pause / stop the pipeline). Assess blast radius: with a rolling update, old-version pods are usually still serving — the surge/unavailable settings limited damage. Then choose: roll back to the last known-good version (kubectl rollout undo, redeploy previous artifact, or flip blue-green traffic back) — this is the default; or roll forward with a fix only if rollback is impossible (e.g., irreversible migration). This is why you design for it in advance: versioned artifacts kept in the registry, backward-compatible expand/contract DB migrations so N-1 still works, feature flags as an instant kill-switch, and health-gated rollouts that abort automatically. Afterwards: blameless postmortem.

Interviewers listen for
  • ✓Pause first; assess blast radius
  • ✓Rollback to known-good as default; roll-forward as exception
  • ✓DB backward compatibility makes rollback possible
  • ✓Flags as kill-switch; automate abort on health checks
Q16How do you keep CI fast as the codebase and team grow?+

Treat pipeline speed as a feature with a budget (e.g. <5–10 min to merge signal). Levers: caching (dependencies, Docker layer cache, build caches like Gradle/Turborepo), parallelization (shard test suites across runners), test selection (run only tests affected by the change; full suite on a schedule), pushing coverage down the test pyramid so E2E stays thin, right-sized runners, and splitting pipelines so the merge-blocking path is minimal while slower suites (full E2E, performance, nightly security) run asynchronously. Also organizational: fix flaky tests immediately — reruns are hidden minutes — and track pipeline duration on a dashboard so regressions are visible.

Interviewers listen for
  • ✓Explicit time budget, tracked
  • ✓Caching + parallel shards + affected-only testing
  • ✓Thin blocking path, heavy suites async
  • ✓Flakes treated as build breakage
Q17What branching strategy would you choose for continuous delivery, and why?+

For CD, trunk-based development: short-lived branches (hours to ~2 days) merged into main frequently, with main always releasable. Long features ship incrementally behind feature flags. Contrast with GitFlow — develop/release/hotfix branches add merge ceremony and delay integration, which fights CI's whole purpose; it fits versioned, packaged software with parallel supported releases more than services. Release management on trunk: tag/cut releases from main, and if a fix is needed, fix on main and cherry-pick — never let release branches live long. The honest nuance interviewers like: the strategy follows the delivery model; trunk-based is the fit for continuous delivery specifically.

Interviewers listen for
  • ✓Trunk-based + short-lived branches for CD
  • ✓Flags decouple merge from release
  • ✓Why GitFlow conflicts with continuous integration
Q18How does GitOps (Argo CD / Flux) differ from push-based CD pipelines?+

In push-based CD, the pipeline has cluster credentials and imperatively applies changes at the end of a run. In GitOps, a controller inside the cluster continuously pulls desired state from a Git repo and reconciles reality against it. Consequences: Git is the audit log and rollback mechanism (revert = rollback), drift is auto-detected and corrected (a manual kubectl edit gets reverted), cluster credentials never leave the cluster, and the same model handles disaster recovery (point a fresh cluster at the repo). CI still builds/test/pushes images; "deployment" becomes a PR updating an image tag in the config repo. Trade-offs to mention: secret management needs extra tooling (sealed-secrets, external-secrets), and orchestrating multi-step releases needs progressive-delivery extensions (Argo Rollouts, Flagger).

Interviewers listen for
  • ✓Pull + continuous reconciliation vs one-shot push
  • ✓Git as source of truth: revert = rollback, full audit
  • ✓Drift auto-corrected; credentials stay in-cluster
  • ✓CI builds images; CD = config repo change
Q19Monorepo vs polyrepo — how does the choice change your pipelines?+

A monorepo gives atomic cross-service changes, one dependency graph, and easy code sharing — but naive CI rebuilds everything on every commit, so it demands path filtering / affected-graph builds (Bazel, Nx, Turborepo) and per-service deploy pipelines triggered selectively. A polyrepo gives natural isolation, simple per-repo pipelines and clear ownership — but cross-cutting changes span multiple PRs, and shared libraries need versioning/publishing discipline. Either way the goal is the same: each service builds, tests, and deploys independently. The mature answer names the tooling cost honestly: monorepos need investment in build tooling; polyrepos need investment in dependency and release coordination.

Interviewers listen for
  • ✓Monorepo → affected-only builds, selective triggers
  • ✓Polyrepo → versioned shared libs, multi-PR changes
  • ✓Independent deployability regardless of layout

Terraform & IaC

7 questions

Q20How does Terraform build and execute its dependency graph?+

Terraform parses all configuration and builds a DAG (directed acyclic graph) of resources. Edges come from implicit dependencies — any expression referencing another resource's attribute (subnet_id = aws_subnet.a.id) — plus explicit depends_on. During plan, it walks the graph comparing desired config against state (after refreshing real-world data); during apply it executes in topological order, running independent branches in parallel (default up to 10 concurrent operations). Destroys traverse the graph in reverse. Cycles are errors. This is why you should prefer attribute references over depends_on — they document the real data dependency and give Terraform accurate ordering.

VPCSubnet ASubnet BEC2 · AEC2 · BLB targetparallel branchesreverse on destroy
Implicit references form the DAG; independent branches apply in parallel
Interviewers listen for
  • ✓DAG from implicit references + depends_on
  • ✓Topological order; parallel independent branches
  • ✓Reverse order for destroy; cycles are errors
Q21What is Terraform drift, and how do you handle it?+

Drift is when real infrastructure no longer matches state/config — usually from manual console changes, emergency fixes, or external automation. Detection: terraform plan (with refresh) shows unexpected diffs; mature teams run scheduled plans in CI and alert on any non-empty diff. Handling depends on which side is right: if the manual change was wrong, apply to converge reality back to code; if it was a legitimate fix, codify it — update the configuration (and possibly terraform import resources created outside). Prevention is process: no console write access in prod, all changes through the pipeline, and drift alerts so exceptions surface in hours not months.

Interviewers listen for
  • ✓Definition: reality ≠ state/config; causes
  • ✓Scheduled plan in CI as drift detection
  • ✓Converge or codify — decide which side is truth
  • ✓Prevent via access restriction + pipeline-only changes
Q22How does Terraform state locking work with S3 and DynamoDB?+

State lives in an S3 backend (versioned, encrypted); locking uses a DynamoDB table. Before any state-modifying operation, Terraform attempts a conditional PutItem on the lock table keyed by the state path — it succeeds only if no lock item exists (an atomic compare-and-set). A second concurrent run's conditional write fails, so it errors with the holder's lock info instead of proceeding. On completion the lock item is deleted. If a run is killed mid-apply, a stale lock remains — inspect the holder, then terraform force-unlock <id> once you're sure nothing is running. S3 versioning matters here too: it's your state backup if corruption happens. (Newer Terraform also supports S3-native locking via conditional writes.)

Engineer AEngineer BDynamoDB lockconditional put (CAS)S3 stateversioned · encrypted1 · acquire ✓1 · acquire ✗ blocked2 · read/write3 · release
Atomic conditional write = mutual exclusion; the loser fails fast with lock info
Interviewers listen for
  • ✓DynamoDB conditional write as atomic lock
  • ✓Concurrent run fails fast with holder info
  • ✓force-unlock for stale locks — carefully
  • ✓S3 versioning + encryption as safety net
Q23How do you design reusable Terraform modules for enterprise projects?+

Design modules like APIs. Scope: one logical capability per module (a "VPC", an "internal service"), not a grab-bag. Interface: minimal required variables with sane, secure defaults; validation blocks; typed objects for grouped options; outputs exposing everything consumers legitimately need. Composition over inheritance: small building blocks composed into opinionated "stack" modules for common patterns. Versioning: publish via a registry or git tags with semver; consumers pin versions so module changes roll out deliberately. Quality: examples folder, README with usage, terraform test/Terratest coverage, and CI on the module repo itself. Enterprise extras: encode governance in the module (encryption on, tags mandatory, private by default) so the paved road is also the compliant road.

Interviewers listen for
  • ✓Single responsibility; API-like variable/output design
  • ✓Semver + pinned versions via registry/tags
  • ✓Tests, examples, docs as part of the module
  • ✓Secure/compliant defaults baked in
Q24How do you manage multiple environments (dev/staging/prod) in Terraform?+

Main options, with trade-offs: directory-per-environment — each env has its own root module and state, composing shared modules with env-specific variables; explicit, isolated blast radius, but some duplication (mitigated by keeping roots thin). Workspaces — same code, multiple states; light for near-identical envs, but risky when environments genuinely differ and it's easy to apply to the wrong workspace. Branch-per-env is an anti-pattern — code drifts between branches. Whatever the layout: separate state (ideally separate accounts/subscriptions) per environment, promote by applying the same module versions with different tfvars, and let the pipeline — not laptops — run prod applies with an approval on the plan.

Interviewers listen for
  • ✓Directories + shared modules vs workspaces trade-off
  • ✓Separate state and accounts per env
  • ✓Same module versions promoted with different vars
  • ✓Branch-per-environment named as an anti-pattern
Q25Terraform state is lost or corrupted — what do you do?+

First, don't panic-apply: with no state, Terraform thinks nothing exists and a plan will propose creating everything again (and duplicating real infra). Recovery, in order of preference: restore from backend versioning — S3 versioned buckets keep every prior state; roll back to the last good version. If truly gone, rebuild state by importing: the config still describes everything, so terraform import (or bulk import blocks in 1.5+) re-associates each real resource with its address, then iterate until plan is clean. For corruption, terraform state subcommands (list/rm/mv) can surgically fix entries. Then fix the root cause: versioning + encryption on the backend, locking enabled, no local state, and restricted write access to the state bucket.

Interviewers listen for
  • ✓Danger: empty state → plan recreates/duplicates everything
  • ✓Backend versioning as first-line recovery
  • ✓terraform import / import blocks to rebuild
  • ✓Prevention: versioning, locking, access control
Q26What exactly happens during terraform plan vs apply?+

Plan: load config and state → refresh (query providers for the real current attributes of tracked resources) → diff three things — config (desired), state (last known), reality (refreshed) → produce an execution plan of creates/updates/replaces/destroys, marking forces replacement where an immutable attribute changed. Plan makes no changes (aside from refresh reads) and can be saved (-out). Apply: execute exactly that plan (or re-plan if none given) in DAG order with parallelism, streaming provider API calls, updating state as each resource completes. Interview gold: explain why a resource shows "must be replaced" — an attribute the provider can't update in place — and that applying a saved plan guarantees you deploy exactly what was reviewed.

Interviewers listen for
  • ✓Refresh reconciles state with reality before diffing
  • ✓Three-way comparison: config vs state vs real world
  • ✓Replace = immutable attribute change
  • ✓Saved plan → apply-what-you-reviewed

Deployment Strategies

2 questions — Short section, but nearly guaranteed to come up.

Q27Explain the difference between Rolling, Blue-Green, and Canary deployments.+

Rolling: replace instances gradually (maxSurge/maxUnavailable in k8s) — no extra environment, zero-downtime by default, but both versions overlap during rollout and rollback means rolling back through the same process. Blue-Green: two full environments; deploy to idle (green), test it, then switch all traffic at once — instant cutover and instant rollback (switch back), at the cost of double capacity and all-users-at-once exposure. Canary: route a small percentage to the new version, compare error/latency metrics against baseline, then progressively widen or roll back — best blast-radius control and the foundation for automated, metric-gated releases, but needs traffic splitting and real observability. All three require N and N+1 to coexist against the same data — backward-compatible schema changes.

Rollingold → new, one at a timeBlue-Greenblue 100%green idleswitch all at onceCanarystable 95%5%→ watch metrics → widen or roll back
Three shapes of risk: gradual overlap · all-at-once switch · measured slice
Interviewers listen for
  • ✓Mechanics + rollback story of each
  • ✓Trade-offs: capacity cost, exposure, complexity
  • ✓Canary = metric-gated, smallest blast radius
  • ✓All need N/N+1 data compatibility
Q28How would you implement a zero-downtime deployment?+

Layer the requirements: traffic layer — a strategy that never drops all healthy backends (rolling with maxUnavailable tuned, or blue-green/canary behind a LB). Instance layer — readiness probes so new pods only receive traffic when actually ready, and graceful shutdown: handle SIGTERM, stop accepting new requests, drain in-flight ones within terminationGracePeriod (plus a preStop sleep to let endpoint removal propagate before the process exits — the classic source of "zero-downtime" 502s). Data layer — expand/contract migrations so old and new code both work mid-rollout. Client layer — retries with idempotency for the odd reset connection. Verification — deploy-correlated metrics with automated abort. Zero downtime is a property of the whole system, not a deployment flag.

Interviewers listen for
  • ✓Readiness gating + graceful SIGTERM drain (preStop)
  • ✓Connection draining / endpoint propagation race
  • ✓Expand/contract DB migrations
  • ✓Automated verification and abort

Security

4 questions

Q29How do you secure secrets across CI/CD, Kubernetes, and cloud services?+

One principle everywhere: secrets never live in code or images; they're injected at runtime from a dedicated store, scoped by identity. Centralize in a secrets manager (Vault, AWS Secrets Manager). CI/CD: pipeline authenticates via short-lived identity — ideally OIDC federation (the CI job exchanges its identity token for temporary cloud credentials; no stored long-lived keys) — secrets injected as masked env vars, never echoed. Kubernetes: External Secrets Operator or Vault injector syncs from the store; native Secrets are base64-not-encryption, so enable encryption at rest and RBAC-restrict access; workloads get cloud access via workload identity / IRSA, not node keys. Everywhere: least privilege per consumer, rotation, audit logging, and secret scanning pre-commit and in CI.

Secrets managerrotation · audit · ACLCI pipelineOIDC → short-lived credsKubernetesESO / injector · IRSACloud servicesworkload identity
One store, identity-scoped runtime injection everywhere — no static keys in code
Interviewers listen for
  • ✓Central store; runtime injection; nothing in git/images
  • ✓OIDC federation replaces long-lived CI keys
  • ✓K8s Secrets = base64 → encrypt at rest + RBAC; ESO/Vault
  • ✓Workload identity, rotation, scanning
Q30Explain IAM Roles vs IAM Policies with a real-world scenario.+

A policy is the permission document — JSON allowing/denying actions on resources under conditions. A role is an assumable identity that policies attach to: it has no long-term credentials; a trusted principal assumes it and receives temporary credentials via STS. Scenario: an app on EC2/EKS needs to read s3://invoices. Bad: create an IAM user, put its access keys in config — static credentials that leak and never rotate. Good: create role invoice-reader with a policy allowing s3:GetObject on that bucket, and a trust policy letting the EC2 instance profile (or the k8s service account via IRSA) assume it. The app gets auto-rotated temporary credentials from the metadata service; permissions and identity stay decoupled, auditable, and revocable. Rule: humans and machines assume roles; access keys are the exception.

Interviewers listen for
  • ✓Policy = permissions doc; role = assumable identity
  • ✓Trust policy vs permission policy distinction
  • ✓STS temporary credentials; no static keys
  • ✓Instance profile / IRSA scenario
Q31A cloud API key was committed to the repository. Walk me through your response.+

Treat it as compromised the moment it touched the remote — clones, CI logs, and scrapers may already have it (public repos get scanned in minutes). Order matters: 1) Revoke/rotate the credential immediately — this is the only real fix; everything else is cleanup. 2) Assess impact: audit logs (CloudTrail) for usage of that key since exposure — unusual calls, new resources, privilege escalation. 3) Clean history (git filter-repo/BFG + force push) so it stops resurfacing — while stating clearly this does NOT un-leak it. 4) Prevent recurrence: pre-commit secret scanning (gitleaks), CI scanning, push protection, and remove the need for static keys at all via OIDC/workload identity. 5) Blameless writeup — committed secrets are a tooling gap, not a character flaw.

Interviewers listen for
  • ✓Rotate first — history rewriting is not remediation
  • ✓Audit-log impact assessment
  • ✓Scanning + push protection + eliminating static keys
  • ✓Blameless framing
Q32How do you apply least privilege to a CI/CD pipeline?+

Pipelines are prime targets — they hold deploy power over everything. Measures: identity per pipeline/job, not one god-credential shared by all — each repo/environment gets its own role with only its needed permissions (this build pushes to this ECR repo, deploys to this namespace). Short-lived credentials via OIDC federation, with trust conditions pinning repo, branch, and environment (only main can assume the prod role). Environment protection: prod deploy jobs require reviews/approvals; secrets scoped to environments. Runner hygiene: ephemeral runners, no privileged Docker socket exposure, pinned third-party actions/plugins (supply chain). Plus audit logging of every assume and deploy. The theme: compromise of one pipeline should be a contained incident, not keys to the kingdom.

Interviewers listen for
  • ✓Per-pipeline identity, scoped permissions
  • ✓OIDC trust conditions on repo/branch/env
  • ✓Protected environments and approvals
  • ✓Ephemeral runners, pinned dependencies

Observability & Performance

5 questions

Q33Difference between logs, metrics, and traces — when do you use each?+

Metrics are cheap aggregated numbers over time (rates, latencies, saturation) — best for detection and alerting: they tell you that something is wrong and trends over time, but not why. Logs are discrete, high-cardinality event records — best for explanation: the exact error, stack trace, the specific request's story; costly at volume, so structure them (JSON) and index wisely. Traces follow a single request across services via propagated context, showing spans and timings per hop — best for localization in distributed systems: which service/call in the chain is slow or failing. Workflow: metric alert fires → trace localizes the failing hop → logs of that service explain it. Correlation (trace IDs in logs, exemplars linking metrics→traces) is what makes the three pillars one system.

METRICSaggregated numbersdetect: "something is wrong"TRACESrequest across serviceslocalize: "where it's wrong"LOGSdiscrete rich eventsexplain: "why it's wrong"
Detect → localize → explain, tied together by trace IDs
Interviewers listen for
  • ✓Metrics detect, traces localize, logs explain
  • ✓Cost/cardinality trade-offs
  • ✓Context propagation; correlation via trace IDs
Q34How do you identify the bottleneck when an application becomes slow in production?+

Be systematic, not random. Define the symptom: which endpoints, since when, p50 vs p99 (tail-only latency points at contention/GC/noisy neighbors; across-the-board points at a shared dependency), and what changed — deploys, config, traffic, data growth. Then follow the request path with traces: find which span eats the time — app CPU, a downstream call, or the database. Check saturation at each layer (USE method: utilization, saturation, errors — CPU throttling, memory/GC, connection pools, disk/network). Databases deserve special suspicion: slow query log, missing index, lock contention, growing table. Confirm with a targeted fix and measure again — one variable at a time. Mention profiling (flame graphs) for CPU-bound cases.

Interviewers listen for
  • ✓Scope symptom; p50 vs p99 reasoning; "what changed?"
  • ✓Traces to find the slow span
  • ✓USE method per layer; pools and throttling
  • ✓DB usual suspects; verify with measurement
Q35How do you troubleshoot intermittent 503 errors in Kubernetes?+

Intermittent 503s mean the proxy sometimes has no healthy backend or backends are refusing. First, find who emits the 503 — ingress controller, mesh sidecar, or the app — from response headers/logs. Correlate timing: do they align with deployments (the classic: pods terminated without graceful drain — no preStop delay, endpoints lag behind pod deletion), with readiness flapping (probe too aggressive, app GC pauses), with HPA scale events, or with load peaks (upstream connection pool/queue limits). Check kubectl get endpoints over time for empty/shrinking endpoint sets, ingress logs for "no route/upstream" vs upstream-returned 503s, and pod restarts/OOMKills. Fixes follow the cause: preStop sleep + graceful shutdown, tuned probes, PodDisruptionBudgets for node drains, capacity/limit tuning.

Interviewers listen for
  • ✓Identify the 503 source layer first
  • ✓Correlate with deploys/scaling — drain race is the classic
  • ✓Readiness flapping, empty endpoints, OOM restarts
  • ✓preStop/graceful shutdown, PDBs, probe tuning
Q36How do you design monitoring and alerting to reduce alert fatigue?+

Principle: page on symptoms, not causes — alert on user-facing SLO violations (error rate, latency, availability), not on every CPU spike; cause-level signals become dashboards or tickets, not pages. Use burn-rate alerts on error budgets (fast-burn pages, slow-burn tickets) instead of static thresholds. Every page must be actionable: linked runbook, clear owner, severity tiers (page vs ticket vs FYI). Reduce noise structurally: deduplicate/group related alerts, inhibit downstream alerts when the upstream cause fires, silence during maintenance. Then treat alert quality as a process: review pages regularly — every non-actionable page gets fixed, re-tiered, or deleted; track pages per on-call shift as a health metric. Noisy alerting isn't a monitoring achievement; it trains people to ignore the one page that matters.

Interviewers listen for
  • ✓Symptom/SLO-based paging; causes → dashboards
  • ✓Burn-rate alerting on error budgets
  • ✓Actionable: runbook, owner, severity tiers
  • ✓Regular alert reviews; noise is a risk
Q37How do you define SLIs, SLOs, and error budgets for a service?+

SLI = the measurement: a ratio of good events to total (successful requests / all requests; requests under 300ms / all). Choose SLIs that reflect user experience, measured as close to the user as practical. SLO = the target on that SLI over a window (99.9% success over 30 days) — set from user needs and business reality, not aspiration; 100% is wrong because it forbids all change. Error budget = 1 − SLO: the allowed unreliability (0.1% ≈ 43 min/month). It converts the speed-vs-stability argument into policy: budget healthy → ship features; budget burning → reliability work takes priority, agreed in advance. SLAs are the external, contractual echo — always looser than internal SLOs.

Interviewers listen for
  • ✓SLI = good/total ratio tied to user experience
  • ✓SLO target over a window; why not 100%
  • ✓Error budget as a decision-making policy
  • ✓SLA vs SLO distinction

Incidents & Behavioral

3 questions — Practice these out loud most; delivery matters as much as content.

Q38Explain a production incident you resolved and your RCA approach.+

Use a STAR-shaped structure with engineering depth: Situation — the impact in user/business terms first ("checkout errors for 40 minutes, ~15% of requests"). Detection — how you knew (alert, not customer complaint, ideally). Triage — stabilizing action before root cause: rollback, feature-flag off, scale up; communicating status while working. Diagnosis — the evidence trail: metrics → traces → logs, hypotheses tested, the actual cause. Resolution + RCA — blameless postmortem, and crucially multiple contributing causes (the bug AND the missing test AND the alert gap), using "5 whys" carefully without stopping at a person. Follow-through — the systemic fixes shipped and what measurably improved. Pick a real incident where you drove it; interviewers probe details, and honest depth beats polished vagueness.

Interviewers listen for
  • ✓Impact stated in user terms; detection story
  • ✓Mitigate first, root-cause second
  • ✓Evidence-driven diagnosis narrative
  • ✓Blameless, multi-cause RCA + completed actions
Q39How do you run a blameless postmortem that actually prevents recurrence?+

Blameless means treating actions as reasonable given what people knew — the question is never "who?", it's "what made this look correct, and which guardrail was missing?" Structure: timeline reconstructed from data (not memory), impact quantified, contributing causes (plural — resist single-root-cause thinking), what went well (detection? mitigation?), and action items that are specific, owned, and deadlined — "improve testing" is not an action; "add contract test X, owner Y, by date Z" is. The part most teams fail: tracking actions to completion — review open postmortem actions in a recurring forum until done, and measure the completion rate. Share postmortems widely; near-misses deserve the same review because they're free lessons.

Interviewers listen for
  • ✓Systems framing over blame; psychological safety
  • ✓Multiple contributing causes, data-based timeline
  • ✓Specific, owned, deadlined actions
  • ✓Tracking to completion; near-miss reviews
Q40Production is down, the cause is unknown, and leadership wants updates. How do you handle the first 30 minutes?+

Show incident-command thinking. Declare and structure: name an incident commander (you), a comms lead, and keep responders focused — one channel, one source of truth. Stabilize before understanding: check the obvious reversibles first — recent deploy? roll it back; recent config/flag change? revert it; dependency down? fail over or degrade gracefully. Restoration beats diagnosis; preserve evidence (logs, a cordoned pod) for later. Communicate on a cadence: short structured updates at promised intervals ("impact, current action, next update at :45") — this stops leadership from pulling responders for status. Escalate deliberately — page the owning team early rather than heroically debugging alone. And afterwards, the postmortem — mentioning it unprompted signals maturity.

Interviewers listen for
  • ✓Incident command roles; single channel
  • ✓Mitigate via recent-change rollback first
  • ✓Cadenced structured comms shielding responders
  • ✓Early escalation; postmortem follow-through