Welcome to the Containers & Kubernetes master collection containing 250 comprehensive interview questions and detailed answers covering Docker, containerd, CRI-O, Kubernetes Core & Control Plane Architecture, Advanced Networking (Gateway API, Cilium/eBPF), Autoscaling (Karpenter, KEDA), Helm, Pod Security Standards, and in-depth Production Troubleshooting.
Answer: A container is an isolated Linux process running on a shared host OS kernel, utilizing Linux Namespaces (for isolation) and cgroups (for resource limitation). A Virtual Machine runs a complete guest operating system on top of a hypervisor layer (Type 1 or Type 2), requiring dedicated virtual hardware, gigabytes of RAM, and minutes to boot. Containers share the host kernel, start in milliseconds, and consume minimal system resources.
Answer:
- PID (Process ID): Isolates the process tree (container sees its main process as PID 1).
- NET (Networking): Isolates network devices, IP routing tables, port numbers, and firewall rules.
- MNT (Mount): Isolates filesystem mount points (container sees its own root filesystem).
- IPC (Inter-Process Communication): Isolates POSIX message queues and shared memory segments.
- UTS (UNIX Timesharing System): Isolates hostname and domain name.
- USER (User IDs): Maps UID/GID inside the container to different UIDs on the host (e.g., container root UID 0 maps to unprivileged host UID 10001).
- CGROUP (Control Group): Isolates cgroup root hierarchy visibility.
3. What are Linux cgroups (Control Groups) and what is the difference between cgroups v1 and cgroups v2?
Answer: cgroups enforce resource allocation and limits (CPU, memory, disk I/O, PIDs) for process groups.
- cgroups v1: Had separate, uncoordinated resource hierarchies for each controller (cpu, memory, blkio). Memory controller could not track buffered writeback I/O.
- cgroups v2 (Modern Standard): Unified single-hierarchy architecture, unified page cache and memory pressure tracking, robust out-of-memory (OOM) handling, and rootless container resource delegation.
Answer:
- High-Level Runtimes (CRI Runtimes): Manage container lifecycle, pull images from registries, unpack filesystems, and configure networks (e.g., containerd, CRI-O).
- Low-Level Runtimes (OCI Runtimes): Interact directly with the Linux kernel to create namespaces, set up cgroups, and spawn the container process (e.g., runc, crun, youki).
Answer: An open governance project under the Linux Foundation that establishes standardized specifications:
- Image Spec: Defines the archive format, manifest JSON, and layer serialization for container images.
- Runtime Spec: Defines the configuration (
config.json) and lifecycle operations (create,start,kill,delete) for running containers. - Distribution Spec: Defines the standard HTTP API for pushing and pulling images from OCI registries.
Answer: The reference implementation of the OCI runtime specification. It is a lightweight CLI wrapper written in Go that configures Linux kernel namespaces and cgroups to execute a container process and then exits.
Answer: containerd is an industry-standard core container runtime managing image transfer, storage, and execution. The containerd-shim is a lightweight daemon that sits between containerd and runc. It stays alive to maintain open standard I/O (stdin/stdout/stderr) streams and exit codes without keeping the main containerd daemon attached, enabling daemon restarts with zero container downtime.
Answer: A lightweight, purpose-built Kubernetes Container Runtime Interface (CRI) runtime developed specifically for Kubernetes, running OCI-compliant containers directly without any Docker or third-party overhead.
Answer:
- gVisor (
runsc): An application kernel written in Go that runs in userspace, intercepting and implementing Linux system calls to provide a strong security boundary between the container and host kernel. - Kata Containers: Runs each container inside its own lightweight hardware-isolated microVM (using QEMU or Cloud Hypervisor) with its own guest kernel.
Answer: A union mount filesystem that combines multiple directory layers into a single unified view:
- LowerDir (Read-Only): The immutable image layers.
- UpperDir (Read-Write): The ephemeral container layer where modifications are made.
- WorkDir: Internal scratch space used for atomic copy-up operations.
- MergedDir: The unified filesystem presented to the container.
Answer: A strategy where files in underlying read-only image layers are shared across all containers. When a container attempts to modify a file, the storage driver copies the file from the lower layer up to the upper read-write layer before modifying it, saving storage and memory.
Answer: A Dockerfile containing multiple FROM instructions. The first stage contains heavy compilers, SDKs, and build tools; subsequent stages copy only the compiled binary into a minimal runtime base image (e.g., Distroless or Alpine), reducing image size from 1GB to
Answer: Container base images containing only the application binary and its direct runtime dependencies. They contain zero package managers, zero shells (/bin/sh, /bin/bash), and zero OS utilities, drastically reducing attack surfaces and CVE counts.
Answer:
ENTRYPOINT: Defines the core executable binary that will always run when the container starts.CMD: Provides default arguments passed to theENTRYPOINT.CMDcan be easily overridden via CLI arguments (docker run image arg1), whereasENTRYPOINTrequires--entrypoint.
Answer:
- Exec Form (JSON Array): Spawns the binary directly as PID 1. Receives Unix signals (
SIGTERM,SIGINT) properly for graceful shutdown. - Shell Form: Wraps the command in
/bin/sh -c. The shell becomes PID 1 and typically swallowsSIGTERM, preventing the application from draining connections and causing abrupt kills (SIGKILL).
Answer: Running the container daemon and container processes inside an unprivileged user namespace without host root privileges. If a container breakout exploit occurs, the attacker gains only unprivileged user permissions on the host OS.
Answer: Fine-grained privileges breaking down the monolithic power of Linux root. By default, Docker drops dangerous capabilities; DevOps engineers should drop all capabilities (--cap-drop=ALL) and explicitly add only what is strictly required (--cap-add=NET_BIND_SERVICE).
Answer: --privileged grants the container access to all host devices, disables AppArmor/SELinux security profiles, and gives the container full host root capabilities, making container escapes trivial.
Answer: A child process that has completed execution but remains in the process table because its parent process did not read its exit status (wait() syscall). If PID 1 inside the container does not reap orphaned child processes, the PID table exhausts, crashing the host.
Answer: Lightweight init systems designed to run as PID 1 inside containers to properly forward Unix signals (SIGTERM) to child processes and reap zombie processes.
Answer: The modern container build engine in Docker providing concurrent multi-stage builds, remote cache export/import, secret mounting without committing secrets to layers (RUN --mount=type=secret), and SSH forwarding.
Answer: Using QEMU emulation and binfmt_misc to compile and package container images for multiple CPU architectures (linux/amd64, linux/arm64, linux/riscv64) into a single multi-arch manifest list.
Answer:
- Manifest: A JSON document specifying the config hash, media types, and layer digests that compose an image.
- Manifest List (Index): A higher-level JSON document mapping different CPU architectures/OS combinations to their specific image manifests under a single unified tag.
Answer:
- Bridge (Default): Creates a private virtual Ethernet bridge (
docker0) on the host; containers get private IPs and communicate via NAT/port forwarding. - Host: Container shares the host network namespace directly with zero network isolation (no port mapping needed, maximum throughput).
- None: Container has no network interface except loopback (
lo), completely air-gapped.
Answer: Docker Swarm is Docker's native clustering tool—simple to set up but limited in advanced routing, autoscaling, and ecosystem integrations. Kubernetes is the industry-standard container orchestrator offering declarative reconciliation, service discovery, auto-scaling, complex storage, and extensive extension APIs.
Answer: An instruction in a Dockerfile defining a command (e.g., curl -f http://localhost:8080/healthz || exit 1) run periodically by the container engine to determine if the internal application process is healthy.
Answer:
- Volume: Managed by Docker on the host storage (
/var/lib/docker/volumes/), isolated from the host filesystem. - Bind Mount: Mounts any arbitrary file or directory from the host OS into the container.
- tmpfs Mount: Mounts temporary storage directly in host memory (RAM), never written to disk.
Answer: A tool for defining and running multi-container Docker applications using a declarative YAML configuration file (docker-compose.yml).
Answer: An open-source tool developed by Google that builds container images from a Dockerfile inside a Kubernetes pod or container without requiring a Docker daemon or privileged root access.
Answer:
- Buildah: Specializes in building OCI container images without a daemon.
- Podman: A daemonless container engine for running and managing OCI containers, pods, and images.
- Skopeo: A CLI utility for inspecting, copying, and signing container images directly between remote registries without downloading them locally.
Answer: Use BuildKit secret mounts: RUN --mount=type=secret,id=mysecret cat /run/secrets/mysecret. Secrets are mounted into memory during build time and are never stored in any image layer.
Answer: Docker caches the result of each Dockerfile instruction. If an instruction and its preceding steps are unchanged, Docker reuses the cached layer. Instructions that change frequently (e.g., COPY . .) must be placed at the bottom of the Dockerfile.
Answer: A file that excludes files and directories (node_modules, .git, .env, build logs) from being sent to the Docker daemon as part of the build context, drastically speeding up builds and preventing secret leaks.
Answer: An industry-standard interface specification that allows third-party storage providers (AWS EBS, Ceph, NetApp) to write plugins for block and file storage systems in Kubernetes without modifying core Kubernetes code.
Answer: A CNCF specification and library providing a standardized interface for configuring network interfaces, IP addresses, and routing in Linux containers and Kubernetes pods (e.g., Calico, Cilium, AWS VPC CNI).
Answer: A Kubernetes API interface allowing the kubelet to communicate with heterogeneous container runtimes (containerd, CRI-O) over gRPC without recompiling Kubernetes binaries.
Answer: Alpine uses musl libc and the apk package manager (can cause subtle C-extension bugs with Python/Node). Distroless uses glibc directly from Debian with zero shell or package manager, offering better compatibility and minimal attack surfaces.
Answer: A systemd generator in Podman that translates declarative container configuration files into native systemd service units, managing containers as standard Linux system services.
Answer: A mechanism that allows switching the target Docker CLI engine between different endpoints (e.g., local daemon, remote SSH server, AWS ECS).
Answer: A system utilizing digital signatures to verify the publisher and integrity of container images pulled from registries using Notary.
Answer: Any non-container file type (Helm charts, WebAssembly modules, OPA policies, SBOMs) packaged and stored in standard OCI container registries using OCI specifications.
Answer: Container security and vulnerability analysis tools that analyze image layers, detect CVEs in OS packages and language dependencies, and provide remediation suggestions.
Answer: The declarative specification consumed by runc defining namespaces, mounts, capabilities, process arguments, and resource limits to spawn a container.
Answer: Flatpak and Snap are desktop application sandbox distribution formats; Docker containers are optimized for headless server processes and microservices with standardized OCI specs.
Answer: A malicious process repeatedly spawning infinite child processes to exhaust the kernel PID table. Prevented using cgroups PID limits (--pids-limit 100 or Kubernetes podPidsLimit).
Answer: A Linux kernel parameter (--memory-swappiness) controlling how aggressively the kernel swaps container anonymous memory pages to disk under memory pressure.
Answer: The Linux kernel Out-Of-Memory Killer terminating the container process when its physical memory usage exceeds the assigned cgroup memory limit (limits.memory).
Answer: The kernel enforcing container CPU limits over a period (usually 100ms). If a container exhausts its allocated quota within the first 20ms, it is throttled (frozen) for the remaining 80ms, causing severe latency spikes.
Answer: A Linux kernel security facility that filters the system calls a container process is allowed to make to the kernel, blocking dangerous syscalls (e.g., reboot, ptrace).
Answer: Mandatory Access Control (MAC) kernel security modules that enforce granular filesystem and process permission profiles on containers regardless of user privileges.
Answer: A CLI command that cleans up unused Docker data: stopped containers, dangling images, unused build caches, and networks.
Answer: An out-of-process extension that adds capabilities to the Docker daemon (custom network drivers, volume drivers, logging drivers).
Answer: The Linux firewall utility Docker configures to route packets between physical interfaces and virtual bridge interfaces (docker0), performing Network Address Translation (NAT).
Answer: The persistent background process that manages Docker objects (images, containers, networks, volumes) and listens for Docker REST API requests.
Answer: An immutable, content-addressable SHA256 cryptographic hash of the image manifest (e.g., image@sha256:abcd...), guaranteeing identical image bits regardless of mutable tag changes.
Answer: When an instruction in a Dockerfile changes, Docker invalidates the cache for that instruction and forces all subsequent instructions to execute from scratch.
Answer: The UNIX domain socket used by the Docker CLI to communicate with the Docker daemon. Mounting this socket into a container grants that container full root control over the host Docker daemon.
Answer: Cryptographically signed statements attached to container images verifying their build provenance, test results, and vulnerability scanning compliance.
Answer: An adapter shim that allows Kubernetes to continue using the legacy Docker Engine as a CRI runtime after Kubernetes removed built-in dockershim.
Answer: An open-source virtualization technology developed by AWS that creates lightweight, secure microVMs in sub-5 milliseconds with minimal memory footprint, powering AWS Lambda and Fargate.
Answer:
kube-apiserver: The centralized REST API gateway and validation hub; the only component that communicates withetcd.etcd: The distributed, highly consistent key-value store holding the entire cluster state.kube-scheduler: Assigns unscheduled Pods to optimal worker nodes based on resource requests, affinity, taints, and topology.kube-controller-manager: Runs core reconciliation control loops (Node Lifecycle, Deployment, ReplicaSet, ServiceAccount controllers).cloud-controller-manager: Integrates with underlying cloud provider APIs for Load Balancers, Route tables, and Block Volumes.
Answer:
kubelet: The node agent that communicates with the API server, watches PodSpecs assigned to the node, and commands the CRI runtime to start containers.kube-proxy: Manages network routing and iptables/IPVS packet filtering rules on each node to implement KubernetesServiceabstractions.- Container Runtime (CRI): The runtime (containerd/CRI-O) executing container processes.
Answer: etcd is an append-only distributed key-value store implementing the Raft consensus algorithm. It requires an odd number of nodes (3 or 5) to form a quorum (
Answer: The smallest deployable computing unit in Kubernetes. A Pod encapsulates one or more co-located containers that share the exact same Network namespace (IP address and port space), IPC namespace, and shared storage volumes.
Answer: A minimal container running an infinite sleep loop spawned first in every Pod. It establishes and holds the shared Network, IPC, and UTS namespaces. Application containers join the pause container's namespaces (--net=container:pause), allowing containers in the pod to communicate over localhost.
Answer:
- Pod: The running instance of containers.
- ReplicaSet: Ensures a specified number of identical Pod replicas are running at any given time.
- Deployment: A higher-level controller managing declarative rolling updates, rollbacks, and version history for ReplicaSets.
Answer: StatefulSet manages stateful applications requiring:
- Stable, unique network identifiers (
pod-0,pod-1). - Ordered, sequential deployment and rolling updates.
- Dedicated, persistent volume bindings via
volumeClaimTemplatesthat persist across pod rescheduling.
Answer: Ensures that all (or selected) worker nodes run exactly one copy of a specific Pod. Used for node-level agents (Cilium CNI, Fluent Bit log shippers, Prometheus node_exporter).
Answer:
- Job: Spawns one or more pods and ensures a specified number of them terminate successfully (run-to-completion batch tasks).
- CronJob: Schedules and executes
Jobobjects periodically based on a standard cron expression.
Answer:
- ClusterIP (Default): Exposes the Service on an internal cluster IP, accessible only within the cluster.
- NodePort: Exposes the Service on a static high port (
30000-32767) on every worker node's physical IP. - LoadBalancer: Provisions an external cloud load balancer (e.g., AWS NLB/ALB) and routes traffic to NodePort/ClusterIP.
- ExternalName: Maps the Service to a CNAME DNS record (e.g.,
my-db.external.com) without proxying.
Answer: A Service that does not allocate a ClusterIP. CoreDNS returns the individual A records of all backing Pod IPs directly, allowing clients to establish direct peer-to-peer connections (used for StatefulSets, Kafka, and Cassandra).
Answer:
- Ingress: A declarative API resource defining Layer 7 HTTP/HTTPS routing rules, TLS termination, and host/path routing.
- Ingress Controller: The actual reverse proxy daemon (Nginx Ingress, Envoy, Traefik) that watches Ingress resources and configures routing rules.
Answer: The modern successor to Ingress providing role-oriented resource separation:
GatewayClass(Cluster Operator defines infrastructure).Gateway(Platform Engineer configures listeners, ports, TLS).HTTPRoute/GRPCRoute(Developer defines routing paths and traffic splitting).
Answer: Virtual clusters within a physical cluster providing logical scoping for object names, RBAC policies, and ResourceQuotas. Namespaces do not provide network isolation by default (requires NetworkPolicies).
Answer:
- ConfigMap: Stores non-confidential configuration data in key-value pairs mounted as environment variables or volume files.
- Secret: Stores sensitive data (passwords, tokens, keys) encoded in base64. Stored encrypted at rest in
etcdwhen KMS encryption is enabled.
Answer:
- PV: A storage resource provisioned in the cluster (e.g., AWS EBS volume) by an administrator or StorageClass.
- PVC: A request for storage by a user specifying size and access modes (
ReadWriteOnce,ReadOnlyMany,ReadWriteMany).
Answer: A mechanism where Kubernetes automatically provisions underlying cloud storage volumes (e.g., creates an AWS gp3 volume via CSI) on-demand when a developer creates a PVC referencing a StorageClass.
Answer:
- RWO (ReadWriteOnce): Volume can be mounted as read-write by a single worker node.
- ROX (ReadOnlyMany): Volume can be mounted as read-only by multiple worker nodes simultaneously.
- RWX (ReadWriteMany): Volume can be mounted as read-write by many worker nodes concurrently (e.g., AWS EFS, NFS).
- RWOP (ReadWriteOncePod): Volume can be mounted as read-write by a single Pod instance only.
Answer: An identity assigned to Pods to authenticate against the Kubernetes API Server. Tokens are mounted automatically at /var/run/secrets/kubernetes.io/serviceaccount/token.
Answer:
- Role: Grants API permissions (verbs:
get,list,create) scoped strictly to a single namespace. - ClusterRole: Grants cluster-wide permissions across all namespaces and cluster-scoped resources (Nodes, PVs).
- RoleBinding: Assigns a Role or ClusterRole to subjects within a specific namespace.
- ClusterRoleBinding: Assigns a ClusterRole to subjects cluster-wide.
Answer: Plugins that intercept API Server requests after authentication and authorization, but before the object is persisted to etcd.
- Mutating Admission Controllers: Modify objects (e.g., injecting sidecars).
- Validating Admission Controllers: Enforce security policies and accept or reject objects (e.g., Kyverno, OPA).
Answer:
- Startup Probe: Validates if legacy slow-booting applications have started. Disables liveness and readiness checks until it passes.
- Liveness Probe: Checks if the application process is running and healthy. If it fails,
kubeletrestarts the container. - Readiness Probe: Checks if the container is ready to accept incoming network traffic. If it fails, the Pod's IP is removed from Service endpoints.
Answer: Specialized containers that run and complete sequentially before primary application containers start. Used for blocking tasks (waiting for database availability, running migrations).
Answer: Temporary containers injected into existing running Pods for interactive debugging and network troubleshooting (kubectl debug) without restarting the Pod.
Answer: A policy limiting the number of Pod replicas of an application that can be unavailable simultaneously during voluntary disruptions (node draining, cluster upgrades).
Answer:
- Requests: The guaranteed minimum CPU and memory allocated for the Pod. Used by
kube-schedulerto place the Pod on a node. - Limits: The hard upper bound of CPU and memory the Pod is allowed to consume. CPU exceeding limits causes CFS throttling; memory exceeding limits triggers OOMKill.
Answer:
- Guaranteed: Every container has equal CPU and Memory Requests and Limits. Evicted last during resource exhaustion.
- Burstable: Requests are set and lower than Limits. Evicted when Guaranteed pods require resources.
- BestEffort: Zero requests and limits set. Evicted first during node memory pressure.
Answer:
- Node Affinity: Constrains which nodes a Pod can schedule on based on node labels (
requiredDuringScheduling...orpreferredDuringScheduling...). - Pod Anti-Affinity: Prevents multiple replicas of the same service from scheduling on the same node or AZ to ensure high availability.
Answer:
- Taint: Applied to a Node to repel Pods (
key=value:NoSchedule,NoExecute). - Toleration: Applied to a Pod allowing it to schedule on matching tainted nodes (e.g., scheduling GPU workloads on dedicated GPU nodes).
Answer: Rules that control how Pods are evenly spread across failure domains (Availability Zones, racks, nodes) based on label selectors (maxSkew: 1).
Answer: A proactive subsystem in kubelet that monitors node memory, disk, and PID usage. When thresholds are breached (memory.available < 100Mi), it evicts lower-priority pods to prevent node failure.
Answer: Defines the relative importance of Pods (value: 1000000). If a high-priority Pod cannot schedule due to lack of resources, the scheduler preempts (evicts) lower-priority pods to free up capacity.
Answer: A control loop that automatically adjusts the number of Pod replicas in a Deployment based on observed CPU/memory utilization or custom external metrics.
Answer: A controller that automatically recommends or updates CPU and memory resource requests/limits based on historical runtime analysis.
Answer: An autoscaler that scales Kubernetes workloads from 0 to hundreds based on external event sources (Kafka consumer lag, AWS SQS queues, RabbitMQ, Prometheus metrics).
Answer: An open-source, high-performance node autoscaler developed by AWS that bypasses native Auto Scaling Groups, provisioning right-sized EC2 instances directly via EC2 Fleet API in seconds.
Answer: An automated feature where Karpenter continuously evaluates cluster compute waste, drains pods from underutilized nodes, and replaces them with cheaper, right-sized single instances.
Answer: A traditional autoscaler that adds or removes worker nodes by modifying the desired capacity of underlying cloud Auto Scaling Groups (ASGs) when Pods enter Pending state.
Answer: The internal DNS server running as a cluster service that resolves internal service names (my-svc.my-ns.svc.cluster.local) and external domains for Pods.
Answer: A DaemonSet that runs a local DNS caching agent on each worker node, reducing CoreDNS network queries, UDP packet drops, and glibc ndots:5 latency storms.
Answer:
- Endpoints: A monolithic object tracking all pod IPs backing a service; updating 1 pod in a 5,000-pod service rewritten the entire object, overloading etcd.
- EndpointSlice: Splits backing pod endpoints into scalable chunks of 100 endpoints each, vastly improving cluster scalability.
Answer:
- iptables: Linear sequential rule evaluation ($O(N)$). As services reach thousands, packet processing latency degrades significantly.
- IPVS (IP Virtual Server): Linux kernel hash table lookup ($O(1)$). Provides consistent sub-millisecond throughput regardless of service scale.
Answer: A next-generation CNI plugin that replaces kube-proxy and iptables with eBPF programs loaded directly into Linux kernel socket layers, providing
Answer: A declarative firewall resource that controls traffic flow between Pods at Layer 3 and Layer 4 using label selectors.
Answer: A policy selecting all pods in a namespace with empty ingress/egress rules, blocking all incoming and outgoing network traffic until explicitly whitelisted.
Answer: The built-in replacement for legacy PodSecurityPolicies (PSP) that enforces three security profiles (Privileged, Baseline, Restricted) via namespace labels.
Answer: An admission controller that enforces custom organizational governance policies written in Rego across Kubernetes manifests.
Answer: A Kubernetes-native policy engine written in pure YAML that validates, mutates, and generates resources and verifies container image signatures.
Answer: The package manager for Kubernetes. Helm Hooks (pre-install, post-upgrade, pre-delete) allow executing specific jobs (database migrations) at defined lifecycle points during chart releases.
Answer: Defining subcharts in Chart.yaml and locking their exact semantic versions and tarball hashes in Chart.lock.
Answer: A JSON Schema file bundled inside a Helm chart that validates user-provided values.yaml inputs before template rendering.
Answer: A template-free configuration customization tool built directly into kubectl (kubectl apply -k) that uses declarative overlays on top of base YAML files.
Answer: A patching mechanism that merges overlay YAML blocks into base manifests based on resource names and keys rather than overwriting arrays.
Answer: An API extension mechanism that allows developers to define custom object kinds (e.g., kind: KafkaTopic) that the Kubernetes API Server stores and manages like native resources.
Answer: A software pattern combining CRDs with a custom control loop that watches state and executes domain-specific automation (managing database failovers, taking automated backups).
Answer: Go SDKs and scaffolding tools providing code generators, webhook handlers, and controllers for authoring production Kubernetes operators.
Answer: A consensus mechanism using Lease API objects in coordination.k8s.io to ensure only one active replica of a controller manager executes reconciliation loops at a time.
Answer:
- Compaction: Discards historical revision logs from etcd storage.
- Defragmentation: Reclaims disk space and reorganizes on-disk database pages to prevent reaching etcd's 8GB storage quota.
Answer:
alpha: Disabled by default; experimental; may be dropped without notice.beta: Enabled by default; tested; schema backward-compatibility guaranteed until deprecation.v1(GA): Stable; long-term support and backward compatibility.
Answer: Stable APIs cannot be removed without being deprecated for at least 12 months (or 3 consecutive minor releases).
Answer: Allows extending the Kubernetes API by registering independent secondary API servers (e.g., Metrics Server) behind the main API Server.
Answer: An in-memory, cluster-wide aggregator of resource usage data (CPU, memory) scraped from Kubelets, consumed by kubectl top and HPA.
Answer: An operator that automates deploying and managing Prometheus monitoring stacks on Kubernetes using CRDs (ServiceMonitor, PrometheusRule).
Answer: A declarative CRD that discovers and configures Prometheus scrape targets based on Kubernetes Service labels and endpoints.
Answer: An operator that syncs secrets from external systems (AWS Secrets Manager, HashiCorp Vault) into native Kubernetes Secret resources.
Answer: A Kubernetes add-on that automates the issuance and renewal of TLS certificates from Let's Encrypt, HashiCorp Vault, and private PKIs.
Answer: The mechanism where deleting an owner resource (Deployment) automatically deletes all its dependent children (ReplicaSets, Pods) via ownerReferences.
Answer: Pre-delete hooks in resource metadata that prevent an object from being deleted from etcd until external cleanup operations (releasing cloud storage) complete.
Answer: RollingUpdate with maxUnavailable (e.g., maxUnavailable: 1) updates DaemonSet pods on worker nodes one at a time.
Answer:
- Cordon (
kubectl cordon): Marks the node as unschedulable; existing pods remain running. - Drain (
kubectl drain): Cordons the node and safely evicts all running pods (respecting PDBs) so the node can be rebooted or terminated.
Answer:
exec: Spawns a new process inside the container.attach: Attaches standard input/output streams to the existing running primary process (PID 1).
Answer: Establishes a direct TCP tunnel from a local workstation port to a specific Pod or Service port inside the private Kubernetes cluster.
Answer:
Ignore: If the webhook server is unreachable, the API Server admits the resource anyway.Fail: If the webhook is unreachable, the API Server rejects the resource (mandatory for security admission controllers).
Answer: The time Kubernetes allows a Pod to cleanly shut down (drain connections, complete transactions) after sending SIGTERM before forcibly killing it with SIGKILL (default: 30 seconds).
Answer: A script or HTTP call executed inside the container before the SIGTERM signal is sent, used to initiate graceful connection draining.
Answer: A field-level management mechanism where the API Server tracks which controller or user owns specific YAML fields via managedFields, preventing accidental field overwrites.
Answer: Mounting a specific sub-directory or single file from a volume into a container path rather than mounting the entire root volume directory.
Answer: Combining multiple volume sources (Secrets, ConfigMaps, DownwardAPI, ServiceAccountTokens) into a single unified directory mount inside the container.
Answer: Exposing Pod and container metadata (Pod IP, Node name, namespace, resource limits) to the application via environment variables or volume files.
Answer: Ephemeral records stored in etcd (retained for 1 hour) capturing lifecycle state changes, scheduling decisions, warnings, and errors across cluster objects.
141. Scenario: A Pod is stuck in CrashLoopBackOff. Walk through your step-by-step diagnostic workflow.
Answer:
- Check Pod status and restart count:
kubectl get pod <name> -o wide. - Inspect lifecycle events and exit codes:
kubectl describe pod <name>. - Check application logs:
kubectl logs <name>. - If the container crashed instantly, check previous crash logs:
kubectl logs <name> --previous. - Common Root Causes: Missing environment variables/secrets, database connection timeouts, wrong command/entrypoint syntax, unhandled exception in initialization code.
Answer:
-
Reason: The container was killed by
SIGKILL($128 + 9 = 137$ ) sent by the Linux kernel Out-Of-Memory (OOM) Killer because physical memory usage exceeded the cgrouplimits.memorythreshold. -
Diagnostics: Verify via
kubectl describe pod(look forLast State: Terminated, Reason: OOMKilled). -
Fix: Increase container memory limits, fix Java JVM heap flags (
-XX:MaxRAMPercentage=75.0), or profile application memory leaks using continuous profiling tools.
Answer: The container received a graceful SIGTERM signal (terminationGracePeriodSeconds window.
Answer:
- Incorrect image tag or typo in image repository URL.
- Missing or misconfigured
imagePullSecretsfor authenticating to private registries. - Network timeout or firewall blocking egress from worker node to registry.
- Hitting public Docker Hub rate limits (fixed via ECR Pull Through Cache).
Answer: Run kubectl describe pod <name> and check Events:
- Insufficient CPU / Memory: No single worker node has enough unreserved capacity to satisfy the Pod's
requests. - NodeAffinity / Taints: Pod lacks tolerations for tainted worker nodes.
- PVC Unbound: PersistentVolumeClaim cannot bind to any available PV or StorageClass provisioner fails.
- NodeSelector Mismatch: No nodes match the requested labels.
Answer:
- Check for blocking Finalizers:
kubectl get pod <name> -o jsonpath='{.metadata.finalizers}'. - Check if a CSI volume unmount is hanging on a dead worker node.
- If node is dead and volume is safe, remove finalizer or force delete:
kubectl delete pod <name> --grace-period=0 --force.
147. Scenario: Explain the Linux ndots:5 DNS resolution issue in Kubernetes and how it causes latency storms.
Answer:
- By default,
/etc/resolv.confinside pods setsoptions ndots:5. - If a query has fewer than 5 dots (e.g.,
api.stripe.comhas 2 dots), the resolver sequentially appends all search domains:api.stripe.com.default.svc.cluster.local(NXDOMAIN)api.stripe.com.svc.cluster.local(NXDOMAIN)api.stripe.com.cluster.local(NXDOMAIN)api.stripe.com(SUCCESS)
- Impact: Generates 3 failed queries per external lookup, overwhelming CoreDNS.
- Fix: Deploy NodeLocal DNSCache, set
ndots:2in poddnsConfig, or append a trailing dot (api.stripe.com.).
148. Scenario: How do you achieve Zero-Downtime deployments for high-throughput HTTP services during rolling updates?
Answer:
- Add
preStopSleep: Ingress controllers take 1–3 seconds to update iptables endpoints. Add apreStophook (sleep 5) so the old pod continues servicing in-flight requests while being removed from endpoints. - Handle
SIGTERMGracefully: Ensure the application process interceptsSIGTERMand drains active HTTP connections. - Configure Accurate Readiness Probes: Ensure readiness probe fails only after connection draining starts.
- Tune Deployment Strategy: Set
maxSurge: 25%andmaxUnavailable: 0.
Answer: An open-source service mesh providing traffic management, zero-trust security (mTLS), and observability:
istiod(Control Plane): Translates declarative routing rules and distributes certificates to proxies.- Envoy Proxy (Data Plane): High-performance sidecar proxy intercepting all inbound and outbound network traffic.
Answer: A sidecarless architecture splitting mesh processing into:
ztunnel(Zero Trust Tunnel): Per-node Layer 4 daemon handling mutual TLS encryption.- Waypoint Proxies: Optional per-namespace Envoy instances handling Layer 7 routing and authorization, reducing memory overhead by over 70%.
Answer: A lightweight, ultralow-latency service mesh written in Rust (data plane micro-proxy) and Go (control plane), designed for simplicity and minimal CPU/memory footprint compared to Envoy-based meshes.
Answer: Sidecarless service mesh leveraging Linux kernel eBPF to manage Layer 7 routing, mTLS, and observability directly in the kernel, bypassing userspace proxy hops for Layer 4 traffic.
Answer: An observability platform running on top of Cilium and eBPF providing real-time service dependency graphs, network flow logs, and HTTP/DNS latency metrics.
Answer: Connects multiple independent Kubernetes clusters at the network layer, providing cross-cluster pod IP routing, global service discovery, and cross-cluster failover.
Answer: Running at least 3 stacked master nodes across independent Availability Zones, deploying a highly available etcd quorum, and placing an external Layer 4 Load Balancer (NLB) in front of the API Servers.
Answer: When network partitions isolate etcd nodes into two groups that both elect leaders, corrupting data. Prevented by requiring an odd number of nodes (3 or 5) and enforcing strict majority quorum (
Answer:
- Snapshot:
etcdctl snapshot save /backup/etcd-snapshot.db - Restore:
etcdctl snapshot restore /backup/etcd-snapshot.db --data-dir=/var/lib/etcd-restored
Answer: Protects the API Server from request floods by classifying incoming requests into priority levels and flow schemas, queuing and shedding non-essential requests to guarantee critical control loops continue running.
Answer: Automatic renewal of client and server TLS certificates used by Kubelets to authenticate against the API Server before expiration, managed via CSRs.
Answer: A declarative configuration file managing system resource reservations (--system-reserved, --kube-reserved), container log sizes, and eviction thresholds.
Answer:
system-reserved: Reserves CPU and RAM for OS daemons (systemd, sshd, udev).kube-reserved: Reserves CPU and RAM for Kubernetes daemons (kubelet, containerd).- Without reservations, application pods can consume 100% of node RAM, crashing the kubelet.
Answer: Anti-Affinity is binary (schedule or do not schedule on the same node); Topology Spread Constraints allow configuring proportional distribution (maxSkew: 1) across zones.
Answer: A Kubernetes CRD (VolumeSnapshot) that triggers point-in-time storage array snapshots of persistent volumes via the underlying CSI driver.
Answer: Increasing the capacity of an existing PersistentVolumeClaim dynamically (spec.resources.requests.storage: 100Gi) if the StorageClass has allowVolumeExpansion: true.
Answer: Ephemeral storage backed by dynamic volume provisioners that create dedicated scratch disks for pods and delete them automatically when the pod terminates.
Answer: Configuring an EncryptionConfiguration file on the API Server using a KMS provider (AWS KMS, HashiCorp Vault) to encrypt Secret data before writing to etcd.
Answer: Chronological security records capturing every request made to the API Server (who requested what, when, from which IP, and whether it was authorized).
Answer: An open-source tool that checks whether a Kubernetes cluster is configured securely according to the Center for Internet Security (CIS) Kubernetes Benchmark.
Answer: An open-source penetration testing tool that hunts for security vulnerabilities and exposed ports in live Kubernetes clusters.
Answer: Mounting the host root filesystem (hostPath: /) into a container allows an attacker to modify /etc/shadow, install cron backdoors on the host, or escape namespaces.
Answer: Mounting /var/run/docker.sock allows code inside the container to command the host Docker daemon to spawn privileged sibling containers with host root access.
172. What is Kubernetes Workload Identity Federation (AWS IRSA / GCP Workload Identity / Azure Pod Identity)?
Answer: Projecting OIDC-signed ServiceAccount tokens into Pods, allowing them to exchange tokens with cloud STS for short-lived IAM credentials without static access keys.
Answer: AWS EKS native agent-based identity mapping that binds Kubernetes ServiceAccounts directly to AWS IAM roles without needing OIDC provider trust configurations.
Answer: A framework that turns Kubernetes into a universal control plane, allowing platform teams to manage cloud infrastructure (S3 buckets, RDS databases) using native Kubernetes CRDs.
Answer: Running fully isolated virtual Kubernetes control planes inside namespaces of an underlying host cluster, providing multi-tenancy with separate API servers and CRD support.
Answer: A Kubernetes sub-project that uses declarative APIs to automate provisioning, upgrading, and operating multiple Kubernetes clusters across AWS, Azure, GCP, and vSphere.
Answer: Modular plugins (e.g., kubernetes, forward, cache, errors, health) configured in the Corefile that determine how DNS requests are processed.
Answer: Dynamically scaling CoreDNS pod replicas using cluster-proportional-autoscaler based on the total number of worker nodes and cores in the cluster.
Answer:
Exact: Matches URL path exactly case-sensitively (/api/v1).Prefix: Matches URL paths sharing a prefix (/apimatches/api/usersand/api/orders).
Answer: Sequential processing stages (TLS Inspector, HTTP Connection Manager, Rate Limit, RBAC, Router) executed for each network connection in Envoy.
Answer: Automatically injecting required metadata (labels, annotations, resource requests) into Kubernetes manifests upon submission.
Answer: Kyverno rules targeting Pods automatically generate identical validation rules for higher-level controllers (Deployments, StatefulSets, DaemonSets).
Answer: A tool that connects a local workstation development process directly into a remote Kubernetes cluster network, proxying cluster traffic to local code.
Answer: A Google CLI tool that automates the continuous local build, push, and deployment loop to local (Minikube/Kind) or remote Kubernetes clusters.
Answer: A local development environment tool that monitors file changes and executes fast multi-container sync updates directly to running pods.
Answer: A tool for running local multi-node Kubernetes clusters where each node is a Docker container, heavily used in CI testing.
Answer: A lightweight, fully compliant certified Kubernetes distribution packaged as a single
Answer: A zero-ops, lightweight upstream Kubernetes distribution packaged as a snap for developer workstations and edge devices.
Answer: A local single-node Kubernetes tool designed for learning and local testing running inside a VM or container.
Answer: APIs are transitioned from v1alpha1 v1beta1 v1, with deprecation warnings emitted in API responses before final removal in subsequent releases.
Answer: The official plugin manager for kubectl allowing engineers to discover and install community extensions (kubectl ctx, kubectl ns, kubectl neat).
Answer: A terminal-based UI that provides a real-time dashboard to monitor, manage, and debug Kubernetes cluster resources.
Answer: A wrapper for kubectl that colorizes CLI terminal outputs for pods, services, and events to improve readability.
Answer: Automatically terminating and recycling worker nodes after a defined lifespan (e.g., 30 days) to enforce mandatory OS security patching.
Answer: Policies that limit how many nodes Karpenter can terminate or consolidate simultaneously during voluntary cluster optimization.
Answer: Allocating both IPv4 and IPv6 addresses to Pods and Services across the entire cluster.
Answer: A bare-metal load-balancer implementation for Kubernetes clusters running on-premise without cloud provider LoadBalancer integration.
Answer: Allowing Cilium to advertise Pod and Service CIDR IP blocks directly to physical data center Top-of-Rack (ToR) BGP routers.
Answer: Creating a new PersistentVolume populated with the exact data duplicate of an existing PVC on the same storage array.
Answer: Attaching ephemeral storage directly inside the spec.volumes block without needing standalone PVC manifests.
201. Scenario: You deploy a new version of an app and CPU usage spikes to 100%, causing pods to crash. How do you triage?
Answer:
- Check HPA and deployment events:
kubectl describe hpa <name>. - Grab CPU profile or thread dump from a running pod:
kubectl exec -it <pod> -- jstack 1or use continuous profiling (Pyroscope). - Roll back deployment immediately:
kubectl rollout undo deployment/<name>. - Investigate root cause: Infinite loops, un-indexed database queries holding locks, or regex backtracking.
Answer:
- Scale up CoreDNS memory limits in
kube-dnsdeployment. - Enable NodeLocal DNSCache DaemonSet to absorb 80% of DNS queries on worker nodes.
- Deploy
cluster-proportional-autoscalerto scale CoreDNS replicas dynamically with cluster node count. - Fix
ndots:5configuration in application PodSpecs.
203. Scenario: Worker nodes are intermittently running out of disk space on /var/lib/containerd. How do you resolve it?
Answer:
- Clean up unused images immediately:
crictl rmi --prune. - Configure image garbage collection thresholds in
kubelet-config.yaml:imageGCHighThresholdPercent: 80 imageGCLowThresholdPercent: 65
- Set container log rotation limits:
containerLogMaxSize: 50MiandcontainerLogMaxFiles: 3.
204. Scenario: A developer accidentally deletes a production Namespace. How does Kubernetes behave and how do you prevent this?
Answer:
- Behavior: Kubernetes deletes all resources within the namespace (Pods, Services, PVCs, Secrets) via cascading deletion.
- Prevention:
- Deny
delete namespacesvia RBAC. - Enforce Kyverno / OPA Gatekeeper validating admission webhooks blocking namespace deletion for production namespaces.
- Set
deletionProtection: trueon critical cloud resources via Crossplane/Terraform.
- Deny
205. Scenario: Inter-pod communication fails between nodes, but works between pods on the same node. What is the root cause?
Answer:
- Security Group / Firewall: Cloud Security Group is blocking cross-node encapsulation traffic (Geneve/VXLAN port 6081 or 8472).
- CNI Node Routing: CNI overlay routing table on the host is misconfigured or missing route entries to the peer node's Pod CIDR.
- MTU Mismatch: Overlay network MTU (1450 bytes) exceeds host network MTU (1500 bytes), dropping fragmented packets.
Answer:
- Check API Priority and Fairness metrics in Prometheus (
apiserver_flow_control_rejected_requests_total). - Identify misbehaving microservices or buggy controllers polling the API in tight loops without informers/watchers.
- Scale up API Server replicas and tune
max-requests-inflightsettings.
Answer: etcd leader heartbeats time out, triggering frequent leader elections, dropping consensus, and causing API Server requests to hang and Kubelets to fail status updates. Requires dedicated high-IOPS NVMe storage.
208. Scenario: How do you migrate 1,000 workloads from legacy Ingress to Kubernetes Gateway API with zero downtime?
Answer:
- Deploy Gateway API CRDs and Gateway Controller (Envoy Gateway / Cilium).
- Deploy
Gatewayinstance sharing the same external Load Balancer IP or DNS. - Incrementally deploy
HTTPRouteresources alongside existingIngressresources. - Shift DNS traffic gradually using Route 53 weighted records.
- Decommission Ingress resources once 100% traffic is verified on Gateway API.
209. Scenario: How do you troubleshoot a pod whose logs show nothing and process terminates immediately?
Answer:
- Run with interactive override:
kubectl run debug --image=<image> -it -- /bin/sh. - Inspect exit code:
kubectl describe pod <name>(look for Exit Code 127 = missing binary, or Exit Code 126 = permission denied). - Verify library dependencies (
ldd binary_name).
Answer:
- On control plane nodes:
kubeadm certs check-expiration. - Renew certificates:
kubeadm certs renew all. - Restart control plane static pods (
kube-apiserver,kube-controller-manager,kube-scheduler). - Update
~/.kube/configwith renewed client credentials.
Answer: A DaemonSet that monitors node health issues (kernel deadlocks, corrupted filesystems, hardware errors) and reports them as Node Conditions to the API Server.
Answer: A cleanup controller that automatically deletes expired, orphaned, or non-production Kubernetes resources based on TTL annotations (janitor/ttl: 7d).
Answer: A cluster sanitizer CLI tool that scans live clusters and reports misconfigurations, dead resources, and over-allocated requests.
Answer: A controller that automatically watches ConfigMaps and Secrets and triggers rolling restarts on dependent Deployments when secret values change.
Answer: A utility that uses VPA in recommendation mode to generate optimal CPU and memory resource request/limit baselines.
Answer: A cluster optimizer that evicts running Pods that no longer satisfy scheduling criteria (e.g., node underutilization, broken affinity rules, high pod count) so the scheduler can balance them.
Answer: A cloud-native chaos engineering platform for Kubernetes that injects faults (killing pods, adding network latency, corrupting I/O) via CRDs.
Answer: An add-on that listens to the Kubernetes API Server and generates Prometheus metrics about the health and state of objects (deployment replicas, pod phases, resource limits).
Answer: Standard pods get isolated private IPs from the CNI; Host-Network pods (hostNetwork: true) bind directly to the host OS network interface, bypassing CNI virtualization.
Answer: A feature allowing different Pods in the same cluster to use different container runtimes (e.g., standard pods use runc, untrusted pods use gVisor / kata).
Answer: Configuring kernel parameters per pod namespace (e.g., net.core.somaxconn=1024) via spec.securityContext.sysctls.
Answer:
Immediate: PV is provisioned immediately upon PVC creation.WaitForFirstConsumer: PV is provisioned only after a Pod using the PVC is scheduled, ensuring the storage volume is created in the exact Availability Zone where the Pod lands.
Answer: Ensuring CSI storage drivers allocate block storage volumes in the exact cloud Availability Zone matching the scheduled worker node.
Answer: Setting requests and limits for container writable layers and emptyDir volumes (ephemeral-storage: 2Gi) to prevent pods from filling node root disks.
Answer: A CSI driver that mounts secrets from AWS Secrets Manager, Azure Key Vault, or HashiCorp Vault directly into pod filesystems as in-memory files without storing them in etcd.
Answer: A webhook that converts Custom Resources between different schema versions (e.g., v1alpha1 v1) on-the-fly during API reads/writes.
Answer:
- Typed Client (
clientset): Works with compiled static Kubernetes Go structs. - Dynamic Client: Works with unstructured JSON objects (
unstructured.Unstructured), essential for managing arbitrary CRDs.
Answer: A local in-memory cache of API Server objects that watches etcd events and triggers event handler callbacks (OnAdd, OnUpdate, OnDelete) without polling the API Server.
Answer: A thread-safe, rate-limiting queue that buffers reconcile requests and deduplicates multiple updates for the same object key.
Answer: An idempotent function: Reconcile(Request) -> (Result, error) that compares current cluster state against desired state in Git/CRD and executes CRUD operations to bring them into alignment.
Answer: Every object in etcd has a resourceVersion. When updating an object, the API Server rejects writes if the client's resourceVersion does not match etcd's current version (HTTP 409 Conflict).
Answer: Controllers do not expect immediate state synchronization; they continuously retry failed operations until the system reaches desired state.
Answer: Dedicated API endpoints for updating specific parts of an object (e.g., updating /status does not trigger spec mutating admission webhooks).
Answer: Metadata pointing a child object to its parent controller, ensuring garbage collection deletes child pods when parent deployments are deleted.
Answer: CoreDNS returns SRV records containing port numbers and hostnames for each named port of backing pods in a Headless Service.
Answer: Directing requests from the same client IP to the same backing pod replica for a configurable timeout duration (sessionAffinityConfig.clientIP.timeoutSeconds).
Answer:
Cluster(Default): Node receiving traffic SNATs packet and forwards to any node running backing pods (extra network hop, obscures client IP).Local: Node routes traffic only to local pods on the same node (preserves client source IP, zero extra hop, requires health check node ports).
Answer: Routing Service traffic to endpoints that are topologically closest to the caller (same node
Answer: The subsystem in CNI plugins responsible for allocating and tracking available IP address blocks for nodes and pods.
Answer: Allocating IPv4 /28 subnets (16 IPs) to each ENI instead of individual secondary IPs, increasing the maximum number of pods per EC2 instance by over 4x.
Answer:
- Azure CNI: Every pod gets a real IP from the Azure VNet (fast, but consumes large VNet CIDR space).
- Kubenet: Pods get private IPs behind NAT; routes configured in Azure route tables.
Answer: Replaces standard Linux iptables routing with high-performance eBPF programs, providing native WireGuard encryption and scalable NetworkPolicies.
Answer: Automatically encrypting all node-to-node and pod-to-pod network traffic using kernel WireGuard with zero sidecar proxy overhead.
Answer: Running dummy pause pods with low PriorityClass that hold reserve compute capacity. When real pods schedule, they preempt the pause pods, launching immediately while nodes scale in the background.
Answer: Configuring Karpenter to dynamically mix diverse EC2 instance types, sizes, and generations across Spot and On-Demand pools to maximize availability and minimize cost.
Answer: A standardized Kubernetes API (ServiceExport, ServiceImport) allowing services in Cluster A to discover and communicate with services in Cluster B using my-svc.my-ns.svc.clusterset.local.
Answer: An open-source tool creating direct IPsec/WireGuard VPN tunnels between pods in disparate Kubernetes clusters across clouds and data centers.
Answer: An admission policy rule that verifies container image cryptographic signatures against Rekor transparency logs before allowing pods to schedule.
Answer: A CNCF security tool that intercepts kernel system calls (execve, openat, socket) to detect anomalous runtime behavior (shell spawning, privilege escalation) inside containers.
Answer: Combining vcluster (isolated control planes), Cilium NetworkPolicies (isolated L3-L7 networks), gVisor / Kata Containers (sandboxed kernel runtimes), and ResourceQuotas to run untrusted multi-tenant workloads safely on a shared physical cluster.