Linux Container Internals: cgroups v2 & Namespaces
Demystify Docker and Kubernetes from the ground up: the 8 Linux Namespaces (`CLONE_NEWPID`, `CLONE_NEWNET`, `CLONE_NEWNS`, `CLONE_NEWUSER`), cgroups v2 unified resource hierarchy (`cpu.max`, `memory.max`, `io.weight`), and creating isolated containers using `unshare` and `pivot_root`.
What You Will Learn in This Lesson
- Why containers are NOT virtual machines: containers are ordinary Linux processes isolated by kernel Namespaces
- The 8 Linux Namespaces: PID (Process IDs), Mount (Filesystems), Net (Network routing), User (Rootless mapping), IPC, UTS, Cgroup, Time
- Managing CPU and Memory resource ceilings using cgroups v2 (`/sys/fs/cgroup`)
- Constructing a secure, rootless container environment using `unshare` and custom chroot/pivot_root
Introduction & Core Concept
Understanding cgroups v2 and namespaces allows you to debug container memory limit throttling, resolve Kubernetes OOMKilled errors, and secure multi-tenant cloud platforms.
Syntax & Structure
unshare --mount --uts --ipc --net --pid --fork --user --map-root-user chroot /container-root /bin/shecho "50000 100000" > /sys/fs/cgroup/mygroup/cpu.maxCreating an Isolated Linux Container from Scratch with unshare and cgroups v2
bash123456789101112131415161718192021222324252627282930313233343536#!/usr/bin/env bash# Building a Pure Linux Container from Scratch (No Docker required!)set -euo pipefailecho "=== Linux Container Architecture: Namespaces & cgroups v2 ==="# 1. Create a dedicated cgroups v2 resource restriction folderCGROUP_DIR="/sys/fs/cgroup/kwas_sandbox"sudo mkdir -p "$CGROUP_DIR"# Configure Memory Limit (100MB max before OOM Killer triggers)echo "104857600" | sudo tee "$CGROUP_DIR/memory.max" > /dev/null# Configure CPU Quota: 50,000us per 100,000us period (50% of 1 CPU core)echo "50000 100000" | sudo tee "$CGROUP_DIR/cpu.max" > /dev/nullecho "[1] Configured cgroups v2 limits: Max Memory = 100MB | CPU = 50% core"# 2. Launch an isolated process inside private PID, Mount, UTS, and Network Namespacesecho "[2] Spawning process inside isolated Namespaces via unshare..."sudo unshare --pid --mount --uts --net --fork bash -c '# Attach this container process PID to our cgroups v2 sliceecho $$ > /sys/fs/cgroup/kwas_sandbox/cgroup.procs# Set private container hostname (UTS Namespace)hostname "kwas-isolated-node-01"echo "Inside Container -> Hostname: $(hostname)"echo "Inside Container -> My PID: $$ (Visible as PID 1 inside namespace!)"echo "Inside Container -> Running under cgroup constraints."'# 3. Clean up cgroup slicesudo rmdir "$CGROUP_DIR"echo -e "\n✅ Container execution completed and cgroups cleaned up."
Line-by-Line Technical Breakdown
Try It Yourself (Interactive Editor)
Modify the code in real-time and click Run to test live browser output and console logs.
Common Mistakes & How to Avoid Them
#1: Setting Kubernetes CPU limits without understanding cgroup CFS throttling, resulting in severe latency spikes.
cgroup CPU limits enforce hard CFS quota slice throttling. If a multi-threaded process exhausts its quota within 10ms, it is frozen for the remaining 90ms of the period.
resources:
limits:
cpu: "500m" # Restricts process to 50ms per 100ms CFS period, causing periodic stallsresources:
requests:
cpu: "1000m" # Rely on requests or test CPU burstabilityIndustry Best Practices & Professional Standards
- Use cgroups v2 on all modern Linux distributions (Ubuntu 22.04+ default).
- Deploy Rootless Podman / Docker to prevent container escape privilege escalation.
- Use `pivot_root` instead of `chroot` for secure filesystem root isolation.
Lesson Summary & Core Takeaways
- Linux containers are normal processes governed by Namespaces and cgroups.
- Namespaces isolate system visibility (PID, Mount, Network, Hostname, User).
- cgroups v2 enforces strict CPU, Memory, and Disk I/O bandwidth boundaries.