How Containers Actually Work: Namespaces, Cgroups, and Images

•6 min read•

Containers changed how software gets shipped, and Docker made them easy enough that millions of developers use them without ever asking what is underneath. The common mental model is that a container is a lightweight virtual machine, a small computer running inside your computer. That model is wrong in a way that matters. A container is not a separate machine at all. It is a normal process running directly on the host's Linux kernel, fenced off by a few features that make it believe it has the machine to itself.

Not a virtual machine

A virtual machine runs a full guest operating system on top of virtualized hardware[1]. There is a hypervisor, a complete second kernel, virtual disks, and real overhead. That isolation is strong, which is why you can run Windows inside Linux this way, but it is heavy and slow to start.

A container shares the host's kernel. There is no guest operating system and no virtualized hardware. When you run a process in a container, the host kernel is still the one scheduling it, handling its system calls, and managing its memory. The container is just that process, plus a set of restrictions on what it can see and use. This is why containers start in milliseconds and why you can run dozens of them on a laptop. It is also why a Linux container needs a Linux kernel, which is what the Linux virtual machine inside Docker Desktop quietly provides on a Mac or Windows box[2].

Namespaces: what the process can see

The first half of the disguise is Linux namespaces[3]. A namespace wraps a global system resource so that the processes inside it see their own private version of it. The kernel supports several kinds, and a container uses most of them at once.

The PID namespace gives the container its own process tree. The first process inside becomes PID 1, and it cannot see any of the host's other processes. The mount namespace gives it its own view of the filesystem. The network namespace gives it its own network interfaces, routing table, and ports, which is why two containers can both bind port 8080 without a fight. The UTS namespace lets it have its own hostname. The user namespace can map the root user inside the container to an unprivileged user on the host, so root in the container is not root on the machine.

You can see this with a plain kernel command, no Docker required.

# run a shell in its own PID and mount namespaces
unshare --pid --mount --fork bash
ps aux   # the process tree looks nearly empty

Docker is doing a more elaborate version of exactly this.

Cgroups: what the process can use

Namespaces control what a process sees. They do nothing to control how much it consumes. A container with no limits could still eat all the memory and every CPU cycle on the host and starve everything else.

That job belongs to control groups, or cgroups, the second half of the disguise. A cgroup is a kernel mechanism that measures and caps resource usage for a group of processes. You can tell a cgroup that this container gets at most half a gigabyte of memory and the equivalent of one and a half CPUs, and the kernel enforces it. Go over the memory limit and the kernel's out of memory killer steps in and terminates the process. When you pass --memory or --cpus to docker run, you are configuring a cgroup.

Namespaces plus cgroups are the heart of it. Isolation of view comes from namespaces, isolation of resources comes from cgroups, and both are features of the ordinary Linux kernel that have been there for years.

Images: layers stacked with a union filesystem

The last piece is the filesystem the container runs on top of, and this is where images come in. A container image is not a single blob. It is a stack of read-only layers, each one a set of filesystem changes, stacked on top of each other with a union filesystem that presents them as a single directory tree.

When you write a Dockerfile, each instruction that changes the filesystem produces a new layer.

FROM node:20          # base layer
WORKDIR /app
COPY package.json .   # a layer
RUN npm install       # a layer
COPY . .              # a layer

Those layers are the reason image builds and pulls are fast. Layers are content addressed and cached, so if your package.json has not changed, the expensive npm install layer is reused instead of rebuilt, and a layer shared by ten images is stored and downloaded once. When the container runs, the kernel adds a thin writable layer on top of the read-only stack, so changes the process makes during its life do not alter the image. Throw the container away and the image is untouched, ready to start an identical one.

Why "works on my machine" gets solved

The old problem was that software depended on its surroundings. The right version of a language runtime, a specific system library, an environment variable, a particular directory. Those things differed between a developer's laptop, the test server, and production, so code that ran in one place failed in another.

An image packages the process and its entire userland filesystem together, the libraries and the runtime and the files it expects, as one versioned, content addressed artifact. The same image runs on the laptop, in continuous integration, and in production, on the same shared kernel interface. The surroundings travel with the application instead of being assembled fresh and differently in each place. That is the real win, and it comes from the image format as much as from the isolation.

The takeaway

A container is a process the kernel has been told to show a private view of the system through namespaces, to limit through cgroups, and to run on top of a stack of read-only image layers. There is no second operating system and very little magic. Understanding this changes how you reason about the hard parts, because container security is really Linux kernel security, container performance is really host performance, and a container that misbehaves is really just a process you now know exactly how to inspect.

Sources (3)
  1. Wikipedia: Virtual machine
  2. Wikipedia: Docker (software)
  3. Wikipedia: Linux namespaces