Lobsters - 02 Oct 2026
Page 1 of 3
The other week I was looking into some Docker shenanigans, specifically SOCI (Seekable OCI). In a nutshell, SOCI builds an index of the contents of your container image layers, so that a container can start before the whole image has been downloaded, and the files it needs are lazily fetched as they are accessed. (Spoiler: there will probably be an AWS Bites episode about SOCI soon, so stay tuned there if you are curious.)
While reading about compressed layers, indexes and lazy loading, I started poking at my own understanding of Docker images. And I didn't love what I found.
I have been using Docker for years. I could happily tell you that "an image is made of a stack of immutable layers" and I would probably even draw you a nice diagram with some boxes stacked on top of each other. But if you asked me what those layers actually contain, my answer would get hand-wavy pretty quickly.
One question in particular got stuck in my head:
If Docker layers are basically tar archives applied on top of one another to produce a filesystem, how can one layer delete a file created by a previous layer?
Think about it for a second. A tar archive can say "here's a file called foo". A later tar archive can say "here's another version of foo". But tar doesn't have a generic operation that says "please delete foo from the archive that came before me".
And, of course, once I started asking that question, I had to find out.
Spoiler: the answer involves magic file names, "opaque" directories, a whole namespace of perfectly valid Linux file names that you can't faithfully put in a container image, and a Docker image that behaves differently after you export it and import it again. Let's jump down the rabbit hole together!
Tar: a brilliant, boring choice
Before we start digging, let me share a thought that has been bouncing around my head since I started this investigation.
Using tar as the foundation for image layers feels, in retrospect, like a brilliantly pragmatic engineering choice. Tar is:
- ubiquitous: every Unix-like system has tooling for it;
- simple: it's a sequence of entries, nothing fancy;
- streamable: you can process it as it arrives, without random access;
- extremely well understood: it has been around since the late 70s;
- decent at describing Unix filesystems: permissions, ownership, timestamps, symlinks, and more.
But tar was designed to describe a bunch of files. It was not designed to describe a diff between two filesystems.
My feeling (and I want to be clear that this is my interpretation, not a historical account of why the Docker and OCI folks made these choices) is that this is what often happens when you reuse an existing technology for something slightly beyond its original purpose: it works great, until you hit an impedance mismatch. And then you need a convention, or a workaround, to bridge the gap.
Whiteouts, which we are about to meet, feel like exactly one of those moments.
By the end of this article, I'd love for you to tell me whether you think this is an elegant extension of tar or a slightly hacky workaround. I can see arguments for both.
So what is a Docker image layer, really?
Let's start from something familiar:
You have probably heard that "every Dockerfile instruction creates a layer". That's not quite right. The steps that change the filesystem (like RUN, COPY and ADD) produce filesystem changes that end up as image layers, while other instructions (like ENV, CMD, EXPOSE or LABEL) only tweak the image configuration.
In fact, a container image is more than just filesystem data. Roughly speaking, it's made of:
- a manifest, which lists all the pieces that make up the image;
- an image configuration, with things like environment variables, the default command, the working directory, and so on;
- an ordered list of filesystem layers.
For this article, we mostly care about the last one: the filesystem layers.
For our Dockerfile, the stack of layers looks something like this:
The order matters. When a container runs, it doesn't see a folder called "Layer 0", another folder called "Layer 1" and so on. It sees the combined result of applying all the layers, one after the other, in order:
A quick note on vocabulary: since what really matters is the order in which layers are applied, in this article I'll talk about earlier and later layers. Elsewhere (including the OCI spec and the OverlayFS docs) you'll often see them called lower and upper layers, because they're usually drawn as a stack. Same thing, different metaphor.
The other important property is that layers are immutable. Once a layer is created, it never changes. A new layer doesn't edit a previous one: it describes another set of changes that gets applied after the previous ones.
This immutability is what makes a lot of the Docker magic possible:
- caching: if nothing changed, a build step can reuse the existing layer;
- sharing: 20 images based on alpine:3.20 can share the same base layer, on disk and over the network;
- content-addressed distribution: each layer is identified by the hash of its content, so registries and clients can tell whether they already have it.
OK, so far nothing new. But what's inside one of these layers?
A layer is a filesystem changeset
Here's where we need to be a little more precise than "a layer is a tar archive".
The OCI Image Specification (the standard that describes the format of container images, which Docker images follow these days) calls a layer an image layer filesystem changeset. When an image is distributed (for instance, when you push it to or pull it from a registry), each layer is represented as a tar-based changeset. The tar payload is usually compressed too, most commonly with gzip or zstd (the spec defines media types such as application/vnd.oci.image.layer.v1.tar+gzip and application/vnd.oci.image.layer.v1.tar+zstd).
This doesn't mean that Docker keeps a pile of .tar.gz files around and extracts them every time you start a container. Locally, the storage driver (for example, overlay2) keeps each layer as an extracted directory and uses a union filesystem to stack them. The tar representation is how layers are serialized and moved around. Keep this distinction in mind, because it's going to come back to bite us later!
From now on, I'll casually talk about "the layer tar", but remember: that's the serialized form, not necessarily what's sitting on your disk.
The key idea: A layer is not a complete filesystem. It is a filesystem changeset that only has meaning when applied after the layers that came before it.
So what does a changeset contain? According to the OCI spec, there are three types of change:
- Additions
- Modifications
- Removals
Additions and modifications are easy to represent in tar. Removals are the interesting one. But let's take things in order.
A tar archive is basically a sequential stream of entries. Each entry has a header with the path and some metadata, followed by the file contents (where applicable). Among the metadata that a layer entry carries, you'll find:
- the permissions (mode);
- the owner (UID and GID);
- the modification time;
- the link target, for symlinks and hard links;
- extended attributes, where supported.
So a layer tar might conceptually look like this:
When this changeset is applied on top of the previous filesystem:
- paths that don't exist yet are added;
- paths that already exist are replaced;
- directories that exist on both sides merge into one.
The spec actually makes a point of saying that layer changesets are applied, rather than simply extracted as tar archives. This sounds like a pedantic distinction right now, but hold that thought.
Adding files is the easy part
Let's say Layer A contains:
And Layer B contains:
The merged filesystem will be:
At the tar level, Layer B simply contains an entry for app/config.json (and, typically, one for the parent directory app/). No special trick needed. Tar already knows how to say "here's a file".
The same goes for directories. When a directory exists in an earlier layer and a directory with the same path shows up in a later layer, their contents compose into a single directory in the final filesystem. (If you are wondering: the spec says that the later directory's attributes, like permissions and ownership, replace those of the existing one. The children are merged, not replaced.)
So far so good. This sounds simple enough, right?
Changing a file doesn't change the old layer
What about modifying an existing file?
Let's say Layer 1 contains /app/config.json:
And Layer 2 contains another /app/config.json:
The container will see the Layer 2 version, with debug set to true.
This is how a modification is represented: the new layer simply ships a new, complete version of the file. The OCI spec is pretty explicit about it: "Additions and Modifications are represented the same in the changeset tar archive". There's no binary patch, no "change byte 12 from f to t". Changed one character in a 200 MB file? Congratulations, your new layer contains a brand new 200 MB file!
And, crucially, the original version of the file is still there, inside Layer 1. The newer layer doesn't touch the bytes of the older one (layers are immutable, remember?). It just provides an entry that takes precedence.
This has a practical consequence that you might have stumbled into:
The second RUN removes the file from the final filesystem, but it doesn't magically shrink the layer created by the first RUN. That layer is immutable and it still contains the whole file. This is why "I deleted the giant file in the next RUN instruction" doesn't save you from shipping the giant file's layer (and why you often see downloads, extraction and cleanup chained in a single RUN, or multi-stage builds).
But wait a second... we just said that the second RUN "removes the file". What does that layer actually contain?
How do you delete a file from a layer?
This is the question that sent me down the rabbit hole in the first place.
Let's try to recap what we know so far and see if we can come up with some kind of educated guess...
Let's say Layer 1 contains:
And we want the result after Layer 2 to be:
What should Layer 2 contain?
- It can't modify Layer 1, because layers are immutable.
- It can't just "not mention" old-config.json, because absence in a layer simply means "this layer has no change for that path". If absence meant deletion, every layer would have to list every single file in the filesystem to keep it alive, and we'd lose the whole point of layers.
- And tar has no "negative file". There's no entry type that means "the thing that used to be here, please make it go away".
Take a moment to think about how you would solve this. If you were designing the format, what would you do?
Take your time, I'll be here waiting...
Got an idea?
Great, now let's see how OCI does it.
Meet the whiteout
To remove /app/old-config.json, the newer layer contains an entry called:
That's it. I swear, that's the trick.
The .wh. prefix (short for whiteout) means: "when applying this layer, remove the path from earlier layers whose name is whatever follows .wh.".
Note that .wh.old-config.json is present in the layer archive, but it is not supposed to become a regular file in the final filesystem. It's an instruction encoded as a specially named tar entry. After the layer is applied:
- the earlier old-config.json is gone from the merged view;
- the whiteout itself is also hidden (the spec says: "Once a whiteout is applied, the whiteout itself MUST also be hidden").
Now, here's the part I find most interesting. Tar itself has no idea that .wh.old-config.json is special. As far as tar is concerned, it's just a regular file with a weird name. In fact, the spec even says so: "regardless of the path being deleted, the whiteout file is a regular file in the archive".
It's the OCI consumer, the thing that applies the layer, that looks at the name and says: "Ah! You don't actually want this file. You want me to remove old-config.json from an earlier layer."
Whiteouts are not a tar feature. They are an OCI convention encoded using specially named tar entries.
So that's what "applied, rather than simply extracted" means. If you just ran tar -xf on a layer, you would end up with a bunch of .wh.* files lying around, and nothing would be deleted.
So apparently we solved this one by inventing files that aren't really files.
I'll leave it to you to decide whether you think that's elegant or hacky. It surely is clever, but otherwise I have mixed feelings myself.
Anyway, a couple more rules from the spec are worth knowing, because they'll matter later:
- Whiteouts only apply to earlier layers. A whiteout can't delete a file that was added in the same layer. Quoting the spec: "Files that are present in the same layer as a whiteout file can only be hidden by whiteout files in subsequent layers."
- A .wh. entry with nothing after the prefix is invalid, and implementations "SHOULD return an error when encountering such an entry".
- The spec describes a whiteout as an empty file. Keep this one in your back pocket too.
What about directories?
The same mechanism works for directories. If an earlier layer contains:
A following layer containing:
removes the entire cache directory, along with everything inside it.
Nice and consistent. But sometimes we want something slightly different.
There's an even stranger whiteout
What if we don't want to delete the directory, but we want to say: "keep this directory, but forget about everything it inherited from earlier layers"?
This can happen, for instance, when a build step deletes a directory and recreates it with completely new contents. The directory still exists, but none of the old children should show up.
One option is to add a whiteout for every single child. That works, but imagine you have a large build folder with hundreds or even thousands of child files or folders (yes, like a node_modules ), it wouldn't be convenient to have to create a whiteout for each one of them, right?
In fact, OCI has a dedicated marker for this case, and it's the weirdest file name in this whole article:
Yes, that's .wh. twice, followed by .opq, which stands for opaque. An opaque whiteout inside a directory means: "for this directory, don't merge in the children inherited from earlier layers".
Let's see an example to better understand this concept.
Suppose we have a directory called /node_modules in an earlier layer, with two subdirectories:
Later layer:
Result:
The /node_modules directory stays, the earlier left-pad and event-stream directories disappear, and the new colors directory from the same layer as the marker survives. The spec also clarifies that the opaque marker is processed before the other entries of that directory in the same layer, regardless of the order in which they appear in the archive, so it only ever hides stuff from earlier layers.
(Fun fact: the spec says that implementations SHOULD generate layers using explicit per-file whiteouts, but MUST accept opaque ones.)
OK, at this point I felt pretty good about my new mental model. And then my brain did the thing it always does.
Wait... what if my file is actually called .wh.foo?
BTW, am I the weird one, or did your brain come up with the same question?
Anyway... Linux is perfectly happy with a file called .wh.foo:
There's nothing invalid about that name on a normal Unix filesystem. It's a hidden file (it starts with a dot) with a slightly odd name. Yes, but it's still a perfectly valid file, that's all.
But in an OCI layer, an entry called .wh.foo already means "delete foo from earlier layers". How would you tell the two apart?
- There's no flag in the tar entry saying this_is_a_literal_file = true.
- There's no escaping convention, like .wh.literal.wh.foo.
- There's no PAX header or extended attribute defined by OCI to distinguish "a regular file named .wh.foo" from "a whiteout for foo".
So the same bytes can only mean one thing to an OCI consumer. And the spec is very honest about the consequence:
As files prefixed with .wh. are special whiteout markers, it is not possible to create a filesystem which has a file or directory with a name beginning with .wh..
Let me restate that more precisely, because it's easy to overgeneralise. Linux doesn't forbid .wh.* names. Your laptop doesn't care. The limitation belongs to the OCI image layer representation: a serialized OCI image cannot faithfully represent a regular filesystem entry whose name begins with .wh.. The whole .wh.* namespace is effectively reserved:
- .wh.<name> means "delete <name>";
- .wh..wh..opq means "this directory is opaque";
- .wh. on its own is invalid.
So yes, saying "these are file names you can't use in a container image" is a bit of a shorthand. As we are about to discover, these names can exist in some places along the way. They just can't survive the trip through an OCI layer as regular files.
Of course I had to try it
Reading the spec answered the theoretical question. But now I had another one: what does Docker actually do if you try this?
Does docker build reject the file? Does it silently turn it into a whiteout? Does it escape it somehow? Does the file exist in a later RUN step? And in a container started from the image? What does the layer tar look like? What happens after exporting and re-importing the image? And what about .wh..wh..opq and the invalid .wh.?
At this point, of course, there was only one sensible thing to do: write some Dockerfiles and try to break them.
As the saying goes, "in theory there's no difference between theory and practice. In practice, there is." ...And, as we're about to find out, with Docker there can even be a difference between practice and practice after a docker load.
The result is a small repository with a very honest name: lmammino/broken-dockerfile. Its description sums it up: "It works on my machine. Then you push it." You can go and check it out... Hey, but only after you finish reading here!
Next page | More Lobsters | Headlines
Original: https://loige.co/hidden-design-compromises-of-docker-layers/