Lobsters - 02 Oct 2026
Page 2 of 3
The repository contains a build context with files literally called foo, .wh.foo, .wh..wh..opq, .wh. and c. Each file contains some text naming itself (for example, .wh.foo contains CONTENT-OF-.wh.foo: I am a regular file literally named .wh.foo), so we can always tell what we are looking at. There are four Dockerfiles, one per experiment, and a script that pushes each one through a bunch of different paths:
- a regular docker build, with probes in RUN steps;
- a docker run of the resulting image;
- a docker image save, with a small Python script that lists every entry in the layer tars (type, mode, size, content, PAX headers);
- OCI and filesystem exports via docker buildx;
- and, most importantly, two ways of forcing a fresh unpack of the serialized layers: exporting the image and loading it back with docker image load, and exporting an OCI layout and using it as the base of a new BuildKit build.
To see what's going on, each Dockerfile has one or more small "probe" steps that list the directory and then check each file of interest:
The same check also runs inside containers started from the resulting images. So in the outputs below you'll see lines like these:
Printing the content too makes it obvious which file we're actually looking at. I also run the builds with --progress=plain, so the output of a build step is prefixed by BuildKit's step number and timing (for example #9 0.117).
I tested this with Docker Engine 29.4.0 and BuildKit (0.29.0 on the default builder, 0.32.2 on a temporary docker-container builder), running on OrbStack on an arm64 Mac, with the overlay2 storage driver and the containerd image store disabled. That last detail matters, and I'll come back to it. The repository contains the full version matrix, the scripts and all the raw outputs, so you don't have to take my word for any of this.
Let's see what happens.
BuildKit says "sure, why not?"
Let's start with the simplest possible case:
Drumroll... the build succeeds. No error. No warning. A later RUN step can see the file, content and all. And if I docker run the image I just built, the file is there:
So... did we just prove the OCI spec wrong?
Nope. So far, we have only built the image and run it on the same machine, straight from the files that BuildKit wrote to disk. We haven't yet forced the image to go through an OCI layer unpack: nothing has taken the serialized layer tars and applied them one by one, following the OCI rules. And that's exactly what happens when an image travels, for example when you export it and load it somewhere else.
The local snapshot and the layer archive are not quite the same thing
Remember when I said that the tar representation is how layers are serialized, not necessarily how they're stored locally? This is where that distinction stops being pedantic.
BuildKit doesn't work with tarballs internally. It works with filesystem snapshots. When it executes COPY .wh.foo /tmp/.wh.foo, it copies a regular file into a snapshot, and nothing about a snapshot cares about .wh. prefixes. So later RUN steps happily see the file.
When I build with the default Docker setup in my environment (overlay2, no containerd image store), the image ends up in overlay2 directories that BuildKit wrote directly. When I then docker run that image, the daemon stacks those directories and starts the container. At no point does anything take a serialized layer tar and apply it according to the OCI rules. Even docker image save reads from those directories and produces tar files: it writes layer tars, but it doesn't read them back.
So we actually have (at least) two different representations of "the same" layer:
They are supposed to describe the same thing. For almost every file name in the universe, they do. But for this pathological file name, their meaning can diverge.
I want to be careful here: this is what I observed in my environment. I did not test the containerd image store (which is the default on newer Docker installs), other snapshotters, or other storage drivers, and some of those could well behave differently locally. The part that is not environment-specific is what comes next: what the serialized layer means.
What does the layer tar look like?
Here's how BuildKit serialized that COPY, as listed by the inspection script (this one is from the second experiment, which we'll look at in a moment):
Look at that entry closely:
- it's a plain regular file (REG);
- it has the original mode (0644);
- it has the original content (64 bytes);
- there's no PAX header, no extended attribute, nothing that says "hey, I'm a literal file, not a whiteout".
But from the point of view of OCI layer semantics, tmp/.wh.foo is unambiguously a whiteout for tmp/foo. Remember the "whiteouts are empty files" rule I asked you to keep in your back pocket? This one has 64 bytes of content, and (spoiler) the unpackers I tested didn't care one bit. The name is all that matters.
In other words:
Same tar entry. Different semantic layer on top.
BuildKit managed to represent my local filesystem snapshot as tar bytes, but those same bytes acquire whiteout semantics as soon as someone interprets them as an OCI filesystem changeset. I think this is a beautiful example of the difference between a serialization format (tar) and the protocol that interprets it (OCI layer application).
Time to put on the lab coat and do some science!
Experiment 1: foo and .wh.foo in the same layer
The first Dockerfile copies both files in a single instruction, so they end up in the same layer:
During the build, and when running the locally built image, both files are there:
The serialized layer contains both entries, as regular files:
Now let's force a fresh unpack. To do that, I build the image again with a separate BuildKit instance (a docker-container builder), export it straight to a tarball instead of the local Docker image store, and then load that tarball into Docker:
Since the Docker daemon doesn't already have the layers created by our COPY, it has no choice but to unpack the serialized layer tars itself. (I also re-imported the image into BuildKit as an OCI layout, which forces a fresh unpack in a different way: check the repo for that one.) This is what we get:
.wh.foo is gone, but foo survived! Why?
Because whiteouts only apply to earlier layers. The unpacker saw tmp/.wh.foo, decided it was a whiteout, and (as the spec requires) did not materialise it as a file. But foo was added in the same layer as the whiteout, and a whiteout can't hide a sibling from its own layer. So foo stays.
So the only casualty was .wh.foo itself: foo survived because it lives in the same layer as the whiteout. But what if foo lived in an earlier layer instead?
Experiment 2: put foo in the previous layer and everything changes
This is the experiment that I consider the real smoking gun.
(The probe steps are the ones described earlier; check the repo for the exact Dockerfile.)
This time foo lives in an earlier layer, and .wh.foo comes later. During the build, both files exist:
docker run on the locally built image shows both files too. And if we peek into the overlay2 storage, .wh.foo is sitting there as an ordinary file in the later layer's directory, while foo lives in the earlier one:
Everything looks fine. It works on my machine!
Now let's force a fresh unpack of the serialized layers, exactly like we did before:
Both files are gone.
Here's exactly what happened:
- the earlier layer added /tmp/foo;
- the later layer was serialized with a regular tar entry called tmp/.wh.foo;
- the unpacker (both dockerd and BuildKit, in my tests) interpreted that entry as a whiteout;
- so it removed /tmp/foo coming from the earlier layer;
- and, as the spec requires, it did not materialise the whiteout itself;
- therefore, neither path exists.
Let that sink in for a moment.
A file that I never asked to delete (foo) is gone, deleted by a file that I simply asked to copy.
A filesystem state that BuildKit can represent internally is not necessarily a filesystem state that an OCI layer can faithfully serialize and reconstruct.
"It works on my machine" has rarely been this literal.
But wait, who exports images to tarballs anyway? The way most of us ship images is docker build followed by docker push. So I also tried exactly that: a plain docker build with the default builder, a docker push to a local registry, and then a docker pull + docker run on two brand new Docker daemons (running in docker:dind containers) that had never seen these layers. One of them used the containerd image store (the default on new installs) and the other one the classic overlay2 storage driver.
Same result on both: /tmp/foo and /tmp/.wh.foo are both gone. And the layer stored in the registry contains the very same regular-file entry tmp/.wh.foo that we saw before. So yes, it really works on my machine, then you push it, and it doesn't work anywhere else.
Experiment 3: opaque whiteouts, for real
Is this specific to .wh.foo? Let's try the opaque marker. The earlier layers create a directory with two files:
During the build (and on the locally built image), the marker is just a regular file and everything is visible:
The serialized layer, once again, contains a plain regular-file entry:
And after a fresh unpack:
That's the opaque whiteout doing exactly what the spec says:
- the children inherited from the earlier layer (a and b) are gone;
- c, which came in the same layer as the marker, survives;
- the marker itself is hidden.
So the collision isn't a quirk of one file name. It's the whole reserved namespace.
Experiment 4: the .wh. that shouldn't exist
Finally, the edge case of the edge cases: a file called just .wh., which the spec says is an invalid whiteout.
BuildKit doesn't reject it. The build succeeds, the image runs locally, and /tmp/.wh. shows up with its content. The layer tar contains a regular file entry tmp/.wh..
The trouble starts when something tries to unpack that layer. The BuildKit re-import fails, pretty explicitly:
And docker image load fails with a lower-level error:
I won't try to reverse-engineer that mknod error in detail here (more on character devices in a second), but it looks like the loader tried to turn the whiteout into an overlay whiteout for... the empty name, i.e. the parent directory itself, which already exists. The useful takeaway is simpler: BuildKit let me create a local image that can't subsequently be consumed as a valid OCI filesystem layer. Build succeeds, export succeeds, load fails.
(If you're curious, the BuildKit error comes from containerd's archive package, which checks that the target of a whiteout lives inside the whiteout's directory. .wh. targets the directory itself, so it fails that check.)
To be fair to BuildKit, the rule that makes .wh. invalid is very fresh. It was only added to the spec in May 2026 (PR #1314), after someone asked in image-spec#1301 what a bare .wh. is supposed to do. Until then, the spec simply didn't say, and different tools did different things: in that thread, one of the maintainers noticed that umoci treated it as an opaque whiteout! The ink is barely dry on this one.
For extra weirdness: the legacy builder
The repository also runs every experiment with the legacy builder (DOCKER_BUILDKIT=0), which you probably shouldn't be using in 2026 anyway, but it's an interesting comparison.
The legacy builder applies whiteout semantics much earlier: at COPY time, during the build itself. So:
- in experiment 1, .wh.foo disappears immediately and foo stays;
- in experiment 2, both foo and .wh.foo are gone already in the next RUN step;
- in experiment 3, a, b and the marker vanish, and c stays;
- in experiment 4, the build fails at the COPY step with the same mknod error we saw before.
The funny part is that docker image save then fails for every one of these images, with errors like:
So the legacy builder produces a filesystem that matches the OCI semantics from the start, but then trips over the missing file when it tries to serialize the image. Different parts of the Docker stack clearly disagree about when these names should acquire their special meaning. The repository has the complete matrix if you want the gory details.
OCI whiteouts are not OverlayFS whiteouts
In experiment 2 with the legacy builder, the overlay2 upper directory contained this:
That's not a .wh.foo file. It's a character device with device number 0/0, which is how the Linux OverlayFS represents a whiteout on disk (opaque directories, instead, are marked with an extended attribute). That's also likely why mknod(..., S_IFCHR, 0) showed up in the .wh. error: the loader was translating OCI whiteouts into overlay whiteouts.
This is a good reminder that the .wh.* convention belongs to the OCI layer format. A storage driver is free to represent "this file is deleted" in whatever way its filesystem supports, and translate between the two when importing or exporting layers. The whole mess we've seen happens at the boundary between those representations.
Is this a bug?
I'll admit that naming the repository broken-dockerfile was a bit provocative. So, let's be fair.
The OCI behaviour is not a bug. It's explicitly specified. The spec tells you that .wh.* names are whiteouts and that a filesystem containing them can't be represented. Every unpacker I tested did exactly what the spec says.
Whether BuildKit should reject or warn about literal .wh.* paths before producing an image that can't round-trip cleanly is a separate question. It's a tooling and API-design choice, and there may be good reasons (performance, compatibility, "garbage in, garbage out") for not checking every path.
The limitation itself is well known upstream. It has been discussed since before OCI 1.0: image-spec#24 ("Any chance of changing the whiteout file approach?"), opened in 2016 and still open, points out that with this scheme "base images can no longer contain arbitrary data". But I couldn't find any BuildKit issue about rejecting or warning on these names at build time (as of September 2026). And, reading the code, the layer writer that BuildKit uses (again, containerd's archive package) only generates .wh. names for deletions: added files are written under whatever name they have, with no check. So I don't want to claim that BuildKit is "broken". It's just a case that nothing currently guards against.
The way I'd put it is:
The format is doing exactly what the spec says. The surprising part is that the build pipeline lets us create a state whose serialized meaning is different.
Also, a couple of things I didn't test, which might be fun follow-ups: building with the containerd image store enabled (I only used it to pull), and creating the file from a RUN step (for example RUN touch /tmp/.wh.foo) rather than with COPY. If you try them, let me know what happens!
So what did we actually learn?
Let's zoom back out. Here's what I'm taking home from this rabbit hole.
1. Layers are changesets, not snapshots of the whole filesystem
Each layer only makes sense when applied after the ones that came before it. This is also why extracting a single layer tar somewhere doesn't give you the filesystem that a container sees (and why just running tar -xf on all of them isn't enough either: you'd need to apply the whiteouts).
2. Adding and modifying are easy. Deleting is not
Tar already knows how to carry files, directories and their metadata. Additions and modifications are just entries. Deletion is "negative" state, and tar has no concept of it, so OCI had to invent a convention: whiteouts.
"Easy", though, doesn't mean "efficient". A modification isn't optimised to save bytes in any way: change a single byte in a file, and the new layer contains the entire file again, with that one byte changed. There's no delta encoding, like the one git uses to pack objects or rsync uses to transfer files. Just a brand new full copy of the file.
3. Deleting doesn't delete bytes from old layers
A whiteout hides a file from the merged filesystem, but the old layer, and all its bytes, are still part of the image. Same for modifications: a new version of a file doesn't shrink the old one. This explains a lot of "why is my image so big?" moments.
It's also why, if you want to delete files to keep your image small (think package manager caches or temporary downloads), you need to do it in the same RUN instruction that creates them:
A layer only captures the difference between the filesystem before and after its build step. If a file is created and deleted within the same step, it's simply not there at the end, so it never makes it into any layer (and no whiteout is needed either). Delete it in the next RUN, instead, and you get a whiteout on top of a layer that still carries all those bytes.
4. .wh.* isn't just an odd implementation detail
It creates a genuine representational limitation. A perfectly valid Unix filesystem can contain .wh.foo, but an OCI image layer can't encode it as a regular file, because that name already has protocol-level meaning.
Thankfully, the naming scheme is awkward enough that you're very unlikely to ever give a real file a name starting with .wh.. At the very least, in over 10 years of using Docker, I have never bumped into an issue caused by this limitation (if you consider one person's experience a statistically significant sample, that is ).
5. The builder's internal state and the serialized image can differ
This was the most surprising discovery for me. In the setup I tested, BuildKit's snapshots (and the local image built from them) can contain states that the OCI layer format simply cannot round-trip. Everything works locally, and then the image means something different somewhere else.
6. Formats inherit the compromises of what they're built on
This is probably the most important takeaway from a systems design perspective. Tar was a pragmatic foundation. Whiteouts are the extra convention that lets a tar-based changeset express something tar itself was never designed to express. And conventions like that tend to have sharp edges in the corners.
Sure, whiteouts could have been built on PAX headers instead (as proposed back in 2016), which would arguably have been a better fit: a PAX record lives in the entry's metadata, so it wouldn't reserve any file names. But either way, it would still be a convention layered on top of tar, a workaround for something tar simply can't express on its own.
Next page | Previous page | More Lobsters | Headlines
Original: https://loige.co/hidden-design-compromises-of-docker-layers/